Method and device for automatically adjusting parameters of adaptive optical system based on reinforcement learning, storage medium and electronic equipment

By applying reinforcement learning technology in adaptive optical systems and automatically adjusting system parameters, the problem of slow response speed of manual adjustment is solved, and imaging quality and robustness are improved.

CN120125796AActive Publication Date: 2025-06-10SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202510194102.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

When facing rapidly changing atmospheric turbulence and target brightness, the parameter adjustment relies on manual experience and is slow to respond, so it is unable to adapt to changes in the external environment in time.

Method used

Using reinforcement learning technology, the characteristics of the wavefront sensor image and far-field image are extracted through the cross attention mechanism, and combined with the wavefront slope, turbulence intensity and wavefront corrector voltage, the current system status representation is constructed. Use the policy network to predict action parameters and optimize them in combination with the environment model and reward function to achieve automatic adjustment of AO system parameters.

Benefits of technology

Automatic optimization and adjustment of AO system parameters is achieved in complex dynamic environments, which significantly improves imaging quality and system robustness and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125796A_ABST
    Figure CN120125796A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for automatically adjusting parameters of an adaptive optical system based on reinforcement learning, a storage medium and electronic equipment, and relates to the field of reinforcement learning and adaptive optics, and the method comprises the steps that an AO system observes a target and collects related data; extracting features of the wavefront sensor image and the far-field image by using a cross attention mechanism; the extracted features and the collected data are spliced to form current system state representation; using a reinforcement learning strategy network and predicting action parameters according to the current state; predicting the next state based on the action of the strategy network and the current system state by combining the environment model, and calculating a corresponding reward value by using a reward function; and collecting a complete interaction track to update the strategy network and environment model. According to the technical scheme, the key parameters of the AO system can be automatically adjusted in a dynamic environment, the imaging quality of the system under complex conditions is improved, and the requirement for manual intervention is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of reinforcement learning and adaptive optics. Specifically, it relates to a method, device, storage medium, and electronic device for automatically adjusting the parameters of an adaptive optics system based on reinforcement learning. Background Art

[0002] Adaptive Optics (AO) technology uses optoelectronic devices to measure the dynamic error of the wavefront in real time, uses a fast electronic system for calculation and control, and uses active devices for real-time wavefront correction, enabling the optical system to automatically adapt to changes in external conditions and always maintain a good working state. It has important applications in high-resolution imaging observations and high-concentration laser energy transmission. The core goal of an adaptive optics system is to overcome the wavefront distortion caused by atmospheric turbulence, thereby improving the resolution and imaging quality of the optical system. It mainly consists of three core components: a wavefront sensor, a wavefront corrector, and a controller. The wavefront sensor is responsible for measuring the wavefront distortion and providing wavefront slope information for subsequent correction. The wavefront corrector can dynamically adjust the surface shape or refractive index of the optical element and is responsible for correcting the wavefront distortion. The controller calculates the optimal correction signal based on the data provided by the wavefront sensor. Currently, the settings of the wavefront sensor rely on manual operation or are based on pre-written rules. Although the imaging quality can be improved, the interference of atmospheric turbulence is highly non-linear and difficult to describe with an accurate mathematical model, making it difficult to quickly adapt to rapid changes in atmospheric conditions or target brightness.

[0003] To improve the imaging quality of an adaptive optics system, one of the main existing methods is to use deep learning technology. Traditional Shack-Hartmann sensors calculate the local slope of the wavefront by analyzing the position of the light spot on the microlens array. However, due to noise interference, errors may occur in slope calculation. To solve this problem, some methods propose using a convolutional neural network (CNN) to predict the slope from the wavefront image or directly predict the phase difference of the wavefront. In addition, for phase recovery, existing technologies have also achieved more efficient and accurate phase information reconstruction through the powerful non-linear fitting ability of machine learning models.

[0004] Although these deep learning methods can significantly improve the performance of an adaptive optics system, they still have some limitations. First, deep learning methods highly rely on the quality of training data and require a large amount of labeled data for training, which poses high requirements for data acquisition. Especially in complex environments such as atmospheric turbulence, obtaining sufficient high-quality labeled data is a challenge. In addition, although these methods can achieve good results on the training set, their generalization ability is poor when facing different environmental conditions (such as changes in turbulence intensity), and they often show a significant performance decline.

[0005] These problems have limited the popularization of deep learning methods in practical applications. Although they can improve the performance of AO systems under certain conditions, when the imaging effect of the system is poor, the adjustment relying on automated technology still cannot completely replace manual intervention. In scenarios with strong atmospheric turbulence or rapid changes in target brightness, manually adjusting the parameters of the AO system remains a common solution. However, in such cases, due to the rapid change of the environment, manual adjustment often fails to respond in a timely manner and cannot fully adapt to the rapidly changing external environmental conditions. Summary of the Invention

[0006] Embodiments of the present application provide a method, device, storage medium, and electronic device for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning to solve the problems that parameter adjustment in existing AO systems relies on manual experience and has a slow response speed. Through the introduction of reinforcement learning technology and the design of corresponding adjustment mechanisms, automatic optimization and adjustment of the parameters of the AO system are achieved, improving the imaging quality and system robustness in complex dynamic environments.

[0007] Other features and advantages of the present application will become apparent through the following detailed description or will be partially learned through the practice of the present application.

[0008] According to the first aspect of the embodiments of the present application, a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning is provided, including the steps of:

[0009] Observing the target and collecting relevant data, where the relevant data includes wavefront sensor images, wavefront slopes, turbulence intensities, wavefront corrector voltages, and far-field images;

[0010] Using a cross-attention mechanism to extract the features of wavefront sensor images and far-field images;

[0011] Concatenating the extracted features with the wavefront slope, turbulence intensity, and wavefront corrector voltage to form a representation of the current system state;

[0012] Using the policy network of reinforcement learning to predict action parameters based on the current state;

[0013] Combining with the environment model, predicting the next state based on the action of the policy network and the current system state and calculating the corresponding reward value using the reward function;

[0014] Collecting interaction trajectories to update the policy network and the environment model, thereby achieving continuous optimization.

[0015] In some embodiments of the present application, based on the foregoing solution, the observing the target and collecting relevant data are implemented through an AO system, where the collected relevant data is used to comprehensively represent the current environment.

[0016] In some embodiments of the present application, based on the foregoing solution, the action parameters represent all the parameters required in the control process of the AO system, and include the working frame rate and working gain of the wavefront sensor, the exposure time, gain, and number of restoration modes of the imaging camera.

[0017] In some embodiments of the present application, based on the foregoing solution, calculating the corresponding reward value using the reward function includes the following sub-steps:

[0018] Taking the Strehl ratio as the reward function and defining it in a selected calculation method in combination with the selected evaluation index.

[0019] In some embodiments of the present application, based on the foregoing solution, the interaction trajectory includes state, action, and reward information.

[0020] In some embodiments of the present application, based on the foregoing solution, the updating of the policy network and the environment model specifically includes the following sub-steps:

[0021] Using offline data update: Pre-training the policy network and the environment model using the offline dataset constructed from historical observation data.

[0022] Using online update: In actual operation, dynamically updating the policy network and the environment model based on the interaction trajectory collected in real time.

[0023] According to the second aspect of the embodiments of the present application, there is provided a device for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning, including:

[0024] A data acquisition unit for acquiring relevant data, where the relevant data includes wavefront sensor images, wavefront slopes, turbulence intensities, wavefront corrector voltages, and far-field images;

[0025] A feature extraction unit for extracting the features of the wavefront sensor images and far-field images using the cross-attention mechanism;

[0026] A feature splicing unit for splicing the extracted features with the wavefront slope, turbulence intensity, and wavefront corrector voltage to form the current system state representation;

[0027] An action output unit for predicting action parameters according to the current state using the policy network of reinforcement learning;

[0028] A state prediction unit for predicting the next state based on the action of the policy network and the current system state in combination with the environment model and calculating the corresponding reward value using the reward function;

[0029] An optimization unit for collecting the interaction trajectory to update the policy network and the environment model, so as to achieve continuous optimization.

[0030] According to a third aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer instructions are stored. When the computer instructions run on a computer, the computer is caused to execute the method described in the first aspect above.

[0031] According to a fourth aspect of the embodiments of the present application, there is provided an electronic device, including: a memory and a processor;

[0032] The memory is used for storing computer instructions;

[0033] The processor is used for calling the computer instructions stored in the memory, so that the electronic device executes the method described in the first aspect above.

[0034] The beneficial effects of the present application include:

[0035] The technical solution of the present application collects data such as wavefront sensor images, wavefront slopes, turbulence intensities, wavefront corrector voltages, and far-field images, and uses a cross-attention mechanism to extract image features to form an accurate system state representation. Based on this, in a further inventive concept, the reinforcement learning policy network can automatically predict and adjust AO system parameters, and continuously update and optimize the policy in combination with the environment model and the reward function, so that the system can continuously adaptively adjust under a rapidly changing environment, significantly improve the imaging quality under complex conditions, and effectively reduce the need for manual intervention.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0038] Figure 1 It shows a schematic flowchart of a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning according to an embodiment of the present application;

[0039] Figure 2 It shows a schematic diagram of implementing a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning using specific modules according to an embodiment of the present application;

[0040] Figure 3The block diagram of a device for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning according to an embodiment of the present application is shown;

[0041] Figure 4 The block diagram of an electronic device according to an embodiment of the present application is shown.

[0042] Figure 5 The schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. Detailed implementation manners

[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0044] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0045] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0046] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings.

[0049] To solve the technical problems existing in the prior art, the embodiments of the present application propose a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning. Specifically, to solve the problem that the manual adjustment of the parameters of the AO system cannot adapt to the rapidly changing external environmental conditions, the technical solution of the present application uses the method of reinforcement learning to encode information such as the sensor image, far-field image, and slope of the AO system as the state in the reinforcement learning environment, and uses the reinforcement learning agent to output actions, including the working frame rate, exposure time, camera gain, restoration mode number, etc., so as to replace the manual adjustment of the AO system parameters and ensure that it can respond to changes in the external environment in a timely manner.

[0050] See Figure 1 , which shows a schematic flow chart of a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning according to an embodiment of the present application.

[0051] As Figure 1 shown, a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning is presented, including steps S100 to S600.

[0052] Step S100, the AO system observes the target and collects relevant data, including wavefront sensor images, slopes, turbulence intensities, wavefront corrector voltages, and far-field images.

[0053] It can be understood that the target type of the observed target is a point source or an extended target, including simulated KJ targets, solar granules, etc.

[0054] In some feasible embodiments, based on the foregoing solution, the AO system observes the target and collects relevant data, where the collected relevant data is used to comprehensively represent the current environment, including but not limited to wavefront sensor images, slopes, turbulence intensities, wavefront corrector voltages, and far-field images.

[0055] Continuing to refer to Figure 1 , step S200, uses the cross-attention mechanism to extract the features of the wavefront sensor image and the far-field image.

[0056] It can be understood that in the embodiments of the present application, using the cross-attention mechanism to extract the features of the wavefront sensor image and the far-field image can fuse the information of the wavefront sensor image and the far-field image to achieve a more accurate modeling of the environmental state.

[0057] Continuing to refer to Figure 1 , step S300, splices the extracted features with the wavefront slope, turbulence intensity, and wavefront corrector voltage to form the current system state representation.

[0058] It should be noted that in some embodiments, the extracted features are not necessarily spliced with the wavefront slope, turbulence intensity, and wavefront corrector voltage. Instead, other collected parameters that can reflect the system state are selected, which depends on the task requirements and specific implementation effects.

[0059] Continuing to refer to Figure 1 , in step S400, the policy network using reinforcement learning predicts a set of action parameters based on the current state, including the working frame rate and working gain of the wavefront sensor, the exposure time and gain of the imaging camera, and the number of restoration modes.

[0060] It can be understood that in the embodiments of the present application, the input of the policy network is the current state of the system, and the output is the action to be taken currently.

[0061] It should be noted that in some embodiments, the specifically selected action parameters can be adjusted according to the task requirements, and are not limited to the working frame rate and working gain of the wavefront sensor, the exposure time and gain of the imaging camera, or the number of restoration modes.

[0062] Continuing to refer to Figure 1 , in step S500, in combination with the environment model, based on the action of the policy network and the current system state, the next state is predicted and the corresponding reward value is calculated using the reward function.

[0063] In some embodiments of the present application, based on the foregoing solution, the combination of the environment model, predicting the next state based on the action of the policy network and the current system state, and calculating the corresponding reward value using the reward function includes:

[0064] The Strehl ratio is used as the reward function.

[0065] In combination with other evaluation indicators, such as the score of the signal-to-noise ratio, it is defined in a weighted or other calculation manner.

[0066] It can be understood that the environment model represents the state transition function P(s′|s,a) of the system, which can predict the next state corresponding to the state and action parameters, and does not require actual execution of operations, which can accelerate policy optimization and ensure the safety of the system at the same time. The Strehl ratio is defined as the ratio of the peak light intensity of the actual optical system to the peak light intensity of the ideal diffraction-limited optical system, and is used to measure the wavefront quality of the optical system.

[0067] It should be noted that in some embodiments, the environment model is not necessary. Whether to use the environment model depends on the cost of actual interaction. When the cost of actual interaction is high, the environment model can be selected for simulation; otherwise, interaction can directly rely on the real environment. In addition, the reward function is not fixed. The core role of the reward function is to comprehensively evaluate the quality of actions to guide the optimization of the policy. Besides the Strehl ratio, other evaluation metrics such as the signal-to-noise ratio score can also be combined and defined in a weighted or other calculation manner.

[0068] Continue to refer to Figure 1 , step S600, collect a complete interaction trajectory, including state, action, and reward information, to update the policy network and the environment model, thereby achieving continuous optimization.

[0069] It should be noted that in the embodiments of this application, by executing step S100 to step 500, a tuple (s t , a t , r t , s t+1 ) at time t can be obtained, where s t represents the state at time t, a t represents the action taken at time t, and r t represents the reward obtained at time t. By repeatedly executing step S100 to step S500 in chronological order, multiple tuples at different times can be continuously obtained, ultimately forming a complete interaction trajectory.

[0070] In some embodiments of this application, based on the foregoing solution, the collection of the complete interaction trajectory, including state, action, and reward information, to update the policy network and the environment model, thereby achieving continuous optimization includes:

[0071] Use offline data for update: Use the offline dataset constructed from historical observation data to pre-train the policy network and the environment model.

[0072] Use online update: In actual operation, dynamically update the policy network and the environment model based on the real-time collected interaction trajectory.

[0073] In summary, the technical solution of this application proposes an adaptive control scheme for AO system parameters based on reinforcement learning. This scheme can automatically adjust the key parameters of the AO system in a dynamic environment, improve the imaging quality of the system under complex conditions, and significantly reduce the need for manual intervention. At the same time, the present invention combines the cross-attention mechanism with reinforcement learning to fully explore the correlation between data, further improving the efficiency and accuracy of optimization.

[0074] Specifically, an example of implementing the method of this application using specific modules is provided below.

[0075] See Figure 2 , which shows a schematic diagram of a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning using specific modules according to an embodiment of the present application.

[0076] As Figure 2 shown, first, the relevant data collected is encoded for the state. The collected data includes but is not limited to wavefront sensor images, slopes, turbulence intensities, wavefront corrector voltages, and far-field images. The state encoding module is used to encode the collected data to generate a state representation of the system. First, the wavefront sensor image and the far-field image are divided into small patches to reduce the time complexity when calculating attention and ensure computational efficiency. Then, the cross-attention between the wavefront sensor image and the far-field image is calculated to extract their relevant feature information. Finally, the extracted image features are concatenated with non-image data such as slopes, turbulence intensities, and wavefront corrector voltages to generate a state representation of the current environment, that is, s t , for subsequent processing.

[0077] As Figure 2 shown, the state s t of the current environment has three different uses. First, it is input into a pre-set reward function to calculate the reward r t of the current state to evaluate the correction effect. Second, the state s t is used as the input of the policy network module π(s t ), and the output action a t is obtained, that is, the AO system parameters that need to be adjusted, including frame rate, gain, number of restoration modes, etc. Finally, the state s t and the action a t act together with the state update module, that is, the environment model. The specific process can be expressed by the following formula:

[0078] s t+1 = f(s t , a t ) + ∈

[0079] f(s t , a t ) represents the environment model, which takes the state s t and the action a t at time t as inputs and predicts the next state s t+1 . Among them, ∈ represents a random noise term used to simulate the uncertainty in the real system.

[0080] Figure 2 The state s t , action a t , reward r t and next state s t+1Data will be collected. A complete set of data collected at consecutive time intervals is called a trajectory, which is used to update the policy network.

[0081] The following describes the device embodiments of the present application, which can be used to execute a method for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application above.

[0082] Refer to Figure 3 As shown in

[0083] A device 300 for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning according to an embodiment of the present application includes: a data acquisition unit 301, a feature extraction unit 302, a feature splicing unit 303, an action output unit 304, a state prediction unit 305, and an optimization unit 306.

[0084] Refer to Figure 4 As shown in

[0085] Since the electronic device introduced in this embodiment is the device used in the apparatus for automatically adjusting the parameters of an adaptive optical system based on reinforcement learning in the embodiments of the present application, based on the method introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of the present application falls within the scope of protection of the present application.

[0086] In the specific implementation process, when the computer program 411 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0087] Figure 5 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0088] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0089] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503, such as executing the method described in the above embodiments. In the RAM 503, various programs and data required for system operation are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0090] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0091] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, various functions defined in the system of the present application are performed.

[0092] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0094] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation on the units themselves in some cases.

[0095] As another aspect, the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for automatically adjusting the parameters of the adaptive optical system based on reinforcement learning described in the above embodiments.

[0096] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by an electronic device, the electronic device implements the method for automatically adjusting the parameters of the adaptive optical system based on reinforcement learning described in the above embodiments.

[0097] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0098] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented in software or in the form of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0099] Other embodiments of the present application will be readily contemplated by those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for automatically adjusting parameters of an adaptive optical system based on reinforcement learning, characterized in that: Includes steps: Observe a target and collect relevant data, wherein the relevant data includes a wavefront sensor image, a wavefront slope, a turbulence intensity, a wavefront corrector voltage, and a far-field image; The features of wavefront sensor image and far-field image are extracted using cross-attention mechanism; The extracted features are spliced ​​with the wavefront slope, turbulence intensity and wavefront corrector voltage to form a representation of the current system state; A policy network using reinforcement learning predicts action parameters based on the current state; Combined with the environment model, the next state is predicted based on the action of the policy network and the current system state, and the corresponding reward value is calculated using the reward function; Collect interaction trajectories to update the policy network and environment model for continuous optimization.

2. The method according to claim 1, characterized in that The observation of the target and collection of relevant data are achieved through the AO system, wherein the collected relevant data are used to comprehensively describe the current environment.

3. The method according to claim 2, characterized in that The action parameters represent all parameters required in the AO system control process, and include the working frame rate and working gain of the wavefront sensor, the exposure time, gain and number of restoration modes of the imaging camera.

4. The method according to claim 1, characterized in that The method of calculating the corresponding reward value using the reward function includes the following sub-steps: The Strehl ratio is used as a reward function and is defined in a selected calculation method in combination with a selected evaluation metric.

5. The method according to claim 4, characterized in that The interaction trajectory includes state, action and reward information.

6. The method according to claim 5, characterized in that The updating strategy network and environment model specifically includes the following sub-steps: Update using offline data: Use offline datasets constructed using historical observation data to pre-train the policy network and environment model. Use online updates: In practice, the policy network and environment model are dynamically updated based on interaction trajectories collected in real time.

7. A device for automatically adjusting parameters of an adaptive optical system based on reinforcement learning, characterized in that: include: A data acquisition unit, used to acquire relevant data, wherein the relevant data includes a wavefront sensor image, a wavefront slope, a turbulence intensity, a wavefront corrector voltage, and a far-field image; A feature extraction unit, for extracting features of the wavefront sensor image and the far-field image using a cross-attention mechanism; A feature concatenation unit, used for concatenating the extracted features with the wavefront slope, the turbulence intensity and the wavefront corrector voltage to form a current system state representation; An action output unit, which is used to predict action parameters based on the current state using a reinforcement learning policy network; The state prediction unit is used to combine the environment model, predict the next state based on the action of the policy network and the current system state, and calculate the corresponding reward value using the reward function; The optimization unit is used to collect interaction trajectories to update the policy network and environment model to achieve continuous optimization.

8. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 6.

9. An electronic device, characterized in that: Including: memory and processor; The memory is used to store computer instructions; The processor is used to call the computer instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Camera adaptive adjustment method and device and camera

    CN108803328A

  • CPS system reinforcement learning control method based on attention mechanism

    CN114527666A

  • Automatic focusing method and system for electro-hydraulic focusable lens, and electronic equipment

    CN117156272A

  • Cross-layer routing method based on reinforcement learning agent exploration optimization and related equipment

    CN117354227A

  • Self-adaptive image signal processing method and device and storage medium

    CN118096491A

Cited By

  • Adaptive optical dynamic modeling and control method and device based on model

    CN121480603A