Optical anti-shake compensation calculation method and device and medium
By introducing a reinforcement learning model in the optical anti-shake system, and using gyroscope data and lens position to construct the Markov decision-making process, the problem of poor adjustment of lens compensation parameters in the prior art is solved, and high-quality imaging is achieved under different vibration environments.
Patent Information
- Application Number
- CN202510102417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
The existing optical anti-shake technology has poor problems in adjusting lens compensation parameters, which makes it difficult to ensure the imaging quality under different vibration environments.
Using reinforcement learning model, the Markov decision-making process is constructed by collecting gyroscope data and lens position, and the model is trained using a proximal strategy optimization algorithm, lens compensation parameters are generated, and lens adjustments are adjusted to achieve the best compensation effect.
Effectively reduce image blur and jitter, improve image stability, adapt to complex and changeable vibration environments, and ensure high-quality imaging effects.
Smart Images

Figure CN119946431A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical image stabilization, and in particular relates to an optical image stabilization compensation calculation method, device and medium. Background Art
[0002] With the development of the times, optical image stabilization (OIS) technology has been widely used in many fields. Smartphones, cameras, and various optical instruments are inseparable from the support of OIS technology. Its core function is to accurately compensate for the jitter detected by the gyroscope through the lens, thereby effectively reducing image blur and significantly improving image quality.
[0003] Existing optical image stabilization technology is mainly based on classical calculation methods. In this mode, its performance is highly dependent on manually set control parameters. For a long time, this method has exposed many disadvantages in practical applications, specifically:
[0004] 1. Parameter adjustment is extremely complex. The gain, threshold and other parameters in traditional control methods need to be manually adjusted according to different devices and diverse usage scenarios. This process often takes a lot of time and the debugging cycle is long, resulting in extremely low work efficiency.
[0005] 2. Poor adaptability: When faced with different vibration frequencies and amplitudes, fixed parameters are difficult to flexibly adapt to dynamic scenes. This makes it difficult to achieve the expected compensation effect and ensure the imaging quality in actual use.
[0006] 3. The ability to handle nonlinear and complex environments is very limited. When traditional control methods process nonlinear vibration signals, over-compensation or under-compensation problems are very likely to occur, which not only affects the stability of imaging, but may also cause serious distortion of the image.
[0007] In recent years, with the continuous innovation of technology, data-driven methods such as deep learning and reinforcement learning have gradually emerged and are widely used in fields such as industrial control and image processing. With their powerful self-learning capabilities, these emerging technologies have far surpassed traditional methods in complex environments. However, the introduction of reinforcement learning into OIS system optimization is still an emerging technology. In particular, there are still problems to be solved in the adjustment of lens compensation parameters. Summary of the invention
[0008] The technical problem solved by the present invention is to provide an optical image stabilization compensation calculation method, device and medium to solve the problem of poor adjustment of lens compensation parameters in the existing optical image stabilization technology.
[0009] The basic solution provided by the present invention is an optical image stabilization compensation calculation method, comprising:
[0010] S1: Collect the real-time data of the gyroscope in different environments at the current moment and generate a sample data set;
[0011] S2: Build a reinforcement learning model and define the state, action and reward function of the reinforcement learning model by modeling the optical image stabilization compensation calculation process as a Markov decision process;
[0012] S3: Input the sample data set into the reinforcement learning model, use the proximal strategy optimization algorithm to train the reinforcement learning model, and output the trained reinforcement learning model;
[0013] S4: Deploy the trained reinforcement learning model to the MCU, output the lens compensation parameters, and adjust the lens through the OIS driver according to the lens compensation parameters generated by the reinforcement learning model.
[0014] Further, the S2 includes:
[0015] S2-1: Build a reinforcement learning model and perform quantification based on the definition of the current environment state, the collection of real-time data from the gyroscope and the range of compensation actions that the lens can perform;
[0016] S2-2: Extract the angular velocity and angular acceleration from the real-time gyroscope data, extract the lens position and historical SR value, and determine the state vector of the reinforcement learning model. The expression is:
[0017] s t =[ω t ,α t ,p t ,SR t-1 ]
[0018] Among them, s t represents the state vector, ω t represents the angular velocity of the gyroscope, α t is the angular acceleration of the gyroscope, p t Indicates the lens position, SR t-1 Indicates the historical SR value;
[0019] S2-3: Extract the lens displacement adjustment amount and determine the execution action of the reinforcement learning model. The expression is:
[0020] a t =p t -p t-1
[0021] S2-4: Extract the quantitative results of the range of compensation actions that can be performed by the lens and determine the reward function, which is expressed as:
[0022] r t =SRt -SR t-1
[0023]
[0024] Among them, r t represents the reward function, SR t represents the SR value at time t, SR t-1 represents the SR value at time t-1, W off Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned off; W on Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned on; W static Indicates the line width calculated from a static image captured when the optical image stabilization algorithm is turned off.
[0025] Further, the S3 includes:
[0026] S3-1: Based on the optical image stabilization algorithm scenario, the proximal strategy optimization algorithm is used to design the reinforcement learning model processing logic;
[0027] S3-2: Input the collected sample data into the reinforcement learning model for training, and continuously optimize the strategy until the reinforcement learning model with the optimal strategy is output.
[0028] Further, the S3-1 includes:
[0029] S3-1-1: Based on the policy function from the current state s t Predict action a t ;
[0030] S3-1-2: Use the value function to evaluate the current state s t of income;
[0031] S3-1-3: Calculate the total revenue by accumulating the reward discounts. The expression is:
[0032]
[0033] Where γ represents the discount factor;
[0034] S3-1-4: The near-end strategy optimization algorithm is used to optimize the optical image stabilization compensation strategy. The target loss function of the near-end strategy optimization algorithm is:
[0035] L PPO (θ) = E t [min(r t (θ)·A t ,clip(r t (θ),1-∈,1+ε)·At )]
[0036]
[0037] A t =R t -V(s t )
[0038]
[0039] Among them, r t (θ) represents the strategy probability ratio, A t represents the advantage function, ∈ represents the cut threshold for limiting strategy changes, V(s t ) is the state value function.
[0040] Further, the S3-2 includes:
[0041] S3-2-1: Build and initialize the training environment, including setting the amplitude and frequency of vibration to simulate vibration scenarios, adjusting the lens using the OIS driver, real-time gyroscope data, and capturing lens images;
[0042] S3-2-2: Obtain the state vector, execution action and reward data in the static environment and vibration environment of the training environment, including the line width W off , line width W static , line width W on , angular velocity ω t , angular acceleration α t , lens position p t ;
[0043] S3-2-3: Input the state vector into the reinforcement learning model;
[0044] S3-2-4: According to the current strategy π θ (a|s) select action a t , and update the lens position p t =p t-1 +a t To generate lens compensation parameters;
[0045] S3-2-5: Capture an image according to the compensated lens and calculate a new SR value of the image;
[0046] S3-2-6: According to the reward function r t =SR t -SR t-1 Update reinforcement learning model parameters;
[0047] S3-2-7: Repeat the above steps until the reinforcement learning model converges, the optimal strategy is obtained, and the trained reinforcement learning model is output.
[0048] Further, the S4 includes:
[0049] S4-1: Deploy the trained reinforcement learning model to the MCU for real-time reasoning;
[0050] S4-2: Receive gyroscope data through MCU, and calculate the current state of the lens according to the current state of the lens. t Call the reinforcement learning model to infer the target lens position compensation action a t ;
[0051] S4-3: Compensate motion a according to target lens position through OIS driver t Adjust the lens position.
[0052] An electronic device comprises a processor and a memory, wherein the processor stores programs or instructions, and the processor executes any one of the above-mentioned optical image stabilization compensation calculation methods by calling the programs and instructions stored in the memory.
[0053] A computer-readable storage medium stores a program or instruction, wherein the program or instruction enables a computer to execute an optical image stabilization compensation calculation method as described in any one of the above items.
[0054] The principle and advantage of the present invention are as follows: the technical solution of the present invention realizes optical image stabilization (OIS) compensation calculation based on the proximal policy optimization algorithm (PPO) in reinforcement learning, and its core is to construct the OIS compensation process as a Markov decision process. In this process, the system state is composed of a state vector composed of gyroscope data, current lens position and historical SR values, and these data fully reflect the vibration of the device and the current imaging state; and the policy function in the reinforcement learning model generates an action of the lens displacement adjustment amount according to the state vector, and the action is converted to determine the target compensation position of the lens, so as to adjust the lens through the OIS driver;
[0055] The reward function is based on the SR value increment. If the compensation causes the SR value to increase, indicating that the image clarity has improved, a positive reward is given, otherwise a negative reward is given, which provides key feedback for model training. During the training phase, the PPO algorithm uses the policy function, value function, reward discount, and specific loss function to continuously optimize the model parameters. The policy function predicts the action from the state, the value function evaluates the long-term benefits of the state to guide the policy optimization, the reward discount accumulates the total benefits to balance the long-term and short-term rewards, and the loss function prevents the policy update from being too large to ensure stable learning. Through continuous iterative training in a simulated vibration environment, the model learns the optimal compensation strategy, and finally realizes the precise adjustment of the lens position according to the real-time vibration and imaging conditions in the actual device to improve image stability.
[0056] Therefore, the advantages of this solution are:
[0057] 1. Through the deep optimization of lens compensation strategy through reinforcement learning, it can effectively reduce image blur and jitter in a vibrating environment, greatly improve image stability, and ensure that the captured images are clearer and sharper, providing users with high-quality imaging effects;
[0058] 2. Dynamically adjust the compensation position based on real-time gyroscope data, which can flexibly respond to various complex and changeable vibration environments, effectively overcome the disadvantage that traditional OIS methods are difficult to adapt to changes in vibration conditions, and maintain good anti-shake performance in different scenarios;
[0059] 3. The system can update the anti-shake compensation in real time, getting rid of the dependence on fixed preset parameters. The compensation strategy is adjusted instantly according to the actual imaging effect to ensure that the vibration effect is always compensated in the best state, avoiding the problem of unclear images caused by fixed parameters, and greatly improving the user's shooting experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of an embodiment of the present invention;
[0061] Figure 2 The figure is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following is further described in detail through specific implementation methods:
[0063] The symbols in the drawings of the specification include: electronic device 400 , processor 401 , memory 402 , input device 403 , and output device 404 .
[0064] The embodiment is basically as shown in the attached Figure 1 As shown: An optical image stabilization compensation calculation method, comprising:
[0065] S1: Collect the real-time data of the gyroscope in different environments at the current moment and generate a sample data set;
[0066] In this embodiment, the gyroscope data collected is a device with OIS technology, such as a smart phone, a camera, an optical instrument, etc., which can compensate for the jitter detected by the gyroscope through the lens according to the OIS technology, thereby reducing image blur and improving image quality. Therefore, the real-time data of the gyroscope at the current moment is collected, including data in static environments and vibration environments, including when the optical image stabilization algorithm of the lens is turned off, the image is collected under vibration conditions and the line width W is calculated off ; When the optical image stabilization algorithm of the lens is turned on, the line width W is calculated by collecting images under vibration conditions on ; The line width W calculated from the image captured in a static state when the optical image stabilization algorithm is turned off static , angular velocity ω t , angular acceleration α t , lens position p t Etc., these data provide convenience for the subsequent training of reinforcement learning models; at the same time, the real-time data of the above-mentioned gyroscope is the real-time data at the current moment, that is, the data collected during the subsequent model training process is also collected in real time, making the sample data set real-time.
[0067] S2: Construct a reinforcement learning model and define the state, action and reward function of the reinforcement learning model by modeling the optical image stabilization compensation calculation process as a Markov decision process; S2 includes:
[0068] S2-1: Build a reinforcement learning model, and perform the definition based on the current environment state, collect the real-time data of the gyroscope, and quantify the range of compensation actions that the lens can perform; wherein, the definition based on the current environment state is to define the vibration state and static state of the current environment by setting the amplitude and frequency of the vibration, and the real-time data of the gyroscope includes angular velocity and angular acceleration;
[0069] S2-2: Extract the angular velocity and angular acceleration from the real-time gyroscope data, extract the lens position and historical SR value, and determine the state vector of the reinforcement learning model. The expression is:
[0070] s t =[ω t ,α t ,p t ,SR t-1 ]
[0071] Among them, s t represents the state vector, ω t represents the angular velocity of the gyroscope, α t is the angular acceleration of the gyroscope, p t Indicates the lens position, specifically the compensation position of the lens, expressed in two-dimensional coordinates, SR t-1Represents the historical SR value, which is used to assist the model in learning to compensate for the impact on image clarity;
[0072] S2-3: Extract the lens displacement adjustment amount and determine the execution action of the reinforcement learning model. The expression is:
[0073] a t =p t -p t-1
[0074] Among them, action a t It is the control instruction output by the reinforcement learning model, which indicates the displacement increment or adjustment amount of the lens at the current time t, and adopts a two-dimensional continuous value;
[0075] S2-4: Extract the quantitative results of the range of compensation actions that can be performed by the lens and determine the reward function, which is expressed as:
[0076] r t =SR t -SR t-1
[0077]
[0078] Among them, r t represents a reward function with the following properties:
[0079] If the lens compensation improves the image clarity, the reward is positive; if the compensation is ineffective or results in a blurry image, the reward is negative;
[0080] SR t represents the SR value at time t, SR t-1 represents the SR value at time t-1, W off Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned off; W on Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned on; W static Represents the line width calculated from the image captured in a static state when the optical image stabilization algorithm is turned off, where W off and W static To obtain before training, W on is calculated during the training process.
[0081] S3: Input the sample data set into the reinforcement learning model, use the proximal strategy optimization algorithm to train the reinforcement learning model, and output the trained reinforcement learning model; S3 includes:
[0082] S3-1: Based on the optical image stabilization algorithm scenario, the near-end strategy optimization algorithm is used to design the reinforcement learning model processing logic; S3-1 includes:
[0083] S3-1-1: Based on the policy function from the current state s t Predict action a t ;
[0084] S3-1-2: Use the value function to evaluate the current state s t of income;
[0085] S3-1-3: Calculate the total revenue by accumulating the reward discounts. The expression is:
[0086]
[0087] Where γ represents the discount factor;
[0088] S3-1-4: The near-end strategy optimization algorithm is used to optimize the optical image stabilization compensation strategy. The target loss function of the near-end strategy optimization algorithm is:
[0089] L PPO (θ) = E t [min(r t (θ)·A t ,clip(r t (θ),1-∈,1+ε)·A t )]
[0090]
[0091] A t =R t -V(s t )
[0092]
[0093] Among them, r t (θ) represents the strategy probability ratio, A t represents the advantage function, ∈ represents the cut threshold for limiting strategy changes, V(s t ) is the state value function.
[0094] S3-2: Input the collected sample data into the reinforcement learning model for training, and continuously optimize the strategy until the reinforcement learning model with the optimal strategy is output. S3-2 includes:
[0095] S3-2-1: Build and initialize the training environment, including setting the amplitude and frequency of vibration to simulate vibration scenarios, adjusting the lens using the OIS driver, real-time gyroscope data, and capturing lens images;
[0096] S3-2-2: Obtain the state vector, execution action and reward data in the static environment and vibration environment of the training environment, including the line width W off , line width W static , line width W on , angular velocity ω t , angular acceleration α t , lens position p t ;
[0097] S3-2-3: Input the state vector into the reinforcement learning model;
[0098] S3-2-4: According to the current strategy π θ (a|s) select action a t , and update the lens position p t =p t-1 +a t To generate lens compensation parameters;
[0099] S3-2-5: Capture an image according to the compensated lens and calculate a new SR value of the image;
[0100] S3-2-6: According to the reward function r t =SR t -SR t-1 Update reinforcement learning model parameters;
[0101] S3-2-7: Repeat the above steps until the reinforcement learning model converges, the optimal strategy is obtained, and the trained reinforcement learning model is output.
[0102] Therefore, the core of this solution is to construct the OIS compensation process as a Markov decision process. In this process, the system state is composed of a state vector composed of gyroscope data, current lens position, and historical SR values. These data fully reflect the vibration of the device and the current imaging state. The policy function in the reinforcement learning model generates an action for the lens displacement adjustment amount based on this state vector. The action is converted to determine the target compensation position of the lens, thereby adjusting the lens through the OIS driver.
[0103] The reward function is based on the increment of SR value. If the compensation makes the SR value increase, indicating that the image clarity has improved, a positive reward is given, otherwise a negative reward is given, which provides key feedback for model training. During the training phase, the PPO algorithm uses the policy function, value function, reward discount, and specific loss function to continuously optimize the model parameters. The policy function predicts the action from the state, the value function evaluates the long-term benefits of the state to guide the policy optimization, the reward discount accumulates the total benefits to balance the long-term and short-term rewards, and the loss function prevents the policy update from being too large to ensure stable learning. Through continuous iterative training in a simulated vibration environment, the model learns the optimal compensation strategy.
[0104] S4: deploying the trained reinforcement learning model to the MCU, outputting lens compensation parameters, and adjusting the lens through the OIS driver according to the lens compensation parameters generated by the reinforcement learning model, wherein S4 includes:
[0105] S4-1: Deploy the trained reinforcement learning model to the MCU for real-time reasoning;
[0106] S4-2: Receive gyroscope data through MCU, and calculate the current state of the lens according to the current state of the lens. t Call the reinforcement learning model to infer the target lens position compensation action a t ;
[0107] S4-3: Compensate motion a according to target lens position through OIS driver t Adjust the lens position.
[0108] Therefore, the technology of this application improves the adaptability and real-time optimization capabilities of the system by introducing a reinforcement learning algorithm into the OIS system, and has a wide range of application potential. It is suitable for fields that require image stability and high-precision imaging, such as smart phones, digital cameras, sports cameras and other devices, and can improve shooting clarity and optimize user experience. Moreover, the reinforcement learning model of this solution is deployed in the MCU to only perform forward reasoning, and has the characteristics of low power consumption and high efficiency. At the same time, this solution can be extended to more complex vibration modes, such as nonlinear vibration, random vibration, etc.; the training and deployment of the reinforcement learning model can adapt to different vibration devices and lens control solutions, and has good versatility and portability.
[0109] like Figure 2 As shown, in another embodiment of the present embodiment, an electronic device is also included, and the electronic device 400 includes one or more processors 401 and a memory 402.
[0110] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.
[0111] The memory 402 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may run the program instructions to implement an optical image stabilization compensation calculation method and / or other desired functions of any embodiment of the present invention described above. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.
[0112] In one example, the electronic device 400 may further include: an input device 403 and an output device 404, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 403 may include, for example, a keyboard, a mouse, etc. The output device 404 may output various information to the outside, including early warning information, braking force, etc. The output device 404 may include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.
[0113] Of course, to simplify, Figure 2 Only some of the components related to the present invention in the electronic device 400 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application conditions, the electronic device 400 may also include any other appropriate components.
[0114] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of an optical image stabilization compensation calculation method provided by any embodiment of the present invention.
[0115] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present invention, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0116] In addition, an embodiment of the present invention may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps of an optical image stabilization compensation calculation method provided by any embodiment of the present invention.
[0117] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. Ordinary technicians in the relevant field know all the common technical knowledge in the technical field to which the invention belongs before the application date or priority date, can obtain all the existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the relevant field can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the relevant field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. An optical image stabilization compensation calculation method, characterized in that: include: S1: Collect the real-time data of the gyroscope in different environments at the current moment and generate a sample data set; S2: Build a reinforcement learning model and define the state, action and reward function of the reinforcement learning model by modeling the optical image stabilization compensation calculation process as a Markov decision process; S3: Input the sample data set into the reinforcement learning model, use the proximal strategy optimization algorithm to train the reinforcement learning model, and output the trained reinforcement learning model; S4: Deploy the trained reinforcement learning model to the MCU, output the lens compensation parameters, and adjust the lens through the OIS driver according to the lens compensation parameters generated by the reinforcement learning model.
2. The optical image stabilization compensation calculation method according to claim 1, characterized in that: The S2 includes: S2-1: Build a reinforcement learning model and perform quantification based on the definition of the current environment state, the collection of real-time data from the gyroscope and the range of compensation actions that the lens can perform; S2-2: Extract the angular velocity and angular acceleration from the real-time gyroscope data, extract the lens position and historical SR value, and determine the state vector of the reinforcement learning model. The expression is: s t =[ω t ,α t ,p t ,SR t-1 ] Among them, s t represents the state vector, ω t represents the angular velocity of the gyroscope, α t is the angular acceleration of the gyroscope, p t Indicates the lens position, SR t-1 Indicates the historical SR value; S2-3: Extract the lens displacement adjustment amount and determine the execution action of the reinforcement learning model. The expression is: a t =p t -p t-1 S2-4: Extract the quantitative results of the range of compensation actions that can be performed by the lens and determine the reward function, which is expressed as: r t =SR t -SR t-1 Among them, r t represents the reward function, SR t represents the SR value at time t, SR t-1 represents the SR value at time t-1, W off Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned off; W on Indicates the line width calculated by capturing an image under vibration conditions when the optical image stabilization algorithm of the lens is turned on; W static Indicates the line width calculated from a static image captured when the optical image stabilization algorithm is turned off.
3. The optical image stabilization compensation calculation method according to claim 2, characterized in that: The S3 includes: S3-1: Based on the optical image stabilization algorithm scenario, the proximal strategy optimization algorithm is used to design the reinforcement learning model processing logic; S3-2: Input the collected sample data into the reinforcement learning model for training, and continuously optimize the strategy until the reinforcement learning model with the optimal strategy is output.
4. The optical image stabilization compensation calculation method according to claim 3, characterized in that: The S3-1 includes: S3-1-1: Based on the policy function from the current state s t Predict action a t ; S3-1-2: Use the value function to evaluate the current state s t of income; S3-1-3: Calculate the total revenue by accumulating the reward discounts. The expression is: Where γ represents the discount factor; S3-1-4: The near-end strategy optimization algorithm is used to optimize the optical image stabilization compensation strategy. The target loss function of the near-end strategy optimization algorithm is: L PPO (θ)=E t [min(r t (θ)·A t ,clip(r t (θ),1-∈,1+ε)·A t )] A t =R t -V(s t ) V(s t )=E πθ [R t |s t ] Among them, r t (θ) represents the strategy probability ratio, A t represents the advantage function, ∈ represents the cut threshold for limiting strategy changes, V(s t ) is the state value function.
5. The optical image stabilization compensation calculation method according to claim 4, characterized in that: The S3-2 includes: S3-2-1: Build and initialize the training environment, including setting the amplitude and frequency of vibration to simulate vibration scenarios, adjusting the lens using the OIS driver, real-time gyroscope data, and capturing lens images; S3-2-2: Obtain the state vector, execution action and reward data in the static environment and vibration environment of the training environment, including the line width W off , line width W static , line width W on , angular velocity ω t , angular acceleration α t , lens position p t ; S3-2-3: Input the state vector into the reinforcement learning model; S3-2-4: According to the current strategy π θ (a|s) select action a t , and update the lens position p t =p t-1 +a t To generate lens compensation parameters; S3-2-5: Capture an image according to the compensated lens and calculate a new SR value of the image; S3-2-6: According to the reward function r t =SR t -SR t-1 Update reinforcement learning model parameters; S3-2-7: Repeat the above steps until the reinforcement learning model converges, the optimal strategy is obtained, and the trained reinforcement learning model is output.
6. The optical image stabilization compensation calculation method according to claim 5, characterized in that: The S4 includes: S4-1: Deploy the trained reinforcement learning model to the MCU for real-time reasoning; S4-2: Receive gyroscope data through MCU, and calculate the current state of the lens according to the current state of the lens. t Call the reinforcement learning model to infer the target lens position compensation action a t ; S4-3: Compensate motion a according to target lens position through OIS driver t Adjust the lens position.
7. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the processor stores programs or instructions, and the processor executes an optical image stabilization compensation calculation method as described in any one of claims 1 to 6 by calling the programs and instructions stored in the memory.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute an optical image stabilization compensation calculation method as described in any one of claims 1 to 6.