Method, device, equipment and medium for adjusting focal length of electro-hydraulic adjustable focus lens
By constructing a focal length adjustment model and target tracking model based on reinforcement learning, and combining it with the radar information of the unmanned boat, the focal length of the electro-hydraulic adjustable focus lens is automatically adjusted, which solves the problems of slow focusing speed and low accuracy of traditional electro-hydraulic adjustable focus lenses in dynamic environments, and achieves a fast and accurate optimal focusing effect.
Patent Information
- Application Number
- CN202411139584.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-08-20
AI Technical Summary
Traditional electro-hydraulic focus-adjustable lenses have slow focusing speed and low accuracy in dynamic environments, making it difficult to achieve optimal focusing quickly and accurately.
A reinforcement learning-based focus adjustment method is adopted. The scene image of the dynamic target is obtained through the camera. A reinforcement learning focus adjustment model and target tracking model are constructed. The radar on the unmanned boat is used to obtain distance information. Combined with the electro-hydraulic adjustable focus lens control network model, the focus is automatically adjusted to achieve optimized image clarity.
It realizes fast and accurate focusing of the electro-hydraulic focusable lens in a dynamic environment, automatically adjusts the focal length to meet the image clarity requirements, and improves focusing efficiency and accuracy.
Smart Images

Figure CN119136052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of lens focusing technology, and in particular to a method, device, equipment and medium for adjusting the focal length of an electro-hydraulic focusable lens. Background Art
[0002] In unmanned underwater vehicle applications, clear image acquisition is crucial. Electro-hydraulic focusable lenses, as advanced imaging devices, can achieve clear images by adjusting the lens' focal length. However, traditional focusing methods often rely on preset parameters or manual adjustments, making it difficult to quickly and accurately achieve optimal focus for dynamic targets in dynamic environments.
[0003] Therefore, the traditional method of focusing the electro-hydraulic adjustable focus lens has the problems of slow focusing speed and low accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, equipment and medium for adjusting the focal length of an electro-hydraulic focus-adjustable lens, which can quickly and accurately achieve optimal focusing of the electro-hydraulic focus-adjustable lens.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a method for adjusting the focal length of an electro-hydraulic adjustable focus lens, the method comprising:
[0007] Acquire a scene image of a dynamic target through a camera, and select an area containing the dynamic target as a target area to be focused; wherein the dynamic target is a movable target on the sea surface;
[0008] Calculating the image clarity of the focus target area;
[0009] When the image clarity is less than an image clarity threshold, calculating the focal length of the electro-hydraulic focus-adjustable lens;
[0010] Constructing a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network, and a target value network, and determining the trained policy network as a control network model of an electro-hydraulic focus-adjustable lens;
[0011] The image clarity and the focal length are input into the electro-hydraulic adjustable focus lens control network model to obtain a focusing current value, and the focal length of the electro-hydraulic adjustable focus lens is adjusted according to the focusing current value, and the adjusted image clarity is determined so that the adjusted image clarity is greater than or equal to an image clarity threshold.
[0012] An optimal focus current value is determined based on the adjusted image clarity, and the focal length of the electro-hydraulic focus-adjustable lens is adjusted according to the optimal focus current value.
[0013] Furthermore, before obtaining the scene image of the dynamic target through the camera, the method further includes:
[0014] Constructing a target tracking model based on reinforcement learning, wherein the target tracking model based on reinforcement learning includes a tracking strategy network, a tracking target strategy network, a tracking value network, and a tracking target value network, and determining the trained tracking strategy network as the unmanned boat target tracking network model;
[0015] Obtaining the distance between the dynamic target and the unmanned boat by using a radar on the unmanned boat;
[0016] When the distance is greater than or equal to the control distance of the electro-hydraulic adjustable focus lens, the distance is adjusted using the unmanned boat target tracking network model;
[0017] When the adjusted distance is less than the control distance of the electro-hydraulic adjustable focus lens, the unmanned boat is controlled to maintain the same speed and acceleration as the dynamic target, so that the unmanned boat and the dynamic target remain relatively stationary.
[0018] Furthermore, the distance is adjusted using the unmanned boat target tracking network model, including:
[0019] Obtaining the current state quantity; the state quantity includes the speed of the dynamic target, the acceleration of the dynamic target, the speed of the unmanned boat, the acceleration of the unmanned boat, and the angle between the unmanned boat bow direction and the line connecting the unmanned boat and the dynamic target;
[0020] The state quantity at the current moment is input into the unmanned boat target tracking network model to obtain the propeller control quantity of the unmanned boat, and the movement of the unmanned boat is controlled according to the propeller control quantity, thereby adjusting the distance so that the distance is less than the control distance of the electro-hydraulic adjustable focus lens.
[0021] Furthermore, the training process of the target tracking model based on reinforcement learning includes:
[0022] Construct an experience replay pool, wherein the experience replay pool stores multiple experience tuples, each of which includes a state quantity at time k, an action quantity at time k, a state quantity at time k+1, and a reward, wherein the action quantity is a propeller control quantity;
[0023] Sampling the experience tuples in the experience replay pool to obtain sampled experience tuples;
[0024] Inputting the state quantity at time k+1 in the sampled experience tuple into the tracking target strategy network to obtain the action quantity at time k+1;
[0025] Based on the target strategy smoothing regularization, the action amount at the k+1 moment is added with noise; the state amount at the k+1 moment and the action amount at the k+1 moment after adding noise are input into the tracking target value network to obtain the state value target value;
[0026] Input the state quantity at time k and the action quantity at time k in the sampled experience tuple into the tracking value network to output the state value evaluation value;
[0027] Minimize the error between the state value evaluation value and the state value target value using a gradient descent algorithm to update the parameters of the tracking value network;
[0028] Inputting the state quantity at time k in the sampled experience tuple into the tracking strategy network to obtain the predicted action quantity; inputting the state quantity at time k and the predicted action quantity into the tracking value network to obtain the predicted state value evaluation value;
[0029] The predicted state value evaluation value is maximized using a gradient ascent algorithm, the parameters in the tracking strategy network are updated, the parameters of the tracking target strategy network are updated according to the parameters of the updated tracking strategy network, and the parameters of the tracking target value network are updated according to the parameters of the updated tracking value network.
[0030] Furthermore, the training process of the focus adjustment model based on reinforcement learning includes:
[0031] Constructing an electro-hydraulic experience replay pool, the electro-hydraulic experience replay pool containing multiple electro-hydraulic experience tuples, each of which includes an electro-hydraulic state quantity at time G, an electro-hydraulic action quantity at time G, an electro-hydraulic state quantity at time G+1, and an electro-hydraulic reward; wherein the electro-hydraulic state quantity is image clarity and focal length, the image clarity includes overall image clarity, image edge clarity, and image contrast, and the electro-hydraulic action quantity is a focusing current value;
[0032] Sampling the electro-hydraulic experience tuple in the electro-hydraulic experience replay pool to obtain a sampled electro-hydraulic experience tuple;
[0033] Inputting the electro-hydraulic state quantity at time G+1 in the sampled electro-hydraulic experience tuple into the target strategy network to obtain the electro-hydraulic action quantity at time G+1;
[0034] Based on the target strategy smoothing regularization, the electro-hydraulic action quantity at the G+1 moment is added with noise; the electro-hydraulic state quantity at the G+1 moment and the electro-hydraulic action quantity at the G+1 moment after adding noise are input into the target value network to obtain the electro-hydraulic state value target value;
[0035] Input the electro-hydraulic state quantity and the electro-hydraulic action quantity at time G into the value network to output the electro-hydraulic state value evaluation value;
[0036] Minimizing the error between the electro-hydraulic state value evaluation value and the electro-hydraulic state value target value using a gradient descent algorithm to update the parameters of the value network;
[0037] Inputting the electro-hydraulic state quantity at time G in the sampled electro-hydraulic experience tuple into the strategy network to obtain the predicted electro-hydraulic action quantity;
[0038] Inputting the electro-hydraulic state quantity and the predicted electro-hydraulic action quantity at time G in the sampled experience tuple into the value network to obtain a predicted electro-hydraulic state value evaluation value;
[0039] The predicted electro-hydraulic state value evaluation value is maximized using a gradient ascent algorithm, the parameters of the strategy network are updated, the parameters of the target strategy network are updated according to the parameters of the updated strategy network, and the parameters of the target value network are updated according to the parameters of the updated value network.
[0040] Furthermore, the calculation formula of the electro-hydraulic reward is:
[0041]
[0042] Among them, r st For electro-hydraulic rewards, ω 11 ,ω 12 ,ω 13 ,ω 14 is a hyperparameter, f t-1 (a) are the image contrast, overall image clarity, image edge clarity and focal length at time G-1, a is the focus target area, f t (a) are the image contrast, overall image clarity, edge clarity and focal length at time G, respectively, and || || is the binary norm.
[0043] Furthermore, determining an optimal focus current value based on the adjusted image clarity, and adjusting the focal length of the electro-hydraulic focus-adjustable lens according to the optimal focus current value, includes:
[0044] Adjusting the focal length in the direction of increasing image clarity according to a preset focusing current step length to obtain a re-adjusted focal length, and determining the re-adjusted image clarity;
[0045] comparing the image clarity after adjustment with the image clarity after re-adjustment, and determining the optimal focus current value according to the comparison result;
[0046] The focal length of the electro-hydraulic focus-adjustable lens is adjusted according to the optimal focus current value.
[0047] In a second aspect, the present application provides a focus adjustment device for an electro-hydraulic focus-adjustable lens, wherein the focus adjustment system for the electro-hydraulic focus-adjustable lens comprises:
[0048] A target area determination module is used to obtain a scene image of a dynamic target through a camera, and select an area containing the dynamic target as a target area to be focused; wherein the dynamic target is a movable target on the sea surface;
[0049] A first calculation module, configured to calculate the image clarity of the focus target area;
[0050] a second calculation module, configured to calculate the focal length of the electro-hydraulic adjustable focus lens when the image clarity is less than an image clarity threshold;
[0051] A model construction module is used to construct a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network, and a target value network, and determine the trained policy network as the electro-hydraulic focus adjustable lens control network model;
[0052] a first focal length adjustment module, configured to input the image clarity and the focal length into a control network model of an electro-hydraulic adjustable focus lens, obtain a focusing current value, adjust the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, and determine an adjusted image clarity, so that the adjusted image clarity is greater than or equal to an image clarity threshold;
[0053] The second focus adjustment module is used to determine an optimal focus current value based on the adjusted image clarity, and adjust the focus of the electro-hydraulic focus-adjustable lens according to the optimal focus current value.
[0054] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the focal length adjustment method of the electro-hydraulic focus-adjustable lens described in any one of the above.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the above-described methods for adjusting the focal length of an electro-hydraulic focus-adjustable lens.
[0056] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0057] The present application provides a method, device, equipment and medium for adjusting the focal length of an electro-hydraulic adjustable focus lens. The present application determines a focus target area based on an image of a scene where a dynamic target is located, and when the image clarity of the focus target area is less than an image clarity threshold, uses an electro-hydraulic adjustable focus lens control network model to adjust the focal length so that the image clarity of the focus target area meets the requirements, thereby achieving automatic focus adjustment without manual intervention. Thereafter, the optimal focus current value is determined using the adjusted image clarity, and the focal length is adjusted again based on the optimal focus current value to further optimize the image clarity and achieve optimal focus of the electro-hydraulic adjustable focus lens. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0059] Figure 1 This is a diagram illustrating an application environment of a method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to an embodiment of the present application;
[0060] Figure 2 This is a flow chart of a method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to an embodiment of the present application;
[0061] Figure 3 A flowchart of a method for determining an optimal focusing current value provided in one embodiment of the present application;
[0062] Figure 4 A schematic diagram of the functional modules of a focus adjustment device for an electro-hydraulic focus-adjustable lens provided in one embodiment of the present application;
[0063] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0065] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0066] The focal length adjustment method of the electro-hydraulic focus-adjustable lens provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the scene image where the dynamic target is located to the server 104. After the server 104 receives the scene image where the dynamic target is located, for the scene image where the dynamic target is located, the server 104 selects the area containing the dynamic target as the target area that needs to be focused; calculates the image clarity of the focus target area; when the image clarity is less than the image clarity threshold, calculates the focal length of the electro-hydraulic adjustable focus lens; inputs the image clarity and the focal length into the electro-hydraulic adjustable focus lens control network model to obtain the focusing current value, and adjusts the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, determines the adjusted image clarity, and makes the adjusted image clarity greater than or equal to the image clarity threshold; determines the optimal focusing current value based on the adjusted image clarity, and adjusts the focal length of the electro-hydraulic adjustable focus lens according to the optimal focusing current value. The server 104 may feed back the adjusted focal length to the terminal 102. Furthermore, in some embodiments, the focal length adjustment method of the electro-hydraulic tunable focus lens may also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 may directly process the scene image where the dynamic target is located, or the server 104 may obtain the scene image where the dynamic target is located from a data storage system and adjust the focal length of the electro-hydraulic tunable focus lens based on the scene image where the dynamic target is located.
[0067] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, and tablet computers. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0068] In an exemplary embodiment, Figure 2 As shown, a method for adjusting the focal length of an electro-hydraulic adjustable focus lens is provided. The method is executed by a computer device. Specifically, the method can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In an embodiment of the present application, the method for adjusting the focal length of an electro-hydraulic adjustable focus lens includes the following steps 101 to 106, wherein:
[0069] Step 101, obtain a scene image of a dynamic target through a camera, and select an area containing the dynamic target as a target area that needs to be focused; wherein the dynamic target includes various ships, cruise ships, unmanned boats and other movable targets on the sea surface.
[0070] Step 102: Calculate the image clarity of the focus target area.
[0071] Step 103: When the image clarity is less than the image clarity threshold, calculate the focal length of the electro-hydraulic adjustable focus lens.
[0072] Step 104: construct a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network, and a target value network, and the trained policy network is determined as the electro-hydraulic focus adjustable lens control network model.
[0073] Step 105: Input the image clarity and the focal length into an electro-hydraulic adjustable focus lens control network model to obtain a focusing current value, adjust the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, and determine the adjusted image clarity, so that the adjusted image clarity is greater than or equal to an image clarity threshold.
[0074] Step 106 : determining an optimal focus current value based on the adjusted image clarity, and adjusting the focal length of the electro-hydraulic focus-adjustable lens according to the optimal focus current value.
[0075] By implementing the above steps 101 to 106 , the present application can improve the efficiency of focal length adjustment and achieve optimal focusing of the electro-hydraulic adjustable focus lens.
[0076] In another exemplary embodiment of the present application, before step 101, the following further comprises: step 100, controlling the unmanned boat so that the unmanned boat and the dynamic target remain relatively stationary. Step 100 specifically comprises:
[0077] Step 1001: Construct a target tracking model based on reinforcement learning, wherein the target tracking model based on reinforcement learning includes a tracking strategy network, a tracking target strategy network, a tracking value network, and a tracking target value network, and determine the trained tracking strategy network as the unmanned boat target tracking network model.
[0078] Step 1002: Obtain the distance between the dynamic target and the unmanned boat through the radar on the unmanned boat.
[0079] Step 1003: When the distance is greater than or equal to the control distance of the electro-hydraulic adjustable focus lens, the distance is adjusted using the unmanned boat target tracking network model.
[0080] Step 1004: When the adjusted distance is less than the control distance of the electro-hydraulic adjustable focus lens, the unmanned boat and the dynamic target are controlled to maintain the same speed and acceleration, so that the unmanned boat and the dynamic target remain relatively stationary.
[0081] Reinforcement learning (RL) is a machine learning method that learns optimal strategies based on reward signals through interaction with the environment. Compared to traditional control methods, RL does not require a clear model and has greater robustness and adaptability. Therefore, applying RL to the control of unmanned aerial vehicles (UAVs) can enable them to autonomously learn optimal motion control strategies in complex marine environments, ensuring that the UAV remains stationary relative to a dynamic target.
[0082] In another exemplary embodiment of the present application, the target tracking model based on reinforcement learning constructed in step 1001 includes a tracking strategy network, a tracking target strategy network, a tracking value network and a tracking target value network, specifically including the tracking strategy network being an Actor network, the tracking target strategy network being a Target Actor network, the tracking value network being a Critic1 network and a Critic2 network, and the tracking target value network being a Target Critic1 network and a Target Critic2 network.
[0083] The network architecture of the Actor network and the Target Actor network is the same, but their input and output are different. Specifically, the input of the Actor network is the state quantity at the current moment, and the output of the Actor network is the action quantity at the current moment. The input of the TargetActor network is the state quantity at the next moment, and the output of the TargetActor network is the action quantity at the next moment.
[0084] The four critic networks can be divided into two parts: Critic1 and Critic2, which are identical; and Target Critic1 and Target Critic2, which are also identical. The main differences between the critic and target critic networks are their input and output. The input of the critic network is the current state-action pair, and the output is the state value assessment of the current state-action pair. A state-action pair is a state-action pair consisting of a state quantity and an action quantity at the same moment. The input of the target critic network is the state-action pair at the next moment, and the output is the state value target value of the state-action pair at the next moment.
[0085] The Actor Network consists of one input layer, four hidden layers, and one output layer. The input layer contains six nodes, corresponding to the six parameters of the state quantity. The hidden layers are constructed using linear layers, with the four hidden layers containing 1024, 1024, 512, and 256 nodes, respectively. The output layer contains two nodes, corresponding to the two parameters of the action quantity.
[0086] The critic network architecture consists of two branches. The first branch, similar to the actor network, consists of one input layer and four hidden layers. The input layer contains six nodes, corresponding to the six parameters of the state. The hidden layers are constructed using linear layers, with the four hidden layers containing 1024, 1024, 512, and 256 nodes, respectively. The second branch consists of one input layer and two hidden layers. The input layer is the actor network output layer and contains two nodes. The two hidden layers contain 512 and 256 nodes, respectively. Both branches have the same size of 256-node hidden layer. The outputs of the two 256-node hidden layers are summed and normalized before being output to the output layer, which contains one node. Therefore, the output layer outputs the evaluation value or target value for the state-action pair.
[0087] In another exemplary embodiment of the present application, the training process of the target tracking model based on reinforcement learning includes the following steps:
[0088] S1. Construct an experience replay pool, in which multiple experience tuples are stored. The experience tuples include the state quantity at time k, the action quantity at time k, the state quantity at time k+1, and the reward. The action quantity is the propeller control quantity.
[0089] Create a simulation environment for unmanned vehicle target tracking, primarily consisting of the unmanned vehicle, radar, dynamic targets, and ocean environment. Randomly define the starting positions of the unmanned vehicle and dynamic targets within the unmanned vehicle target tracking simulation environment. Use the unmanned vehicle's onboard radar to obtain the dynamic target's current position, calculate its current velocity and acceleration, and use the unmanned vehicle's inertial navigation system to confirm its current velocity and acceleration. Also, determine the distance between the unmanned vehicle and the dynamic target, as well as the angle between the unmanned vehicle's bow heading and the line connecting the two.
[0090] The obtained parameters are used as state quantities Among them, the speed of the unmanned boat v u , the acceleration of the unmanned boat a u , the speed v of the dynamic target d , the acceleration of the dynamic target a d , the distance Δl between the unmanned boat and the dynamic target, the angle between the current heading angle of the unmanned boat and the line connecting the unmanned boat and the dynamic target Input the target tracking model based on reinforcement learning to obtain the action amount a that controls the movement of the unmanned boat t , the amount of action performed in, The left and right propeller control quantities are represented respectively, and the state quantity at the next moment is obtained. The left and right propeller control quantities are sent to the unmanned boat to achieve tracking of dynamic targets; at the same time, the reward is calculated according to the obtained state quantity. The target tracking model based on reinforcement learning is an improved twin-delayed deep deterministic policy gradient (TD3) reinforcement learning algorithm neural network.
[0091] Loop through the steps and pack the state s at time k t , the action amount a at time k t , the reward r at time k t And the state quantity s at time k+1 t ', forming a set of experience tuples (s t , a t , r t , s t ') and store it in the experience replay pool, which stores multiple experience tuples. Multiple sets of experience tuples are obtained and stored in the experience replay pool. If the distance between the unmanned boat and the dynamic target is less than the specified distance (the specified distance is the control distance of the electro-hydraulic adjustable focus lens) for more than 30 seconds, or if the distance does not fall below the specified distance within 60 seconds, the current round of training ends. Repeat the steps. If K rounds have been repeated, execute step S2 and restart the count. If the number of repeated rounds is less than K, continue to repeat the steps.
[0092] S2. Sampling the experience tuples in the experience replay pool to obtain sampled experience tuples.
[0093] The priority experience replay algorithm is used to sample the experience tuples in the experience replay pool, and n groups of experience tuples are taken from the experience replay pool.
[0094] Specifically, formula (1) is used to implement sampling of the experience replay pool:
[0095]
[0096] Wherein, P(c) represents the priority of a group of experience tuples screened from the experience replay pool, c represents the sequence number of the currently extracted experience tuple, Indicates the priority of the currently extracted experience tuple, represents the priority of the extracted experience tuple; α represents the preset parameter used to adjust the priority sampling degree of data samples.
[0097] S3. Input the state quantity at time k+1 in the sampled experience tuple into the target tracking strategy network to obtain the action quantity at time k+1.
[0098] S4. Add noise to the action amount at time k+1 based on target strategy smoothing regularization; input the state amount at time k+1 and the action amount at time k+1 after adding noise into the tracking target value network to obtain the state value target value.
[0099] Specifically, the action amount at time k+1 obtained by the tracking target strategy network is added with noise. The calculation formula for adding noise is:
[0100] a t '=μ'(s t '|θ μ' )+∈ (2)
[0101] ∈=clip(N(0,σ),-b,b) (3)
[0102] Among them, θ μ' are the network parameters of the Target Actor network, μ'() represents the Target Actor network, ∈ represents the random noise parameter; the clip function indicates that when N(0, σ) < -b, ∈ = -b, when N(0, σ) > b, ∈ = b, otherwise, ∈ = N(0, σ); N(0, σ) indicates that it satisfies the normal distribution, -b and b represent fixed parameters, and b>0.
[0103] Based on the idea of dual network, the Target Critic network is used to calculate the state action pair (s t ', a t ') state value target value y t , state value target value y t The calculation formula is as follows:
[0104]
[0105] Among them, r t is the reward function, γ is the discounted return rate, is the network parameter of Target Critic1 or Target Critic2 network, Q i '() represents the Target Critic1 or Target Critic2 network.
[0106] Reward t The calculation formula is:
[0107]
[0108] Among them, ω1, ω2, ω3, ω4, ω5 are hyperparameters, which are adjusted according to actual conditions, a is the amplitude height, σ is the standard deviation, b is the position parameter, and r 2tis the reward function within the control distance of the electro-hydraulic adjustable focus lens, and the speed v of the unmanned boat is u , the acceleration of the unmanned boat a u , the speed v of the dynamic target d , the acceleration of the dynamic target a d , the distance Δl between the unmanned boat and the dynamic target, the angle between the current heading angle of the unmanned boat and the line connecting the unmanned boat and the dynamic target
[0109]
[0110] Wherein, R is the control distance of the electro-hydraulic focusable lens.
[0111] S5. Input the state quantity at time k and the action quantity at time k in the sampled experience tuple into the tracking value network to output the state value evaluation value.
[0112] S6. Use the gradient descent algorithm to minimize the error between the state value evaluation value and the state value target value, and update the parameters of the tracking value network.
[0113] Specifically, the gradient descent algorithm is used to minimize the error between the state value evaluation value and the state value target value, and the parameters of the tracking value network are updated. The calculation formula is:
[0114]
[0115] Among them, Q i () represents the Critic1 or Critic2 network, s t and a t Respectively represent the state quantity and action quantity at the current moment, Get the network parameters of the Critic2 network for Critic1.
[0116] S7. Input the state quantity at time k in the sampled experience tuple into the tracking strategy network to obtain the predicted action quantity; input the state quantity at time k and the predicted action quantity into the tracking value network to obtain the predicted state value evaluation value.
[0117] S8. Use the gradient ascent algorithm to maximize the predicted state value evaluation value, update the parameters in the tracking strategy network, and update the parameters of the tracking target strategy network according to the parameters of the updated tracking strategy network, and update the parameters of the tracking target value network according to the parameters of the updated tracking value network.
[0118] Specifically, after the Ctitic1 and Critic2 networks are updated for d steps, the Actor network update is started, and the state quantity s at time k is calculated using the Actor network. t The predicted action amount a tnew , predict the action amount a tnew The calculation formula is:
[0119] a tnew =μ(s t |θ μ ) (8)
[0120] Among them, s t is the state quantity at time k, θ μ is the network parameter of the Actor network, and μ() represents the Actor network.
[0121] After calculating the predicted action amount, there is no need to add noise, because here we hope that the Actor network can be updated towards the maximum value, and adding noise does not make any sense.
[0122] Use Critic1 or Critic2 network to calculate the state action pair (s t , a tnew )’s state value evaluation value q t , state value evaluation value q t The calculation formula is as follows, assuming that the Critic1 network is used:
[0123] q t =Q1(s t , a tnew |θ Q1 ) (9)
[0124] Among them, θ Q1 are the network parameters of the Critic1 network, and Q1() represents the Critic1 network.
[0125] The gradient ascent algorithm is used to maximize the state value evaluation value and update the parameters in the Actor network. The reason why either Critic1 or Critic2 can be used to calculate the Q value here is mainly because the purpose of the Actor network is to maximize the cumulative expected return, and there is no need to use the minimum value.
[0126] The TD3 algorithm uses a soft update method to update the parameters of the tracking target strategy network and the tracking target value network. The parameter update process of the Target Actor network is as follows:
[0127] θ μ' =τθ μ +(1-τ)θ μ' (10)
[0128] Among them, τ is the learning rate (momentum), θ μ is the network parameter of the Actor network, θ μ' The network parameters for the Target Actor network.
[0129] The Target Critic1 and Target Critic2 network update process is as follows:
[0130]
[0131] in, are the network parameters of the Target Critic1 or Target Critic2 network, is the network parameter of Critic1 or Critic2 network, τ is the learning rate (momentum), τ∈(0,1), and is usually set to 0.005.
[0132] During training, when the reinforcement learning-based target tracking model reaches a predetermined number of stable cycles, M, the reward function graph is checked. If the reward has stabilized within cycle M, the loop ends, completing the reinforcement learning-based target tracking model training and obtaining the optimal network model. If the reward has not stabilized, the training process is repeated, and after each predetermined number of repetitions, the reward function in the reward function graph is checked to see if it has stabilized until the reward stabilizes.
[0133] In another exemplary embodiment of the present application, step 1001 is used to obtain a target tracking model based on reinforcement learning, and the trained tracking strategy network is determined as the unmanned boat target tracking network model, and then step 1002 is executed. Step 1002 specifically includes using the radar carried by the unmanned boat to obtain the direction of the dynamic target, calculating the speed and acceleration of the dynamic target, using the unmanned boat's own inertial navigation to determine the speed and acceleration of the unmanned boat, and obtaining the distance between the unmanned boat and the dynamic target, as well as the angle between the direction of the unmanned boat's bow and the line connecting the unmanned boat and the dynamic target.
[0134] In another exemplary embodiment of the present application, step 1003 uses the unmanned boat target tracking network model to adjust the distance, specifically including steps 10031 to 10032:
[0135] Step 10031. Obtain the state quantity at the current moment; the state quantity includes the speed of the dynamic target, the acceleration of the dynamic target, the speed of the unmanned boat, the acceleration of the unmanned boat, and the angle between the unmanned boat bow direction and the line connecting the unmanned boat and the dynamic target.
[0136] Specifically, in the above step 1002, while obtaining the distance between the radar dynamic target on the unmanned boat and the unmanned boat, the speed of the dynamic target, the acceleration of the dynamic target, the speed of the unmanned boat, the acceleration of the unmanned boat, and the angle between the direction of the unmanned boat's bow and the line connecting the unmanned boat and the dynamic target are obtained.
[0137] Step 10032: Input the current state quantity into the unmanned boat target tracking network model to obtain the propeller control quantity of the unmanned boat, control the movement of the unmanned boat according to the propeller control quantity, and then adjust the distance so that the distance is less than the control distance of the electro-hydraulic adjustable focus lens.
[0138] This application controls the distance between the unmanned boat and the dynamic target through an unmanned boat target tracking network model, thereby realizing automatic control and quickly adjusting the distance between the unmanned boat and the dynamic target. At the same time, by controlling the speed and acceleration of the unmanned boat and the dynamic target, the unmanned boat and the dynamic target remain relatively stationary. When the unmanned boat and the dynamic target remain relatively stationary, the scene image of the dynamic target is obtained through the camera, and then the focal length of the electro-hydraulic adjustable focus lens is adjusted.
[0139] In another exemplary embodiment of the present application, step 101 specifically includes:
[0140] An electro-hydraulic focusable lens is placed on a designated camera and mounted on a specific controllable three-degree-of-freedom gimbal to ensure that the dynamic target is always centered in the image. The camera captures an image of the scene containing the dynamic target and uses computer vision technology to automatically select the area containing the dynamic target as the target area to be focused.
[0141] In another exemplary embodiment of the present application, step 102 uses the image clarity formula (12) to calculate the image clarity of the focus target area. In this embodiment, the EOG function is used as the evaluation index of image clarity. The image clarity formula is as follows:
[0142] E0=E EOG =∑ Hight ∑ Widh [I(x+1,y)-I(x,y)] 2 +[I(x,y+1)-I(x,y)] 2 (12)
[0143] Wherein, E0 is the image clarity of the focus target area, and I(x, y) represents the pixel value of the image at the coordinate (x, y).
[0144] In another exemplary embodiment of the present application, in step 103, when calculating the focal length, formula (13) may be used for calculation:
[0145] The focal length is calculated using the lens imaging formula, which is:
[0146]
[0147] Where f is the focal length, u is the object distance, and v is the image distance.
[0148] In another exemplary embodiment of the present application, the reinforcement learning-based focus adjustment model constructed in step 104 specifically includes a policy network, a target policy network, a value network and a target value network, specifically including an Actor network, a Target Actor network, a Critic1 network, a Target Critic1 network, a Critic2 network and a Target Critic2 network.
[0149] The Actor network, Target Actor network, Critic1 network, Target Critic1 network, Critic2 network, and Target Critic2 network in the focal length adjustment model based on reinforcement learning have the same structure as the target tracking model based on reinforcement learning mentioned above, except that the input and output quantities of the Actor network, Target Actor network, Critic1 network, TargetCritic1 network, Critic2 network, and Target Critic2 network are changed respectively. Among them, the nodes of the input layer of the Actor network become 4 nodes, corresponding to the 4 parameters of the electro-hydraulic state quantity; the node of the output layer becomes 1 node, corresponding to the 1 parameter of the electro-hydraulic action quantity.
[0150] In the Critic network architecture, the first branch input layer has four nodes, corresponding to the four electro-hydraulic state variables. The second branch input layer has one node, corresponding to the output of the Actor network.
[0151] In another exemplary embodiment of the present application, the training process of the focus adjustment model based on reinforcement learning includes the following steps:
[0152] S11: Construct an electro-hydraulic experience replay pool, wherein the electro-hydraulic experience replay pool contains multiple electro-hydraulic experience tuples, and the electro-hydraulic experience tuples include the electro-hydraulic state quantity at time G, the electro-hydraulic action quantity at time G, the electro-hydraulic state quantity at time G+1, and the electro-hydraulic reward; wherein, the electro-hydraulic state quantity is image clarity and focal length, the image clarity includes overall image clarity, image edge clarity, and image contrast, and the electro-hydraulic action quantity is the focusing current value.
[0153] A simulation environment for training a focus adjustment model based on reinforcement learning is established, including dynamic targets, electro-hydraulic adjustable focus lenses, corresponding cameras and other equipment. The focus target area obtained in the above step 101 is obtained, and then the image clarity and focal length of the focus target area are calculated. The image clarity includes the overall image clarity, edge clarity and image contrast. The overall image clarity, edge clarity, image contrast and focal length are input into the focus adjustment model based on reinforcement learning as electro-hydraulic state quantities to obtain the electro-hydraulic action quantity, execute the electro-hydraulic action quantity, and realize the focus adjustment of the electro-hydraulic adjustable focus lens. At the same time, the electro-hydraulic reward is calculated based on the obtained electro-hydraulic state quantity. The focus adjustment model based on reinforcement learning is an improved twin-delayed deep deterministic policy gradient (TD3) reinforcement learning algorithm neural network.
[0154] The steps are executed in a loop to package the electro-hydraulic state quantity at time G, the electro-hydraulic action quantity at time G, the electro-hydraulic reward function at time G, and the electro-hydraulic state quantity at time G+1 to form a set of electro-hydraulic experience tuples, which are stored in the electro-hydraulic experience replay pool. The electro-hydraulic state quantity is image clarity and focal length. The image clarity includes overall image clarity, image edge clarity, and image contrast. The electro-hydraulic action quantity is the focusing current value. Multiple sets of electro-hydraulic experience tuples are obtained and stored in the electro-hydraulic experience replay pool. If the image clarity exceeds the image clarity threshold, the training of this round is terminated. Repeat the steps. If K rounds have been repeated, execute step S22 and recount. If the number of repeated rounds is less than K, continue to repeat the steps.
[0155] S22. Sampling the electro-hydraulic experience tuples in the electro-hydraulic experience replay pool to obtain sampled electro-hydraulic experience tuples.
[0156] The specific implementation of this process is the same as that of the above step S2 and will not be repeated here.
[0157] S33: Input the electro-hydraulic state quantity at time G+1 in the sampled electro-hydraulic experience tuple into the target strategy network to obtain the electro-hydraulic action quantity at time G+1.
[0158] S44. Add noise to the electro-hydraulic action quantity at time G+1 based on target strategy smoothing regularization; input the electro-hydraulic state quantity at time G+1 and the electro-hydraulic action quantity at time G+1 after adding noise into the target value network to obtain the electro-hydraulic state value target value.
[0159] S55. Input the electro-hydraulic state quantity at time G and the electro-hydraulic action quantity at time G into the value network to output the electro-hydraulic state value evaluation value.
[0160] S66. Minimize the error between the electro-hydraulic state value evaluation value and the electro-hydraulic state value target value using a gradient descent algorithm, and update the parameters of the value network.
[0161] S77, inputting the electro-hydraulic state quantity at time G in the sampled electro-hydraulic experience tuple into the strategy network to obtain a predicted electro-hydraulic action quantity; inputting the electro-hydraulic state quantity at time G in the sampled electro-hydraulic experience tuple and the predicted electro-hydraulic action quantity into the value network to obtain a predicted electro-hydraulic state value assessment value;
[0162] S88. Use the gradient ascent algorithm to maximize the predicted electro-hydraulic state value assessment value, update the parameters of the strategy network, and update the parameters of the target strategy network according to the parameters of the updated strategy network, and update the parameters of the target value network according to the parameters of the updated value network.
[0163] In another exemplary embodiment of the present application, the process of updating the parameters of the value network through the above steps S33 to S66 is the same as the process of updating the parameters of the tracking value network, except that the input and output of the tracking target strategy network are changed to the input and output of the corresponding target strategy network; the input and output of the tracking target value network are changed to the input and output of the corresponding target value network; the input of the tracking value network is changed to the input of the value network, and the output of the corresponding value network, i.e., the electro-hydraulic state value target value, is obtained. The reward in the calculation formula (4) of the state value target value is changed to the electro-hydraulic reward to obtain the calculation formula of the electro-hydraulic state value target value. The calculation formula (14) of the electro-hydraulic reward is:
[0164]
[0165] Among them, r st For electro-hydraulic rewards, ω 11 ,ω 12 ,ω 13 ,ω 14 is a hyperparameter, f t-1 (a) are the image contrast, overall image clarity, image edge clarity and focal length at time G-1, a is the focus target area, f t (a) are the image contrast, overall image clarity, edge clarity and focal length at time G, respectively, and ‖‖ is the binary norm.
[0166] The parameter update process of the policy network, target policy network and target value network through the above steps S77 to S88 is the same as the update process of the above steps S7 to S8, except that the input and output of the tracking policy network are changed to the input and output of the policy network; the parameter update of the target policy network and the target value network is based on the parameters of the corresponding updated policy network and the parameters of the updated value network. The specific process is not repeated here.
[0167] During the training process, when the number of training rounds of the reinforcement learning-based focus adjustment model reaches a predetermined number of stable cycles M, the number of rounds-electro-hydraulic reward curve is checked. If the electro-hydraulic reward has stabilized at cycle M, the reinforcement learning-based focus adjustment model training is complete, and the optimal network model is obtained. If the electro-hydraulic reward has not stabilized, the training process is repeated, and after each predetermined number of repetitions, the electro-hydraulic reward in the number of rounds-electro-hydraulic reward curve is checked to see if it has stabilized until the electro-hydraulic reward stabilizes.
[0168] In another exemplary embodiment of the present application, Figure 3 As shown, step 106 specifically includes:
[0169] Step 1061 : adjusting the focal length in the direction of increasing the image clarity according to a preset focus current step length, obtaining a re-adjusted focal length, and determining the re-adjusted image clarity.
[0170] In this embodiment, the preset focus current step size is selected as 2 mA, the adjusted image clarity is adjusted again in the direction of increasing image clarity according to the preset focus current step size, and the re-adjusted image clarity is determined.
[0171] Step 1062: Compare the adjusted image clarity with the re-adjusted image clarity, and determine the optimal focus current value according to the comparison result.
[0172] Specifically, the image clarity after adjustment is E1, and the image clarity after further adjustment is E2. Compare E1 and E1. If E1≤E2, return to step 1061. If E1>E2, the focusing current value corresponding to E1 is the optimization result of a single hill climbing, and the focusing current value corresponding to E1 is used as the optimal focusing current value.
[0173] Step 1063: Adjust the focal length of the electro-hydraulic adjustable focus lens according to the optimal focusing current value.
[0174] Specifically, the focal length of the electro-hydraulic adjustable focus lens is adjusted according to the optimal focusing current value I obtained in step 1062 .
[0175] This application uses a camera to obtain a scene image where a dynamic target is located, and based on this, obtains the image clarity of the focused target area. When the image clarity does not meet the clarity requirements, the focal length is adjusted using an electro-hydraulic adjustable focus lens control network model. The electro-hydraulic adjustable focus lens control network model is used to automatically adjust the focal length, quickly obtain an image with image clarity that meets the requirements, and achieve autonomous focusing without human intervention. Afterwards, the image clarity is adjusted again using a single hill climbing optimization algorithm to determine the optimal focusing current value, further optimize the image clarity, and achieve optimal focusing of the electro-hydraulic adjustable focus lens.
[0176] Based on the same inventive concept, embodiments of the present application also provide a device for implementing the aforementioned electro-hydraulic focus-adjustable lens focus adjustment method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the embodiments of one or more electro-hydraulic focus-adjustable lens focus adjustment devices provided below can be found in the aforementioned limitations of the electro-hydraulic focus-adjustable lens focus adjustment method, and will not be further elaborated here.
[0177] In an exemplary embodiment, Figure 4 As shown, a focus adjustment device for an electro-hydraulic focus-adjustable lens is provided, comprising:
[0178] The target area determination module 31 is used to obtain a scene image of a dynamic target through a camera and select an area containing the dynamic target as a target area to be focused; wherein the dynamic target is a movable target on the sea surface.
[0179] The first calculation module 32 is configured to calculate the image clarity of the focus target area.
[0180] The second calculation module 33 is configured to calculate the focal length of the electro-hydraulic adjustable focus lens when the image clarity is less than an image clarity threshold.
[0181] The model construction module 34 is used to construct a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network and a target value network, and determines the trained policy network as the electro-hydraulic adjustable focus lens control network model.
[0182] The first focal length adjustment module 35 is used to input the image clarity and the focal length into the electro-hydraulic adjustable focus lens control network model to obtain a focusing current value, and adjust the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, and determine the adjusted image clarity so that the adjusted image clarity is greater than or equal to an image clarity threshold; wherein the electro-hydraulic adjustable focus lens control network model includes a strategy network, a target strategy network, a value network, and a target value network.
[0183] The second focus adjustment module 36 is configured to determine an optimal focus current value based on the adjusted image clarity, and adjust the focus of the electro-hydraulic focus-adjustable lens according to the optimal focus current value.
[0184] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store scene images where dynamic targets are located. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for adjusting the focal length of an electro-hydraulic adjustable focus lens is implemented.
[0185] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0186] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0187] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0189] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0190] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0191] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0192] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for adjusting the focal length of an electro-hydraulic focus-adjustable lens, characterized in that: The focal length adjustment method of the electro-hydraulic focus-adjustable lens comprises: Acquire a scene image of a dynamic target through a camera, and select an area containing the dynamic target as a target area to be focused; wherein the dynamic target is a movable target on the sea surface; Calculating the image clarity of the focus target area; When the image clarity is less than an image clarity threshold, calculating the focal length of the electro-hydraulic focus-adjustable lens; Constructing a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network, and a target value network, and determining the trained policy network as a control network model of an electro-hydraulic focus-adjustable lens; Inputting the image clarity and the focal length into an electro-hydraulic adjustable focus lens control network model to obtain a focusing current value, adjusting the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, and determining an adjusted image clarity, such that the adjusted image clarity is greater than or equal to an image clarity threshold; Determining an optimal focus current value based on the adjusted image clarity, and adjusting the focal length of the electro-hydraulic focus-adjustable lens according to the optimal focus current value; Before obtaining the scene image of the dynamic target through the camera, it also includes: Constructing a target tracking model based on reinforcement learning, wherein the target tracking model based on reinforcement learning includes a tracking strategy network, a tracking target strategy network, a tracking value network, and a tracking target value network, and determining the trained tracking strategy network as the unmanned boat target tracking network model; Obtaining the distance between the dynamic target and the unmanned boat by using a radar on the unmanned boat; When the distance is greater than or equal to the control distance of the electro-hydraulic adjustable focus lens, the distance is adjusted using the unmanned boat target tracking network model; When the adjusted distance is less than the control distance of the electro-hydraulic adjustable focus lens, the unmanned boat is controlled to maintain the same speed and acceleration as the dynamic target, so that the unmanned boat and the dynamic target remain relatively stationary.
2. The method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to claim 1, wherein: The distance is adjusted using the unmanned boat target tracking network model, including: Obtaining the current state quantity; the state quantity includes the speed of the dynamic target, the acceleration of the dynamic target, the speed of the unmanned boat, the acceleration of the unmanned boat, and the angle between the unmanned boat bow direction and the line connecting the unmanned boat and the dynamic target; The state quantity at the current moment is input into the unmanned boat target tracking network model to obtain the propeller control quantity of the unmanned boat, and the movement of the unmanned boat is controlled according to the propeller control quantity, thereby adjusting the distance so that the distance is less than the control distance of the electro-hydraulic adjustable focus lens.
3. The method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to claim 2, wherein: The training process of the reinforcement learning-based target tracking model includes: Construct an experience replay pool, wherein the experience replay pool stores multiple experience tuples, each of which includes a state quantity at time k, an action quantity at time k, a state quantity at time k+1, and a reward, wherein the action quantity is a propeller control quantity; Sampling the experience tuples in the experience replay pool to obtain sampled experience tuples; Inputting the state quantity at time k+1 in the sampled experience tuple into the tracking target strategy network to obtain the action quantity at time k+1; Based on the target strategy smoothing regularization, the action amount at the k+1 moment is added with noise; the state amount at the k+1 moment and the action amount at the k+1 moment after adding noise are input into the tracking target value network to obtain the state value target value; Input the state quantity at time k and the action quantity at time k in the sampled experience tuple into the tracking value network to output the state value evaluation value; Minimize the error between the state value evaluation value and the state value target value using a gradient descent algorithm to update the parameters of the tracking value network; Inputting the state quantity at time k in the sampled experience tuple into the tracking strategy network to obtain the predicted action quantity; inputting the state quantity at time k and the predicted action quantity into the tracking value network to obtain the predicted state value evaluation value; The predicted state value evaluation value is maximized using a gradient ascent algorithm, the parameters in the tracking strategy network are updated, the parameters of the tracking target strategy network are updated according to the parameters of the updated tracking strategy network, and the parameters of the tracking target value network are updated according to the parameters of the updated tracking value network.
4. The method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to claim 1, wherein: The training process of the reinforcement learning-based focus adjustment model includes: Constructing an electro-hydraulic experience replay pool, the electro-hydraulic experience replay pool containing multiple electro-hydraulic experience tuples, each of which includes an electro-hydraulic state quantity at time G, an electro-hydraulic action quantity at time G, an electro-hydraulic state quantity at time G+1, and an electro-hydraulic reward; wherein the electro-hydraulic state quantity is image clarity and focal length, the image clarity includes overall image clarity, image edge clarity, and image contrast, and the electro-hydraulic action quantity is a focusing current value; Sampling the electro-hydraulic experience tuple in the electro-hydraulic experience replay pool to obtain a sampled electro-hydraulic experience tuple; Inputting the electro-hydraulic state quantity at time G+1 in the sampled electro-hydraulic experience tuple into the target strategy network to obtain the electro-hydraulic action quantity at time G+1; Based on the target strategy smoothing regularization, the electro-hydraulic action quantity at the G+1 moment is added with noise; the electro-hydraulic state quantity at the G+1 moment and the electro-hydraulic action quantity at the G+1 moment after adding noise are input into the target value network to obtain the electro-hydraulic state value target value; Input the electro-hydraulic state quantity and the electro-hydraulic action quantity at time G into the value network to output the electro-hydraulic state value evaluation value; Minimizing the error between the electro-hydraulic state value evaluation value and the electro-hydraulic state value target value using a gradient descent algorithm to update the parameters of the value network; Inputting the electro-hydraulic state quantity at time G in the sampled electro-hydraulic experience tuple into the strategy network to obtain the predicted electro-hydraulic action quantity; Input the electro-hydraulic state quantity and the predicted electro-hydraulic action quantity at time G in the sampled electro-hydraulic experience tuple into the value network to obtain the predicted electro-hydraulic state value evaluation value; The predicted electro-hydraulic state value evaluation value is maximized using a gradient ascent algorithm, the parameters of the strategy network are updated, the parameters of the target strategy network are updated according to the parameters of the updated strategy network, and the parameters of the target value network are updated according to the parameters of the updated value network.
5. The method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to claim 4, wherein: The calculation formula for the electro-hydraulic bonus is: Among them, r st For electro-hydraulic rewards, ω 11 ,ω 12 ,ω 13 ,ω 14 is a hyperparameter, f t-1 (a) are the image contrast, overall image clarity, image edge clarity and focal length at time G-1, a is the focus target area, f t (a) are the image contrast, overall image clarity, edge clarity and focal length at time G, respectively, and || || is the binary norm.
6. The method for adjusting the focal length of an electro-hydraulic focus-adjustable lens according to claim 1, wherein: Determining an optimal focus current value based on the adjusted image clarity, and adjusting the focal length of the electro-hydraulic focus-adjustable lens according to the optimal focus current value, including: Adjusting the focal length in the direction of increasing image clarity according to a preset focusing current step length to obtain a re-adjusted focal length, and determining the re-adjusted image clarity; comparing the image clarity after adjustment with the image clarity after re-adjustment, and determining the optimal focus current value according to the comparison result; The focal length of the electro-hydraulic focus-adjustable lens is adjusted according to the optimal focus current value.
7. A focal length adjustment device for an electro-hydraulic focus-adjustable lens, characterized in that: The focal length adjustment system of the electro-hydraulic focus-adjustable lens comprises: A target area determination module is used to obtain a scene image of a dynamic target through a camera, and select an area containing the dynamic target as a target area to be focused; wherein the dynamic target is a movable target on the sea surface; A first calculation module, configured to calculate the image clarity of the focus target area; a second calculation module, configured to calculate the focal length of the electro-hydraulic adjustable focus lens when the image clarity is less than an image clarity threshold; A model construction module is used to construct a focus adjustment model based on reinforcement learning, wherein the focus adjustment model based on reinforcement learning includes a policy network, a target policy network, a value network, and a target value network, and determine the trained policy network as the electro-hydraulic focus adjustable lens control network model; a first focal length adjustment module, configured to input the image clarity and the focal length into a control network model of an electro-hydraulic adjustable focus lens, obtain a focusing current value, adjust the focal length of the electro-hydraulic adjustable focus lens according to the focusing current value, and determine an adjusted image clarity, so that the adjusted image clarity is greater than or equal to an image clarity threshold; a second focus adjustment module, configured to determine an optimal focus current value based on the adjusted image clarity, and adjust the focus of the electro-hydraulic focus-adjustable lens according to the optimal focus current value; The focal length adjustment system of the electro-hydraulic focus-adjustable lens also includes: A target tracking model construction module is used to construct a target tracking model based on reinforcement learning, wherein the target tracking model based on reinforcement learning includes a tracking strategy network, a tracking target strategy network, a tracking value network, and a tracking target value network, and determines the trained tracking strategy network as the unmanned boat target tracking network model; An acquisition module is used to acquire the distance between the dynamic target and the unmanned boat through a radar on the unmanned boat; An adjustment module, configured to adjust the distance using an unmanned boat target tracking network model when the distance is greater than or equal to the control distance of the electro-hydraulic adjustable focus lens; The control module is used to control the unmanned boat and the dynamic target to maintain the same speed and acceleration when the adjusted distance is less than the control distance of the electro-hydraulic adjustable focus lens, so that the unmanned boat and the dynamic target remain relatively stationary.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for adjusting the focal length of the electro-hydraulic focus-adjustable lens according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for adjusting the focal length of the electro-hydraulic focus-adjustable lens according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Target tracking method of mobile robot based on electrical-hydraulic focusable lens
CN108600620A
Automatic focusing method and system for electro-hydraulic focusable lens, and electronic equipment
CN117156272A