Special-shaped windshield glass mounting method, device, equipment and medium

By initializing the robotic arm control strategy and deploying high-precision sensors, the precise installation of special-shaped windshield glass is solved, the problem of installation difficulty of special-shaped curved glass is improved, the installation accuracy and yield rate are reduced, and production costs and safety risks are reduced.

CN120229368APending Publication Date: 2025-07-01SHENZHEN POLYTECHNIC +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510546804.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The special curved structure of the front windshield of the helicopter increases the difficulty of installation, resulting in inconvenient manual handling and inaccurate installation, which can easily cause glass damage and safety accidents.

Method used

By initializing the control strategy of the robotic arm, deploying high-precision sensors such as lidar and industrial cameras, monitoring the environmental status in real time, analyzing each action command of the robotic arm based on real-time status information and control strategies, quantifying the reward value, updating the value function, optimizing the control strategy, and finally realizing the precise installation of the robotic arm.

Benefits of technology

It significantly improves the installation accuracy of special-shaped windshield glass, improves the product yield rate, shortens the production cycle, reduces production costs, and ensures flight safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120229368A_ABST
    Figure CN120229368A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a special-shaped windshield glass mounting method, device and equipment and a medium. The environment state of the mechanical arm at the current moment is obtained by initializing a control strategy of the mechanical arm provided with the special-shaped windshield glass; determining a reward item of the control strategy based on the environment state at the current moment and the control strategy; updating the value function according to the reward item of the control strategy to obtain an updated value function; optimizing the control strategy according to the updated value function to obtain an optimized control strategy; the step 102 and the step 106 are executed circularly until the updated value function is converged, the optimized control strategy at the moment serves as the final control strategy of the mechanical arm, the mechanical arm is controlled to install the special-shaped windshield according to the final control strategy, and through the configuration mode, the installation precision of the special-shaped windshield can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment and medium for installing a special-shaped windshield glass. Background Art

[0002] As a key component for ensuring flight safety and the pilot's vision, the installation of the front windshield glass of a helicopter cannot be underestimated. Different from ordinary glass, the front windshield glass of a helicopter uses a special-shaped transparent curved glass. This special design aims to optimize the aerodynamic performance, reduce wind resistance and noise during flight, and at the same time provide a wider and clearer vision for the pilot. However, this special-shaped curved structure greatly increases the installation difficulty.

[0003] Currently, the relatively conventional installation method is manual handling and docking installation. In actual operation, the staff needs to manually carry the heavy and irregularly shaped front windshield glass to the installation part of the helicopter, and then carry out detailed docking operations. This process faces many challenges: on the one hand, the shape of the special-shaped curved glass makes manual handling extremely inconvenient. The self-weight of the glass is large, and multiple workers are required to cooperate. Moreover, during the handling process, due to the smooth surface and irregular shape of the glass, it is difficult to achieve stable grasping. Not only does it consume a lot of physical strength, but it is also easy to cause the glass to be bumped or slipped due to improper operation, resulting in glass damage and even safety accidents. On the other hand, in the docking and installation link, due to the lack of accurate positioning and calibration means, workers mainly rely on experience to judge the position of the glass, and it is difficult to ensure the precise matching of the glass and the installation part of the helicopter fuselage. Once the installation position deviates, it will not only affect the sealing performance of the glass, resulting in problems such as air leakage and water leakage during flight, but also there is a probability of local stress concentration, which is affected by factors such as air flow impact and temperature change during flight, causing the glass to crack and seriously endangering flight safety.

[0004] Based on this, the present invention proposes a method and device for installing a special-shaped windshield glass to solve the above technical problems. Summary of the Invention

[0005] The present invention describes a method, device, equipment and medium for installing a special-shaped windshield glass, which can effectively improve the installation accuracy of the special-shaped windshield glass.

[0006] According to the first aspect, the present invention provides a method for installing a special-shaped windshield glass, including:

[0007] Step 100, initialize the control strategy of the robotic arm for installing the special-shaped windshield glass, and obtain the environmental state of the robotic arm at the current moment;

[0008] Step 102, determine the reward item of the control strategy based on the environmental state at the current moment and the control strategy;

[0009] Step 104: Update the value function according to the reward item of the control strategy to obtain the updated value function.

[0010] Step 106: Optimize the control strategy according to the updated value function to obtain the optimized control strategy.

[0011] Step 108: Loop and execute Steps 102 to 106 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robotic arm.

[0012] Step 110: Control the robotic arm to install the special-shaped windshield according to the final control strategy.

[0013] According to a second aspect, the present invention provides a special-shaped windshield installation device, including:

[0014] An acquisition unit configured to initialize the control strategy of a robotic arm for installing a special-shaped windshield and acquire the environmental state of the robotic arm at the current moment.

[0015] A first data processing unit configured to determine the reward item of the control strategy based on the environmental state at the current moment and the control strategy.

[0016] A second data processing unit configured to update the value function according to the reward item of the control strategy and determine the updated value function.

[0017] A third data processing unit configured to optimize the control strategy according to the updated value function to obtain the optimized control strategy.

[0018] A loop unit configured to loop and execute Steps 102 to 106 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robotic arm.

[0019] A control unit configured to control the robotic arm to install the special-shaped windshield according to the final control strategy.

[0020] According to a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method of the first aspect is implemented.

[0021] According to a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.

[0022] According to the installation method, device, equipment and medium of the special-shaped windshield glass provided by the present invention, the control strategy of the robotic arm undertaking the installation task of the special-shaped windshield glass is initialized. At the same time, high-precision sensors such as lidar and industrial cameras are deployed to realize multi-dimensional and real-time monitoring of the environment where the robotic arm is located, obtain the current environmental state data of the robotic arm, analyze each action instruction of the robotic arm based on the real-time state information of the robotic arm and the established control strategy, and quantitatively output the corresponding reward value. After obtaining the reward value, the value function is immediately updated. Subsequently, the current control strategy is deeply optimized to output a more scientific and reasonable control strategy. In order to ensure the optimization of the control strategy, operations such as value function update and control strategy optimization are continuously looped. Finally, the robotic arm carries out the installation work of the special-shaped windshield glass according to this final control strategy. Through the above configuration method, the present invention significantly improves the installation accuracy of the special-shaped windshield glass. After actual measurement, compared with the traditional installation method, the qualified rate of products is greatly improved, the production cycle is effectively shortened, and the production cost of the enterprise is also reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 The flowchart showing the installation method of the special-shaped windshield glass according to one embodiment is shown;

[0025] Figure 2 The schematic block diagram showing the installation device of the special-shaped windshield glass according to one embodiment is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following describes the solution provided by the present invention in conjunction with the drawings.

[0027] Figure 1 The flowchart showing the installation method of the special-shaped windshield glass according to one embodiment is shown. It can be understood that this method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. As Figure 1 shown, this method includes:

[0028] Step 100: Initialize the control strategy of the robotic arm for installing the special-shaped windshield glass, and obtain the environmental state of the robotic arm at the current moment;

[0029] Step 102: Determine the reward item of the control strategy based on the environmental state and the control strategy at the current moment;

[0030] Step 104: Update the value function according to the reward item of the control strategy to obtain the updated value function.

[0031] Step 106: Optimize the control strategy according to the updated value function to obtain the optimized control strategy.

[0032] Step 108: Loop through steps 102 - 106 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy for the robotic arm.

[0033] Step 110: Control the robotic arm to install the special-shaped windshield according to the final control strategy.

[0034] In this embodiment, the control strategy of the robotic arm undertaking the task of installing the special-shaped windshield is initialized. Meanwhile, high-precision sensors such as lidar and industrial cameras are deployed to achieve multi-dimensional and real-time monitoring of the environment where the robotic arm is located, obtain the current environmental state data of the robotic arm, analyze each action instruction of the robotic arm based on the real-time state information and the established control strategy of the robotic arm, and quantitatively output the corresponding reward value. After obtaining the reward value, the value function is immediately updated. Subsequently, the current control strategy is deeply optimized to output a more scientific and reasonable control strategy. To ensure the optimization of the control strategy, operations such as value function update and control strategy optimization will be continuously looped. Finally, the robotic arm installs the special-shaped windshield according to this final control strategy. Through the above configuration method, the present invention significantly improves the installation accuracy of the special-shaped windshield. After actual measurement, compared with the traditional installation method, the yield rate of products is greatly improved, the production cycle is effectively shortened, and the production cost of the enterprise is also reduced.

[0035] In an embodiment of the present invention, based on the environmental state and control strategy at the current moment, determine the reward item of the control strategy, including:

[0036] Based on the environmental state and control strategy at the current moment, determine the execution action of the robotic arm at the current moment.

[0037] Based on the execution action and transition probability at the current moment, determine the environmental state at the next moment.

[0038] Based on the execution action at the current moment and the environmental state at the next moment, determine the reward item of the control strategy.

[0039] In this embodiment, through devices such as lidar and vision sensors deployed in all directions, various parameters of the environment where the robotic arm is located are obtained in real time and accurately. Based on these data, the optimal execution action of the robotic arm at the current moment is quickly screened out. To predict the environmental changes caused by the actions of the robotic arm, a state transition model trained based on a large amount of historical data is used. Referring to the transition probabilities set in the model, the environmental state at the next moment after the execution of the action is effectively predicted, and a new state of the robotic arm, the special-shaped windshield, and the surrounding working environment is simulated. After completing the prediction of the environmental state at the next moment, relying on the pre-established reward evaluation model, the current execution action of the robotic arm and the predicted environmental state at the next moment are substituted into it. The reward term corresponding to the control strategy is quantitatively obtained, providing data support for the subsequent optimization of the control strategy.

[0040] In one embodiment of the present invention, the control strategy is optimized according to the updated value function to obtain the optimized control strategy, including:

[0041] For each actionable action in the environmental state at the current moment, with the help of the updated value function, the value numerical corresponding to each action is determined; wherein, the value numerical corresponding to each action is the expected cumulative reward that can be obtained after executing this action;

[0042] The value numerical corresponding to all actions are sorted from high to low to obtain a sorting table of the value numerical corresponding to the actions;

[0043] Based on the sorting table, the action with the largest value numerical is determined and labeled as the optimal action;

[0044] The action originally selected by the current strategy in this state is replaced with the determined optimal action to obtain the optimized control strategy.

[0045] In this embodiment, optimizing the control strategy to improve the operation efficiency of the robotic arm is the core task. After the value function is updated, these latest data are used to finely adjust the control strategy. The specific optimization steps are as follows: First, for each feasible operation action of the robotic arm in the current environmental state, the present invention deeply evaluates it by means of the updated value function. Through a complex and precise calculation logic, the value corresponding to each action is determined. Here, the value represents the expected cumulative reward that can be obtained in the future period after executing this action, which comprehensively considers the impacts of the action on task completion, efficiency, quality, etc. Then, the present invention collects the values corresponding to all actions and arranges them in descending order. This process generates a detailed action value ranking table, clearly showing the pros and cons order of each action in the current environmental state. Subsequently, based on this ranking table, the present invention can quickly identify the action with the largest value. This action is recognized as the optimal choice in the current environmental state because it indicates that it can bring the largest cumulative reward to the robotic arm and helps to complete the installation task of the special-shaped windshield glass more efficiently and accurately. Finally, the present invention replaces the action originally selected by the current control strategy in this state with the newly determined optimal action. Through this replacement operation, the control strategy is optimized and can better adapt to the environmental state where the robotic arm is currently located, thereby further improving the installation accuracy and efficiency of the special-shaped windshield glass. After such an optimization process, the finally obtained is the optimized control strategy.

[0046] In one embodiment of the present invention, the convergence of the value function is determined through the following steps:

[0047] After each iterative update of the value function, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration;

[0048] From all the calculated differences, determine the difference with the largest absolute value;

[0049] If the difference with the largest absolute value is less than or equal to the preset threshold, it is determined that the value function converges.

[0050] In this embodiment, the difference calculation is as follows: Each time the value function is iteratively updated, the difference calculation program is started. The value function obtained in this iteration is compared with the result of the previous iteration to calculate the difference between the two. Maximum difference screening: After all differences are calculated, the present invention comprehensively sorts out these difference data. With the help of a data screening algorithm, the absolute value of the largest difference is quickly locked among numerous differences. This largest difference can intuitively reflect the most significant part of the change in the value function between two iterations. Convergence determination: The absolute value of the largest difference selected is compared with a preset threshold. If the largest difference is less than or equal to the preset threshold, it means that the change range of the value function between two iterations is within an acceptable range and is in a relatively stable state. At this time, it can be determined that the value function has converged.

[0051] In one embodiment of the present invention, the value function includes a state value function and an action value function.

[0052] In this embodiment, the action value function: It describes the expectation of the long-term cumulative reward that can be obtained by taking a certain action in a specific state following the current policy. The action value function not only considers the state but also clarifies the value difference of taking different actions in that state. In the present invention, the action value function can be used to compare the value of performing different movement, grasping, etc. actions for finally completing the task in the current position and posture of the robotic arm. State value function: It represents the expectation of the long-term cumulative reward that can be obtained in a certain state following the current policy. By calculating the state value function, the quality of different states can be evaluated to help determine which states are more conducive to achieving the goal. In the present invention, it can help determine the long-term value of the robotic arm in a certain position and posture for completing the task of installing a special-shaped windshield glass.

[0053] In one embodiment of the present invention, the state value function is determined by the following formula:

[0054]

[0055] In the formula, v * (s) is the state value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may be generated with a certain probability at the next moment, R is the reward term of the control policy, P is the transition probability, α is the action executed at the current moment, and S is the state space.

[0056] In this embodiment, the state of the state space (S) is jointly determined by the position and attitude of the robotic arm and the readings of the six-dimensional force control sensor. The state space can be divided into the following categories: S1: Initial approach state: The robotic arm carries the glass and arrives near the machine frame, and there are large position and attitude deviations between the glass and the machine frame. S2: Preliminary alignment state: The position and attitude deviations between the glass and the machine frame have decreased, but still do not meet the docking requirements. S3: Fine adjustment state: The position and attitude of the glass and the machine frame are already close to the docking requirements, and fine adjustment is needed. S4: Docking completed state: The glass and the machine frame are perfectly docked, and the docking error is within the allowable range. Action space (A) In each state, the actions that the robotic arm can take are: A1: Translate along the X-axis: Move the robotic arm along the X-axis. A2: Translate along the Y-axis: Move the robotic arm along the Y-axis. A3: Translate along the Z-axis: Move the robotic arm along the Z-axis. A4: Rotate around the X-axis: Rotate the robotic arm around the X-axis. A5: Rotate around the Y-axis: Rotate the robotic arm around the Y-axis. A6: Rotate around the Z-axis: Rotate the robotic arm around the Z-axis. A7: Maintain the current state: Do not perform any movement or rotation operations. The reward term of the control strategy is used to measure the gain or cost obtained by taking a certain action in a certain state. For example: a. Successfully transferring from one state to a state closer to the completed docking state gives a positive reward, such as +10 points. b. Maintaining the current state or transferring to a state farther from the completed docking state gives a negative reward, such as -5 points. c. Completing the docking (transferring to state S4) gives a large positive reward, such as +100 points. d. If the readings of the force control sensor exceed the safe range during the operation, a very large negative reward is given, such as -100 points. The transition probability P represents the probability of transferring to the next state after taking a certain action in the current state. These probabilities can be estimated based on the readings of the six-dimensional force control sensor and the kinematic model of the robotic arm. For example: When taking action A1 (translate along the X-axis) in state S2 (preliminary alignment state), there is a 0.7 probability of transferring to state S3 (fine adjustment state), a 0.2 probability of remaining in state S2, and a 0.1 probability of transitioning back to state S1 (initial approach state) due to overtranslation.

[0057] In an embodiment of the present invention, the action value function is determined by the following formula:

[0058]

[0059] In the formula, q * (s, α) is the action value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may occur with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the action executed at the current moment, α' is the action executed at the next moment, and S is the state space.

[0060] The above describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0061] According to an embodiment of another aspect, the present invention provides a special-shaped windshield glass installation device. Figure 2 A schematic block diagram showing a special-shaped windshield glass installation device according to an embodiment is shown. It can be understood that the device can be implemented by any device, equipment, platform, and cluster of devices with computing and processing capabilities. As Figure 2 shown, the device includes: an acquisition unit 200, a first data processing unit 202, a second data processing unit 204, a third data processing unit 206, a loop unit 208, and a control unit 210. The main functions of each component unit are as follows:

[0062] The acquisition unit 200 is configured to initialize the control strategy of the robotic arm for installing the special-shaped windshield glass and acquire the environmental state of the robotic arm at the current moment;

[0063] The first data processing unit 202 is configured to determine the reward item of the control strategy based on the environmental state at the current moment and the control strategy;

[0064] The second data processing unit 204 is configured to update the value function according to the reward item of the control strategy and determine the updated value function;

[0065] The third data processing unit 206 is configured to optimize the control strategy according to the updated value function to obtain the optimized control strategy;

[0066] The loop unit 208 is configured to loop through steps 102 to 106 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robotic arm;

[0067] The control unit 210 is configured to control the robotic arm to install the special-shaped windshield glass according to the final control strategy.

[0068] As a preferred embodiment, the determining the reward item of the control strategy based on the environmental state at the current moment and the control strategy includes:

[0069] Based on the environmental state at the current moment and the control strategy, determine the execution action of the robotic arm at the current moment;

[0070] Determine the environmental state at the next moment based on the execution action and transition probability at the current moment;

[0071] Determine the reward term of the control strategy based on the execution action at the current moment and the environmental state at the next moment.

[0072] As a preferred implementation manner, the optimizing the control strategy according to the updated value function to obtain the optimized control strategy includes:

[0073] For each actionable action in the environmental state at the current moment, determine the value numerical corresponding to each action by means of the updated value function; wherein, the value numerical corresponding to each action is the expected cumulative reward that can be obtained after executing this action;

[0074] Sort the value numerical corresponding to all actions from high to low to obtain a sorting table of the value numerical corresponding to the actions;

[0075] Based on the sorting table, determine the action with the largest value numerical and label it as the optimal action;

[0076] Replace the action originally selected by the current strategy in this state with the determined optimal action to obtain the optimized control strategy.

[0077] As a preferred implementation manner, the convergence of the value function is determined through the following steps:

[0078] After each iterative update of the value function, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration;

[0079] Determine the difference with the largest absolute value from all calculated differences;

[0080] If the difference with the largest absolute value is less than or equal to a preset threshold, determine that the value function converges.

[0081] As a preferred implementation manner, the value function includes a state value function and an action value function.

[0082] As a preferred implementation manner, the state value function is determined by the following formula:

[0083]

[0084] In the formula, v *(s) is the state value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may be generated with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the execution action at the current moment, and S is the state space.

[0085] As a preferred embodiment, the action value function is determined by the following formula:

[0086]

[0087] In the formula, q * (s, α) is the action value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may be generated with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the execution action at the current moment, α' is the execution action at the next moment, and S is the state space.

[0088] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in combination with Figure 1 what is described.

[0089] According to an embodiment of still another aspect, there is also provided an electronic device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in combination with Figure 1 is implemented.

[0090] Each embodiment in the present invention is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment.

[0091] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0092] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for installing a special-shaped windshield, characterized in that: include: Step 100, initializing the control strategy of the mechanical arm for installing the special-shaped windshield glass, and obtaining the current environmental state of the mechanical arm; Step 102: Determine a reward item of the control strategy based on the current environmental state and the control strategy; Step 104: updating the value function according to the reward item of the control strategy to obtain an updated value function; Step 106: Optimize the control strategy according to the updated value function to obtain an optimized control strategy; Step 108, looping through steps 102 to 106 until the updated value function converges, and using the optimized control strategy at this time as the final control strategy of the robotic arm; Step 110: Control the robotic arm to install the special-shaped windshield according to the final control strategy.

2. The method according to claim 1, characterized in that The step of determining a reward item of the control strategy based on the current environmental state and the control strategy includes: Determining the execution action of the robot arm at the current moment based on the environmental state at the current moment and the control strategy; Determine the state of the environment at the next moment based on the execution action and the transition probability at the current moment; Based on the execution action at the current moment and the environmental state at the next moment, a reward item of the control strategy is determined.

3. The method according to claim 1, characterized in that The step of optimizing the control strategy according to the updated value function to obtain the optimized control strategy includes: For each possible action under the current environmental state, the updated value function is used to determine the value value corresponding to each action; wherein the value value corresponding to each action is the expected cumulative reward that can be obtained after executing the action; Sort the value values ​​corresponding to all actions from high to low to obtain a sorting table of the value values ​​corresponding to the actions; Based on the ranking table, determine the action with the largest value and mark it as the optimal action; Replace the action originally selected by the current strategy in this state with the determined optimal action to obtain the optimized control strategy.

4. The method according to claim 1, characterized in that: The cost function is determined to converge by the following steps: After updating the value function in each iteration, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration; From all calculated differences, determine the difference with the largest absolute value; If the maximum absolute value difference is less than or equal to a preset threshold, it is determined that the cost function converges.

5. The method according to claim 1, characterized in that The value function includes a state value function and an action value function.

6. The method according to claim 5, characterized in that The state value function is determined by the following formula: In the formula, v * (s) is the state value function, γ is the attenuation coefficient, s represents the environment state at the next moment, s′ is the environment state with probability at the next moment, R is the reward item of the control strategy, P is the transition probability, α is the execution action at the current moment, and S is the state space.

7. The method according to claim 5, characterized in that The action value function is determined by the following formula: In the formula, q * (s, α) is the action value function, γ is the attenuation coefficient, s represents the environment state at the next moment, s′ is the environment state with probability at the next moment, R is the reward item of the control strategy, P is the transition probability, α is the execution action at the current moment, α′ is the execution action at the next moment, and S is the state space.

8. A device for installing a special-shaped windshield, characterized in that: include: An acquisition unit is configured to initialize a control strategy of a mechanical arm for installing a special-shaped windshield glass and acquire an environmental state of the mechanical arm at a current moment; A first data processing unit is configured to determine a reward item of the control strategy based on the current environmental state and the control strategy; A second data processing unit is configured to update the value function according to the reward item of the control strategy and determine the updated value function; A third data processing unit is configured to optimize the control strategy according to the updated value function to obtain an optimized control strategy; The loop unit is configured to loop through steps 102 to 106 until the updated value function converges, and the optimized control strategy at this time is used as the final control strategy of the robot arm; The control unit is configured to control the mechanical arm to install the special-shaped windshield according to the final control strategy.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Aircraft windshield mounting device and mounting method

    CN121106734A

  • An aircraft windshield mounting device and method of mounting

    CN121106734B