Man-machine cooperation safety protection method and device for composite assembly robot

By obtaining point cloud and image data of assembly personnel and composite assembly robots, combining safe spacing monitoring and robotic arm optimization control, manual operation problems during the assembly process of special-shaped windshield glass are solved, ensuring safety and accuracy, and improving assembly efficiency and product quality.

CN120396010APending Publication Date: 2025-08-01SHENZHEN POLYTECHNIC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510546805.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

During the assembly process of special-shaped windshield glass, it is difficult to stably grasp large-sized and heavy parts manually, which poses a risk of collision, and existing composite assembly robots have challenges in terms of safety and accuracy.

Method used

By obtaining point cloud data and image data of assembly personnel and composite assembly robots, the safe spacing is determined using a combination of point cloud data and image data, and alarm information is sent when the safety spacing is insufficient, combining the real-time monitoring and optimization control strategies of the robot arm to ensure safety and accuracy.

Benefits of technology

It realizes effective safety protection for assembly personnel during the assembly process of special-shaped windshield glass, improves assembly accuracy and yield rate, and reduces production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120396010A_ABST
    Figure CN120396010A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a man-machine cooperation safety protection method and device for a composite assembly robot. According to the method, when the special-shaped windshield glass is installed, point cloud data of assembly personnel and point cloud data and image data of a composite assembly robot are obtained; wherein the image data comprises assembly personnel and a composite assembly robot; determining a first protection distance based on the point cloud data of the assembly personnel and the point cloud data of the composite assembly robot; determining a second guard distance based on the image data; determining a third protection distance based on the first protection distance and the second protection distance; and when the third protection distance is smaller than or equal to the preset threshold value, alarm information is sent to a user side of the assembly personnel, the personal safety of the assembly personnel in the operation process of the composite assembly robot can be practically and efficiently guaranteed, and a firm and reliable safety protection system is constructed for aviation assembly operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a human-robot collaboration safety protection method and device for a composite assembly robot. Background Art

[0002] In the field of installation of special-shaped windshield glass, its operation mode is unique. Usually, it is not mass-produced, and the assembly link highly relies on manual operation assistance. Although composite robots can theoretically be used to grasp the object to be installed and accurately transport it to the assembly position, so as to achieve stable and reliable assembly of large components, in the actual application process, manual assistance is still required. Especially for the assembly work of large-size and large-weight components, many severe challenges are faced. When operating manually, it is difficult for the operator to stably grasp the object to be installed, and there are also great difficulties in controlling the contact force of the installation interface and ensuring the installation accuracy. Moreover, during the entire installation process, the risk of collision between components remains extremely high, which poses a serious threat to product quality and production safety.

[0003] Applying a composite assembly robot to a complex and extremely safety-demanding working environment such as the installation of special-shaped windshield glass and ensuring the safety of surrounding personnel during its operation have become key problems to be solved urgently.

[0004] Based on this, the present invention proposes a human-robot collaboration safety protection method and device for a composite assembly robot to solve the above technical problems. Summary of the Invention

[0005] The present invention describes a human-robot collaboration safety protection method and device for a composite assembly robot, which can ensure the safety of assembly personnel during its operation.

[0006] According to a first aspect, the present invention provides a human-robot collaboration safety protection method for a composite assembly robot, including:

[0007] When installing special-shaped windshield glass, obtaining the point cloud data of the assembly personnel, the point cloud data and image data of the composite assembly robot; wherein, the image data includes the assembly personnel and the composite assembly robot;

[0008] Based on the point cloud data of the assembly personnel and the point cloud data of the composite assembly robot, determining a first protection distance;

[0009] Based on the image data, determining a second protection distance;

[0010] Based on the first protection distance and the second protection distance, determining a third protection distance;

[0011] When the third protection distance is less than or equal to a preset threshold, sending an alarm message to the user terminal of the assembly personnel.

[0012] According to a second aspect, the present invention provides a human - machine collaboration safety protection device for a composite assembly robot, including:

[0013] An acquisition unit, configured to acquire the point cloud data of an assembly worker, the point cloud data of the composite assembly robot, and image data when installing a special - shaped windshield; wherein, the image data includes the assembly worker and the composite assembly robot;

[0014] A first data processing unit, configured to determine a first protection distance based on the point cloud data of the assembly worker and the point cloud data of the composite assembly robot;

[0015] A second data processing unit, configured to determine a second protection distance based on the image data;

[0016] A third data processing unit, configured to determine a third protection distance based on the first protection distance and the second protection distance;

[0017] A fourth data processing unit, configured to send an alarm message to the user terminal of the assembly worker when the third protection distance is less than or equal to a preset threshold.

[0018] According to the human - machine collaboration safety protection method and device for a composite assembly robot provided by the present invention, when carrying out the installation operation of a special - shaped windshield, three - dimensional laser scanning and other devices are used to synchronously collect the point cloud data of the assembly worker and the composite assembly robot. At the same time, an image acquisition device such as an industrial camera is used to acquire the image data including the assembly worker and the composite assembly robot. First, based on the point cloud data of the assembly worker and the composite assembly robot, algorithms such as spatial coordinate transformation and Euclidean distance calculation are used to obtain the first protection distance. This distance reflects the safety interval under the spatial position relationship between the two presented by the point cloud data. Then, the acquired image data is input into a convolutional neural network (CNN) trained with a large number of samples. The model determines the second protection distance through feature extraction, classification, and localization analysis of the target objects in the image. This distance is the safety interval evaluated based on the image visual information. Subsequently, based on the first protection distance and the second protection distance, the third protection distance is determined. This distance integrates the safety information contained in the point cloud data and the image data, and more comprehensively reflects the safety situation in actual operations. Once the third protection distance is less than or equal to the safety threshold set in advance through risk assessment, an alarm message will be immediately sent to the user terminal of the assembly worker in the form of message push through a wireless communication module. Through the above configuration method, the present invention can effectively and efficiently ensure the personal safety of the assembly worker during the operation of the composite assembly robot, and build a solid and reliable safety protection system for aviation assembly operations. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 The flowchart showing the method for human - machine collaboration safety protection of a composite assembly robot according to an embodiment is presented;

[0021] Figure 2 The schematic block diagram showing the device for human - machine collaboration safety protection of a composite assembly robot according to an embodiment is presented. Detailed implementation manners

[0022] The following describes the solution provided by the present invention in conjunction with the drawings.

[0023] Figure 1 The flowchart showing the method for human - machine collaboration safety protection of a composite assembly robot according to an embodiment is presented. It can be understood that this method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. As Figure 1 shown, this method includes:

[0024] Step 100: When installing a special - shaped windshield glass, obtain the point - cloud data of the assembly personnel, the point - cloud data of the composite assembly robot, and the image data; wherein, the image data includes the assembly personnel and the composite assembly robot;

[0025] Step 102: Determine the first protection distance based on the point - cloud data of the assembly personnel and the point - cloud data of the composite assembly robot;

[0026] Step 104: Determine the second protection distance based on the image data;

[0027] Step 106: Determine the third protection distance based on the first protection distance and the second protection distance;

[0028] Step 108: When the third protection distance is less than or equal to the preset threshold, send an alarm message to the user terminal of the assembly personnel.

[0029] In this embodiment, when carrying out the installation operation of the special-shaped windshield glass, 3D laser scanning and other devices are used to synchronously collect the point cloud data of the assembly personnel and the composite assembly robot. At the same time, image data including the assembly personnel and the composite assembly robot is obtained through image acquisition devices such as industrial cameras. First, based on the point cloud data of the assembly personnel and the composite assembly robot, algorithms such as spatial coordinate transformation and Euclidean distance calculation are used to obtain the first protection distance. This distance reflects the safety interval under the spatial position relationship between the two presented by the point cloud data. Subsequently, the obtained image data is input into a convolutional neural network (CNN) trained with a large number of samples. The model determines the second protection distance through feature extraction, classification, and localization analysis of the target objects in the image. This distance is the safety interval evaluated based on the image visual information. Then, based on the first protection distance and the second protection distance, the third protection distance is determined. This distance integrates the safety information contained in the point cloud data and the image data, and more comprehensively reflects the safety status in the actual operation. Once the third protection distance is less than or equal to the safety threshold set in advance through risk assessment, an alarm message will be immediately sent to the user terminal of the assembly personnel in the form of message push through the wireless communication module. Through the above configuration method, the present invention can effectively and efficiently ensure the personal safety of the assembly personnel during the operation of the composite assembly robot, and build a solid and reliable safety protection system for the aviation assembly operation.

[0030] In an embodiment of the present invention, determining the third protection distance based on the first protection distance and the second protection distance includes:

[0031] The first protection distance and the second protection distance are weighted according to a preset ratio to obtain the third protection distance.

[0032] In this embodiment, regarding determining the third protection distance based on the first protection distance and the second protection distance, the specific operation is as follows: According to the ratio set in advance through a large amount of experimental data and risk assessment analysis, corresponding weight coefficients are assigned to the first protection distance and the second protection distance, and a weighted operation is performed. That is, by multiplying the first protection distance by its corresponding weight and adding the second protection distance multiplied by its corresponding weight, the third protection distance is finally obtained. This calculation method fully integrates the protection distances determined from two different data sources and can more accurately and scientifically reflect the safety interval of the actual operation scenario.

[0033] In an embodiment of the present invention, when installing the special-shaped windshield glass, the following steps are executed:

[0034] Step 200: Initialize the control strategy of the manipulator for installing the special-shaped windshield glass, and obtain the environmental state of the manipulator at the current moment;

[0035] Step 202: Determine the reward items of the control strategy based on the environmental state and control strategy at the current moment;

[0036] Step 204: Update the value function according to the reward items of the control strategy to obtain the updated value function;

[0037] Step 206: Optimize the control strategy according to the updated value function to obtain the optimized control strategy;

[0038] Step 208: Loop through steps 202 - 206 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robotic arm;

[0039] Step 210: Control the robotic arm to install the special-shaped windshield according to the final control strategy.

[0040] In this embodiment, the control strategy of the robotic arm undertaking the human - machine collaboration safety protection task of the composite assembly robot is initialized. Meanwhile, high - precision sensors such as lidar and industrial cameras are deployed to realize multi - dimensional and real - time monitoring of the environment where the robotic arm is located, obtain the current environmental state data of the robotic arm, analyze each action instruction of the robotic arm based on the real - time state information and the established control strategy of the robotic arm, and quantitatively output the corresponding reward value. After obtaining the reward value, the value function is immediately updated. Subsequently, the current control strategy is deeply optimized to output a more scientific and reasonable control strategy. To ensure the optimization of the control strategy, operations such as value function update and control strategy optimization are continuously looped. Finally, the robotic arm installs the special - shaped windshield according to this final control strategy. Through the above configuration method, the present invention significantly improves the installation accuracy of the special - shaped windshield. After actual measurement, compared with the traditional installation method, the qualified product rate of the product is greatly improved, the production cycle is effectively shortened, and the production cost of the enterprise is also reduced.

[0041] In an embodiment of the present invention, determining the reward items of the control strategy based on the environmental state and control strategy at the current moment includes:

[0042] Determine the execution action of the robotic arm at the current moment based on the environmental state and control strategy at the current moment;

[0043] Determine the environmental state at the next moment based on the execution action and transition probability at the current moment;

[0044] Determine the reward items of the control strategy based on the execution action at the current moment and the environmental state at the next moment.

[0045] In this embodiment, various parameters of the environment where the robotic arm is located are obtained in real time and accurately through devices such as omnidirectionally deployed lidar and vision sensors. Based on this data, the optimal execution action of the robotic arm at the current moment is quickly screened out. To predict the environmental changes caused by the actions of the robotic arm, a state transition model trained based on a large amount of historical data is used. Referring to the transition probabilities set in the model, the environmental state at the next moment after the execution of the action is effectively predicted, and a new state of the robotic arm, the special-shaped windshield, and the surrounding working environment is simulated. After the environmental state prediction at the next moment is completed, relying on the pre-established reward evaluation model, the current execution action of the robotic arm and the predicted environmental state at the next moment are substituted into it. The reward item corresponding to the control strategy is quantitatively obtained, providing data support for the subsequent optimization of the control strategy.

[0046] In one embodiment of the present invention, the control strategy is optimized according to the updated value function to obtain the optimized control strategy, including:

[0047] For each actionable action in the environmental state at the current moment, with the help of the updated value function, determine the value numerical corresponding to each action; where the value numerical corresponding to each action is the expected cumulative reward that can be obtained after executing this action;

[0048] Sort the value numerical corresponding to all actions from high to low to obtain a sorting table of the value numerical corresponding to the actions;

[0049] Based on the sorting table, determine the action with the largest value numerical and label it as the optimal action;

[0050] Replace the action originally selected by the current strategy in this state with the determined optimal action to obtain the optimized control strategy.

[0051] In this embodiment, optimizing the control strategy to improve the operation efficiency of the robotic arm is the core task. After the value function is updated, these latest data are used to finely adjust the control strategy. The specific optimization steps are as follows: First, for each feasible operation action of the robotic arm in the current environmental state, the present invention deeply evaluates it with the updated value function. Through a complex and precise calculation logic, the value corresponding to each action is determined. Here, the value represents the expected cumulative reward that can be obtained in the future period after executing this action, which comprehensively considers the impacts of the action on task completion, efficiency, quality, etc. Then, the present invention collects the values corresponding to all actions and arranges them in descending order. This process generates a detailed action value ranking table, clearly showing the pros and cons order of each action in the current environmental state. Subsequently, based on this ranking table, the present invention can quickly identify the action with the largest value. This action is recognized as the optimal choice in the current environmental state because it indicates that it can bring the largest cumulative reward to the robotic arm and helps to complete the installation task of the special-shaped windshield glass more efficiently and accurately. Finally, the present invention replaces the action originally selected by the current control strategy in this state with the newly determined optimal action. Through this replacement operation, the control strategy is optimized and can better adapt to the environmental state where the robotic arm is currently located, thereby further improving the installation accuracy and efficiency of the special-shaped windshield glass. After such an optimization process, the finally obtained is the optimized control strategy.

[0052] In an embodiment of the present invention, the convergence of the value function is determined through the following steps:

[0053] After each iterative update of the value function, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration;

[0054] From all the calculated differences, determine the difference with the largest absolute value;

[0055] If the difference with the largest absolute value is less than or equal to the preset threshold, it is determined that the value function converges.

[0056] In this embodiment, the difference calculation: Each time an iteration update of the value function is completed, the difference calculation program is started. The value function obtained in this iteration is compared with the result of the previous iteration to calculate the difference between the two. Maximum difference screening: After all differences are calculated, the present invention comprehensively sorts out these difference data. With the help of a data screening algorithm, the difference with the largest absolute value is quickly locked among numerous differences. This largest difference can intuitively reflect the part where the value function changes most significantly between two iterations. Convergence determination: The difference with the largest absolute value selected is compared with a preset threshold. If the largest difference is less than or equal to the preset threshold, it means that the change range of the value function between two iterations is within an acceptable range and is in a relatively stable state. At this time, it can be determined that the value function has converged.

[0057] In an embodiment of the present invention, the value function includes a state value function and an action value function.

[0058] In this embodiment, the action value function: It describes the expectation of the long-term cumulative reward that can be obtained by taking a certain action in a specific state following the current policy. The action value function not only considers the state but also clarifies the value difference of taking different actions in this state. In the present invention, the action value function can be used to compare the values of performing different movements, grasping, etc. for finally completing the task in the current position and posture of the robotic arm. State value function: It represents the expectation of the long-term cumulative reward that can be obtained in a certain state following the current policy. By calculating the state value function, the quality of different states can be evaluated to help determine which states are more conducive to achieving the goal. In the present invention, it can help judge the long-term value of the robotic arm in a certain position and posture for completing the task of installing a special-shaped windshield glass.

[0059] In an embodiment of the present invention, the state value function is determined by the following formula:

[0060]

[0061] In the formula, v * (s) is the state value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may occur with a certain probability at the next moment, R is the reward term of the control policy, P is the transition probability, α is the action executed at the current moment, and S is the state space.

[0062] In this embodiment, the state of the state space (S) is jointly determined by the position and orientation of the robotic arm and the readings of the six-dimensional force control sensor. The state space can be divided into the following categories: S1: Initial approach state: The robotic arm carries the glass and arrives near the machine frame, and there are large position and orientation deviations between the glass and the machine frame. S2: Preliminary alignment state: The position and orientation deviations between the glass and the machine frame have decreased, but the docking requirements have not been met yet. S3: Fine adjustment state: The position and orientation of the glass and the machine frame are close to the docking requirements, and fine adjustment is needed. S4: Docking completed state: The glass and the machine frame are perfectly docked, and the docking error is within the allowable range. Action space (A) In each state, the actions that the robotic arm can take are: A1: Translate the X-axis: Move the robotic arm along the X-axis. A2: Translate the Y-axis: Move the robotic arm along the Y-axis. A3: Translate the Z-axis: Move the robotic arm along the Z-axis. A4: Rotate the X-axis: Rotate the robotic arm around the X-axis. A5: Rotate the Y-axis: Rotate the robotic arm around the Y-axis. A6: Rotate the Z-axis: Rotate the robotic arm around the Z-axis. A7: Maintain the current state: Do not perform any movement or rotation operations. The reward term of the control strategy is used to measure the gain or cost obtained by taking a certain action in a certain state. For example: a. Successfully transferring from one state to a state closer to the docking completed state gives a positive reward, such as +10 points. b. Maintaining the current state or transferring to a state farther from the docking completed state gives a negative reward, such as -5 points. c. Completing the docking (transferring to state S4) gives a large positive reward, such as +100 points. d. If the readings of the force control sensor exceed the safe range during the operation, a very large negative reward, such as -100 points, is given. The transition probability P represents the probability of transferring to the next state after taking a certain action in the current state. These probabilities can be estimated based on the readings of the six-dimensional force control sensor and the kinematic model of the robotic arm. For example: When taking action A1 (translate the X-axis) in state S2 (preliminary alignment state), there is a 0.7 probability of transferring to state S3 (fine adjustment state), a 0.2 probability of remaining in state S2, and a 0.1 probability of transitioning back to state S1 (initial approach state) due to overtranslation.

[0063] In an embodiment of the present invention, the action value function is determined by the following formula:

[0064]

[0065] In the formula, q * (s, α) is the action value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may occur with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the action executed at the current moment, α' is the action executed at the next moment, and S is the state space.

[0066] The above describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0067] According to an embodiment of another aspect, the present invention provides a human - machine collaboration safety protection device for a composite assembly robot. Figure 2 A schematic block diagram showing a human - machine collaboration safety protection device for a composite assembly robot according to an embodiment is shown. It can be understood that the device can be implemented by any device, equipment, platform, and cluster of devices with computing and processing capabilities. As Figure 2 shown, the device includes: an acquisition unit 300, a first data processing unit 302, a second data processing unit 304, a third data processing unit 306, and a fourth data processing unit 308. The main functions of each component unit are as follows:

[0068] The acquisition unit 300 is configured to acquire the point cloud data of the assembly personnel, the point cloud data of the composite assembly robot, and the image data when installing a special - shaped windshield glass; wherein, the image data includes the assembly personnel and the composite assembly robot;

[0069] The first data processing unit 302 is configured to determine a first protection distance based on the point cloud data of the assembly personnel and the point cloud data of the composite assembly robot;

[0070] The second data processing unit 304 is configured to determine a second protection distance based on the image data;

[0071] The third data processing unit 306 is configured to determine a third protection distance based on the first protection distance and the second protection distance;

[0072] The fourth data processing unit 308 is configured to send an alarm message to the user of the assembly personnel when the third protection distance is less than or equal to a preset threshold.

[0073] As a preferred embodiment, the determining the third protection distance based on the first protection distance and the second protection distance includes:

[0074] Weighting the first protection distance and the second protection distance according to a preset ratio to obtain the third protection distance.

[0075] As a preferred embodiment, when installing the special - shaped windshield glass, the following steps are performed:

[0076] Step 200: Initialize the control strategy of the robotic arm for installing the special-shaped windshield glass, and obtain the environmental state of the robotic arm at the current moment;

[0077] Step 202: Based on the environmental state at the current moment and the control strategy, determine the reward items of the control strategy;

[0078] Step 204: Update the value function according to the reward items of the control strategy to obtain the updated value function;

[0079] Step 206: Optimize the control strategy according to the updated value function to obtain the optimized control strategy;

[0080] Step 208: Loop through steps 202 - 206 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robotic arm;

[0081] Step 210: Control the robotic arm to install the special-shaped windshield glass according to the final control strategy.

[0082] As a preferred implementation manner, the determining the reward items of the control strategy based on the environmental state at the current moment and the control strategy includes:

[0083] Based on the environmental state at the current moment and the control strategy, determine the execution action of the robotic arm at the current moment;

[0084] Based on the execution action at the current moment and the transition probability, determine the environmental state at the next moment;

[0085] Based on the execution action at the current moment and the environmental state at the next moment, determine the reward items of the control strategy.

[0086] As a preferred implementation manner, the optimizing the control strategy according to the updated value function to obtain the optimized control strategy includes:

[0087] For each actionable action in the environmental state at the current moment, with the help of the updated value function, determine the value numerical corresponding to each action; wherein, the value numerical corresponding to each action is the expected cumulative reward that can be obtained after executing this action;

[0088] Sort the value numerical corresponding to all actions from high to low to obtain the sorting table of the value numerical corresponding to the actions;

[0089] Based on the sorting table, determine the action with the largest value numerical and mark it as the optimal action;

[0090] Replace the action originally selected by the current policy in this state with the determined optimal action to obtain the optimized control policy.

[0091] As a preferred implementation, the value function is determined to converge through the following steps:

[0092] After each iteration of updating the value function, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration;

[0093] Determine the difference with the largest absolute value from all the calculated differences;

[0094] If the difference with the largest absolute value is less than or equal to a preset threshold, it is determined that the value function converges.

[0095] As a preferred implementation, the value function includes a state value function and an action value function.

[0096] As a preferred implementation, the state value function is determined by the following formula:

[0097]

[0098] In the formula, v * (s) is the state value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may occur with a certain probability at the next moment, R is the reward term of the control policy, P is the transition probability, α is the action executed at the current moment, and S is the state space.

[0099] As a preferred implementation, the action value function is determined by the following formula:

[0100]

[0101] In the formula, q * (s, α) is the action value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may occur with a certain probability at the next moment, R is the reward term of the control policy, P is the transition probability, α is the action executed at the current moment, α' is the action executed at the next moment, and S is the state space.

[0102] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0103] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0104] The specific implementation manners described above have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included within the protection scope of the present invention.

Claims

1. A human-machine collaboration safety protection method for a composite assembly robot, characterized in that, Including: When installing a special-shaped windshield, obtain the point cloud data of the assembly worker, the point cloud data and image data of the composite assembly robot; wherein, the image data includes the assembly worker and the composite assembly robot; Based on the point cloud data of the assembly worker and the point cloud data of the composite assembly robot, determine the first protection distance; Based on the image data, determine the second protection distance; Based on the first protection distance and the second protection distance, determine the third protection distance; When the third protection distance is less than or equal to a preset threshold, send an alarm message to the user terminal of the assembly worker.

2. The method according to claim 1, wherein The determining the third protection distance based on the first protection distance and the second protection distance includes: Weight the first protection distance and the second protection distance according to a preset ratio to obtain the third protection distance.

3. The method according to claim 1, characterized in that, When installing the special-shaped windshield, perform the following steps: Step 200: Initialize the control strategy of the robot arm for installing the special-shaped windshield, and obtain the environmental state of the robot arm at the current moment; Step 202: Based on the environmental state at the current moment and the control strategy, determine the reward item of the control strategy; Step 204: Update the value function according to the reward item of the control strategy to obtain the updated value function; Step 206: Optimize the control strategy according to the updated value function to obtain the optimized control strategy; Step 208: Loop through steps 202 to 206 until the updated value function converges, and use the optimized control strategy at this time as the final control strategy of the robot arm; Step 210: Control the robot arm to install the special-shaped windshield according to the final control strategy.

4. The method according to claim 3, wherein The determining the reward item of the control strategy based on the environmental state at the current moment and the control strategy includes: Based on the environmental state at the current moment and the control strategy, determine the execution action of the robot arm at the current moment; Based on the execution action at the current moment and the transition probability, determine the environmental state at the next moment; Based on the execution action at the current moment and the environmental state at the next moment, determine the reward item of the control strategy.

5. The method according to claim 3, characterized in that, The optimizing the control strategy according to the updated value function to obtain the optimized control strategy includes: For each actionable action in the environmental state at the current moment, with the help of the updated value function, determine the value corresponding to each action; wherein, the value corresponding to each action is the expected cumulative reward that can be obtained after executing this action Sort the value corresponding to all actions from high to low to obtain a sorted list of the value corresponding to the actions; Based on the sorted list, determine the action with the largest value and label it as the optimal action; Replace the action originally selected by the current strategy in this state with the determined optimal action to obtain the optimized control strategy.

6. The method according to claim 3, wherein The convergence of the value function is determined through the following steps: After each iterative update of the value function, calculate the difference between the value function obtained in the current iteration and the result of the previous iteration; From all the calculated differences, determine the difference with the largest absolute value; If the maximum absolute difference is less than or equal to a preset threshold, it is determined that the value function converges.

7. The method according to claim 3, wherein The value function includes a state value function and an action value function.

8. The method according to claim 7, wherein The state value function is determined by the following formula: where, v * (s) is the state value function, γ is the attenuation coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may be generated with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the execution action at the current moment, and S is the state space.

9. The method according to claim 7, wherein The action value function is determined by the following formula: where q * (s, α) is the action-value function, γ is the decay coefficient, s represents the environmental state at the next moment, s′ is the environmental state that may be generated with a certain probability at the next moment, R is the reward term of the control strategy, P is the transition probability, α is the execution action at the current moment, α' is the execution action at the next moment, and S is the state space.

10. A human-machine collaboration safety protection device for a composite assembly robot, characterized in that, including: An acquisition unit configured to acquire the point cloud data of the assembler, the point cloud data of the composite assembly robot, and the image data when the special-shaped windshield is installed; wherein, the image data includes the assembler and the composite assembly robot; A first data processing unit configured to determine a first protection distance based on the point cloud data of the assembler and the point cloud data of the composite assembly robot; A second data processing unit configured to determine a second protection distance based on the image data; A third data processing unit configured to determine a third protection distance based on the first protection distance and the second protection distance; A fourth data processing unit configured to send an alarm message to the user of the assembler when the third protection distance is less than or equal to a preset threshold.