Mobile phone assembly feeding method, system and device based on reinforcement learning

By introducing reinforcement learning algorithms and intelligent control systems, the problems of low efficiency, insufficient precision, poor flexibility, and low level of intelligence in the mobile phone assembly and feeding process have been solved, realizing efficient, accurate, and intelligent feeding operations, and improving production efficiency and product quality.

WO2026000240A1PCT designated stage Publication Date: 2026-01-02TSINGHUA UNIVERSITY
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/101600
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing mobile phone assembly and feeding methods are inefficient, lack precision, have poor flexibility and adaptability, and have low levels of intelligence, failing to meet the demands for efficient, precise, and flexible production.

Method used

The system employs reinforcement learning algorithms to collect state data through sensor modules, uses a PLC system to control the robotic arm to perform optimal motion strategies, and combines high-precision cameras and sensors for real-time data feedback and optimization. Through reinforcement learning model generation, it achieves intelligent control, forms a closed-loop system, optimizes the feeding path and strategy, and realizes intelligent operation.

Benefits of technology

It achieves efficient, precise, and flexible intelligent control, significantly improving production efficiency, flexibility, and adaptability. It solves several problems existing in the current technology, improves the accuracy and efficiency of feeding operations, increases the flexibility and intelligence of feeding, and enhances production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024101600_02012026_PF_FP_ABST
    Figure CN2024101600_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a mobile phone assembly feeding method, system and device based on reinforcement learning. The method comprises: using a sensing module to collect state data of a mechanical arm in the process of executing a feeding task, wherein the state data comprises environment state information and an own action state; on the basis of a preset reward function, using a reinforcement learning algorithm to learn the state data to train a reinforcement learning feeding policy model, and generating an optimal action policy on the basis of the trained reinforcement learning feeding policy model; and using a PLC system to perform instruction parsing on the optimal action policy, so that, on the basis of an instruction obtained by parsing, a motor of the mechanical arm and related execution elements are controlled to conduct control operation according to the optimal action policy. The present disclosure aims to solve a plurality of problems in existing mobile phone assembly feeding processes, especially how to improve the accuracy and efficiency of a feeding operation, and increase the flexibility and adaptability of feeding and the degree of intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

A reinforcement learning-based method, system, and device for mobile phone assembly and feeding. Technical Field

[0001] This disclosure relates to the field of reinforcement learning, and in particular to reinforcement learning and mobile phone assembly. Background Technology

[0002] With the rapid development of smartphones and the continuous growth of market demand, the mobile phone manufacturing industry has increasingly higher requirements for production efficiency and product quality. Traditional mobile phone assembly and loading methods mainly rely on manual operation or simple mechanical automation, which have many shortcomings. For example, manual operation is inefficient and easily affected by the skill level and working condition of workers, making it difficult to guarantee product consistency and quality; while traditional mechanical automation systems lack flexibility and are difficult to adapt to changing production needs and complex assembly tasks.

[0003] In recent years, with the rapid development of artificial intelligence technology, especially the maturity of machine learning and deep learning, many advanced algorithms and methods have been gradually applied to industrial production. Reinforcement learning, as a machine learning method that can continuously learn and optimize decisions through interaction with the environment, has demonstrated enormous potential in complex tasks. Reinforcement learning algorithms construct intelligent agents that gradually optimize their strategies through trial and error, thereby achieving adaptive control and efficient execution of complex tasks.

[0004] In the mobile phone assembly and loading process, the introduction of reinforcement learning algorithms can effectively overcome the limitations of traditional methods. Through real-time data acquisition and model training, reinforcement learning algorithms can quickly adapt to different assembly tasks and environmental changes, optimize loading paths and strategies, and improve loading efficiency and accuracy. Furthermore, reinforcement learning-based loading systems also possess self-learning and self-optimization capabilities, enabling continuous performance improvement and enhancement during long-term operation, while reducing maintenance and operating costs.

[0005] Several problems exist in the existing mobile phone assembly and loading process, particularly how to improve the accuracy and efficiency of the loading operation, and increase its flexibility, adaptability, and level of intelligence. Specific technical issues include:

[0006] 1. Low material loading efficiency. Traditional mobile phone assembly and material loading methods mainly rely on manual operation or simple mechanical automation systems, resulting in relatively low operational efficiency that cannot meet the high efficiency requirements of large-scale production. Manual operation is easily affected by factors such as operator fatigue and experience, making it difficult to guarantee the loading speed and consistency.

[0007] 2. Insufficient loading accuracy. Due to the large variety and precise dimensions of the parts involved in mobile phone assembly, traditional loading methods are prone to deviations in positioning and gripping parts, resulting in insufficient assembly accuracy and affecting the quality of the final product. The limited precision of traditional mechanical systems cannot meet the requirements of high-precision assembly.

[0008] 3. Poor flexibility and adaptability. With the continuous updates in mobile phone design and manufacturing processes, the tasks and environment during assembly and loading are also constantly changing. Traditional automated systems typically lack flexibility and adaptability, making it difficult to quickly adjust and adapt to new assembly tasks, resulting in high adjustment and maintenance costs for the production line.

[0009] 4. Low level of intelligence. Existing mobile phone assembly and feeding systems have a low level of intelligence, lack self-learning and self-optimization capabilities, cannot optimize and adjust based on real-time production data, and are unable to cope with various uncertainties and emergencies that occur during the production process.

[0010] Summary of the Invention

[0011] This disclosure presents a method, system, and apparatus for assembling and feeding mobile phones based on reinforcement learning.

[0012] According to one aspect of this disclosure, a mobile phone assembly and loading method based on reinforcement learning is proposed, comprising: collecting state data of a robotic arm during the execution of a loading task using a sensing module; wherein the state data includes environmental state information and its own action state; training a reinforcement learning loading strategy model by learning the state data according to a preset reward function and using a reinforcement learning algorithm, and generating an optimal action strategy based on the trained reinforcement learning loading strategy model; and parsing the optimal action strategy using a PLC system to control the motor and related actuators of the robotic arm to perform control operations according to the parsed instructions.

[0013] According to a second aspect of this disclosure, a mobile phone assembly and loading system based on reinforcement learning is proposed, comprising: a data acquisition module for acquiring state data of a robotic arm during the loading task; wherein the state data includes environmental state information and its own action state; a data processing module for preprocessing the state data sent by the data acquisition module to obtain preprocessed data; a model training module for training a reinforcement learning loading strategy model by learning the state data according to a preset reward function and using a reinforcement learning algorithm, and generating an optimal action strategy based on the trained reinforcement learning loading strategy model; and a loading operation module for parsing the optimal action strategy using a PLC system, so as to control the motor and related actuators of the robotic arm to perform control operations according to the parsed instructions.

[0014] According to a third aspect of this disclosure, a mobile phone assembly and loading device based on reinforcement learning is proposed, comprising: an end effector and a control system, wherein the control system controls the end effector to perform assembly operations; the end effector includes: a pneumatic parallel gripper with a left gripper and a right gripper disposed opposite to each other at their ends; the control system controls the pneumatic parallel gripper to open or close; a positioning pin connected to a fixed plate located on both sides of the pneumatic parallel gripper via a carrier contact baffle; the carrier contact baffle is located between the left gripper and the right gripper and is perpendicularly disposed on one side of the fixed plate; and a camera disposed on the fixed plate to determine the moving position of the positioning pin.

[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0016] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0017] Figure 1 is a flowchart of a mobile phone assembly and feeding method based on reinforcement learning according to an embodiment of this disclosure;

[0018] Figure 2 is a logic diagram of a mobile phone assembly and feeding method based on reinforcement learning according to an embodiment of the present disclosure;

[0019] Figure 3 is a structural diagram of a mobile phone assembly and feeding system based on reinforcement learning according to an embodiment of this disclosure;

[0020] Figure 4 is a schematic diagram of the structure of an end effector provided in an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of Figure 4 from another perspective;

[0022] Figure 6 is a partially enlarged schematic diagram of Figure 4;

[0023] Figure 7 is a structural schematic diagram of a mobile phone assembly carrier provided in an embodiment of this disclosure;

[0024] Figure 8 is a schematic diagram of the structure of a mobile phone assembly and loading platform provided in an embodiment of this disclosure;

[0025] Figure 9 is a schematic diagram of the structure of a mobile phone assembly operation platform provided in an embodiment of this disclosure;

[0026] Figure 10 is a schematic diagram of an end effector gripping a mobile phone assembly carrier provided in an embodiment of the present disclosure;

[0027] Figure 11 is a schematic diagram of a mobile phone assembly carrier moving to a mobile phone assembly operation platform according to an embodiment of the present disclosure;

[0028] In the diagram, 1. End connector; 2. Camera bracket; 3. Industrial camera; 4. Left fixing plate; 5. Carrier contact baffle; 6. Right gripper; 7. Positioning pin; 8. Rubber pad; 9. Left gripper; 10. Pneumatic parallel gripper; 11. Right fixing plate; 12. End effector; 13. Mobile phone assembly carrier; 14. Mobile phone assembly loading platform; 15. Mobile phone assembly operation platform. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] Machine learning (ML) is a branch of computer science that enables computers to learn from data and improve their task performance without explicit programming. In short, machine learning is about designing and developing algorithms that allow computers to learn from experience (i.e., data) and make predictions or decisions based on that learning. Data is fundamental to machine learning; the quality, quantity, and representativeness of the data directly affect the learning effect. A mathematical representation of data is used to process data to make predictions or decisions. Common models include linear regression, decision trees, and neural networks. The process of tuning model parameters using labeled or unlabeled datasets so that the model can learn patterns from the data. Models learn mappings based on known input-output pairs (training data), such as in classification and regression problems. Models search for structures or patterns in unlabeled data, often used for clustering, dimensionality reduction, etc.

[0031] Deep learning (DL) is a new research direction in the field of machine learning (ML). It was introduced into machine learning to bring it closer to its original goal—artificial intelligence. Deep learning learns the inherent laws and hierarchical representations of sample data. The information gained during this learning process greatly aids in the interpretation of data such as text, images, and sound. Its ultimate goal is to enable machines to possess analytical and learning capabilities like humans, capable of recognizing data such as text, images, and sound. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition far exceeding previous related technologies.

[0032] An agent is a core concept in the field of artificial intelligence (AI). It refers to an autonomous entity capable of perceiving its environment, making decisions, and performing actions to achieve specific goals. Agents are widely designed and applied in various complex systems, such as robots, software systems, games, economic simulations, and intelligent network management. Key characteristics and components of an agent include: The agent perceives the state of its environment through sensors or by receiving external input. In a virtual environment, this might involve reading data streams or APIs; in the physical world, it might involve cameras, sound sensors, temperature sensors, etc. Based on the perceived information, the agent uses internal algorithms (such as reinforcement learning, decision trees, genetic algorithms, etc.) to make decisions. These decisions aim to optimize certain objective functions, such as maximizing rewards, minimizing costs, or achieving a specific task. After making a decision, the agent performs corresponding actions to influence the environment. This might involve physical actions (such as a robot moving or grasping objects) or virtual operations (such as sending instructions or modifying data). The agent can learn from its interactions with the environment, optimizing its behavioral strategies through trial and error. This enables the agent to adapt to environmental changes and improve its problem-solving abilities. Intelligent agents are typically designed with specific goals or tasks in mind, and their behavior is directed to achieve those goals.

[0033] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. AI hardware technologies generally include computer vision, speech recognition, natural language processing, and related technologies such as deep learning, big data processing, and knowledge graphs.

[0034] Assembly loading refers to the process in manufacturing, especially in automated production lines or machining, of placing raw materials, components, or semi-finished products into designated locations on production equipment or assembly workstations for subsequent processing, assembly, or handling. This process is crucial for improving production efficiency, reducing costs, ensuring product quality, and achieving production automation. Assembly loading can be broadly categorized into manual loading and automated loading.

[0035] Figure 1 is a flowchart illustrating a reinforcement learning-based mobile phone assembly and feeding method according to an embodiment of this disclosure. As shown in Figure 1, the method includes:

[0036] S1 uses a sensing module to collect status data of the robotic arm during the loading task; the status data includes environmental status information and its own motion status.

[0037] For example, sensing modules, such as vision systems and position and status sensors, can be used to collect environmental status information and the robot's own motion status in real time during the loading task. This data includes the robot's coordinate position, speed, type and position of the object being grasped, etc.

[0038] For example, the position sensor and the status sensor are respectively installed at the end of the robotic arm and at key positions in the assembly station; wherein, the position sensor is used to collect the pose of the end effector of the mobile phone equipment loading device, the target pose of the mobile phone assembly loading platform, and the pose of the mobile phone assembly operation platform; the status sensor is used to collect collision detection information and control cabinet status information.

[0039] Specifically, embodiments of this disclosure collect image data, position data, and status data in real time during the mobile phone assembly process. High-precision cameras, position sensors, and status sensors can be used to ensure the accuracy and real-time nature of the data. The high-precision camera is mounted above the assembly station to ensure that the field of view covers the entire assembly area. Position sensors and status sensors are respectively mounted at the end effector of the robotic arm and at key locations on the assembly station to obtain precise component positions and assembly status.

[0040] Furthermore, this disclosure can also preprocess and extract features from the collected data to ensure data quality, and use reinforcement learning algorithms to train the model to generate an optimized feeding strategy model.

[0041] For example, preprocessing may include data cleaning, noise reduction, feature extraction, etc.

[0042] S2, based on the preset reward function, uses a reinforcement learning algorithm to learn from the state data to train the reinforcement learning feeding strategy model, and generates the optimal action strategy based on the trained reinforcement learning feeding strategy model.

[0043] Specifically, the policy model is trained using a reinforcement learning algorithm based on the reward function. After training, the reinforcement learning-based feeding policy model can generate optimal or near-optimal action policies based on the input state information. The reward function is:

[0044] Specifically, embodiments of this disclosure process and analyze collected data based on reinforcement learning algorithms. This includes a data processing unit and a reinforcement learning algorithm. By constructing an intelligent agent and conducting extensive training in a simulated environment, the material feeding path and strategy are optimized. The reinforcement learning algorithm includes deep Q-networks, policy gradients, and proximal policy optimization, among others.

[0045] S3 utilizes the PLC system to parse the optimal motion strategy into instructions, and controls the robotic arm's motor and related actuators to perform control operations according to the parsed instructions.

[0046] Specifically, a PLC system can be used to parse strategy instructions and control the robotic arm's motors and other actuators to perform precise operations according to the strategy. During the robotic arm's task execution, the system continuously monitors its status and environmental changes, collecting new data in real time and feeding it back to the reinforcement learning model. The model continues to adjust and optimize the strategy based on this new data. Through this iterative process, the entire system can achieve more efficient, precise, and intelligent material handling operations.

[0047] Understandably, during the material loading process, the assembly status is monitored in real time, and the latest operational data is collected and fed back to the data acquisition module. This forms a closed-loop system that can adjust and optimize operational strategies based on problems and results encountered in actual operation, achieving self-learning and improvement. Through this process design, the system can not only automatically complete the material loading task but also continuously optimize its performance based on feedback from actual operation, improving work efficiency and accuracy. This design makes the entire system more intelligent and adaptive.

[0048] Figure 2 is a flowchart of the mobile phone assembly and feeding method based on reinforcement learning disclosed in this disclosure. In summary, this disclosure achieves a significant breakthrough in the mobile phone assembly and feeding process by introducing reinforcement learning algorithms and an intelligent control system, demonstrating remarkable effectiveness. First, by optimizing the feeding path and strategy, feeding efficiency is significantly improved. The reinforcement learning algorithm can quickly adapt to different assembly tasks, reducing time consumption during the feeding process and improving overall production efficiency. Simultaneously, real-time data acquisition and dynamic adjustment ensure high precision in the feeding process. Data feedback from high-precision cameras and sensors allows the reinforcement learning algorithm to precisely control the feeding operation, reducing deviations in component positioning and gripping, thereby improving assembly accuracy and product quality. This disclosure enhances the system's flexibility, enabling it to quickly adapt to changes in mobile phone design and processes, reducing the difficulty of production line adjustment and maintenance. The reinforcement learning algorithm possesses adaptive capabilities, allowing it to adjust the feeding strategy in real time according to different production needs, ensuring high system flexibility. This improves the efficiency and precision of the feeding process, achieving intelligent production. This disclosure simplifies the operation process, reduces reliance on professional technicians, and reduces training and maintenance costs. The application of intelligent control systems allows operators to perform material loading and system maintenance more conveniently, improving system reliability and ease of use. It increases loading efficiency and accuracy, reducing production downtime and resource waste caused by operational errors and equipment malfunctions, thereby effectively lowering production costs. The system's efficient operation and intelligent management help enterprises achieve cost reduction and efficiency improvement. This disclosure also enhances product quality. The high-precision loading process ensures consistent and stable assembly quality, reducing product defect rates. By optimizing loading operations, the quality of the final product is improved, enhancing the enterprise's market competitiveness.

[0049] In summary, this invention significantly improves the efficiency, accuracy, and flexibility of the mobile phone assembly and feeding process by introducing reinforcement learning algorithms and intelligent control systems, while reducing operational complexity and maintenance costs. It provides the mobile phone manufacturing industry with an efficient, intelligent, and reliable feeding solution, which has broad application prospects and significant commercial value.

[0050] The reinforcement learning-based mobile phone assembly and feeding method disclosed in this embodiment solves several problems existing in the current mobile phone assembly and feeding process, especially how to improve the accuracy and efficiency of the feeding operation, and increase the flexibility, adaptability and intelligence of the feeding, so as to achieve more efficient, accurate and intelligent feeding operation, improve feeding efficiency, enhance system flexibility, simplify operation process and improve product quality.

[0051] Corresponding to the reinforcement learning-based mobile phone assembly and feeding methods provided in the above embodiments, one embodiment of this disclosure also provides a reinforcement learning-based mobile phone assembly and feeding system. Since the reinforcement learning-based mobile phone assembly and feeding system provided in this disclosure corresponds to the reinforcement learning-based mobile phone assembly and feeding methods provided in the above embodiments, the implementation of the hand-raising recognition method is also applicable to the reinforcement learning-based mobile phone assembly and feeding system provided in this disclosure, and will not be described in detail in the following embodiments.

[0052] Figure 3 is a schematic diagram of the structure of a mobile phone assembly and feeding system based on reinforcement learning according to an embodiment of the present disclosure. As shown in Figure 3, it includes a data acquisition module, a data processing module, a model training module, and a feeding operation module.

[0053] The data acquisition module is used to collect the status data of the robotic arm during the material loading task; wherein, the status data includes environmental status information and its own action status;

[0054] The data processing module is used to preprocess the status data sent by the data acquisition module to obtain preprocessed data;

[0055] The model training module is used to train a reinforcement learning feeding strategy model by learning the state data according to a preset reward function and using a reinforcement learning algorithm, and to generate the optimal action strategy based on the trained reinforcement learning feeding strategy model.

[0056] The loading operation module is used to use the PLC system to parse the optimal motion strategy, so as to control the motor and related actuators of the robotic arm to perform control operations according to the optimal motion strategy based on the parsed instructions.

[0057] In this embodiment of the disclosure, it is also used for:

[0058] Continuously monitor the real-time status data of the robotic arm during task execution;

[0059] The real-time status data is fed back to the feeding strategy model through the data acquisition module, so that the model can adjust and optimize the action strategy based on the real-time status data.

[0060] In this embodiment of the disclosure, the sensing module includes a vision system, a state sensor, and a position sensor; the action state includes at least the coordinate position and speed of the robotic arm, and the type and position of the object being grasped; the reinforcement learning algorithm includes multiple methods such as deep Q-networks, policy gradients, and proximal policy optimization.

[0061] In this embodiment of the disclosure, the position sensor and the state sensor are respectively installed at the end of the robotic arm and at key positions of the assembly station; wherein, the position sensor is used to collect the pose of the end effector of the mobile phone equipment loading device, the target pose of the mobile phone assembly loading platform, and the pose of the mobile phone assembly operation platform; the state sensor is used to collect collision detection information and control cabinet status information.

[0062] In this embodiment of the disclosure, the reward function is:

[0063] The mobile phone assembly and feeding system based on reinforcement learning proposed in this disclosure solves several problems existing in the current mobile phone assembly and feeding process, especially how to improve the accuracy and efficiency of feeding operations, and increase the flexibility, adaptability and intelligence of feeding operations. It can achieve more efficient, more accurate and more intelligent feeding operations, improve feeding efficiency, enhance system flexibility, simplify operation process and improve product quality.

[0064] Furthermore, a schematic diagram of a reinforcement learning-based mobile phone assembly and feeding device according to an embodiment of the present disclosure is provided.

[0065] As shown in Figures 4 and 5, a mobile phone assembly and feeding device based on reinforcement learning is proposed according to the first aspect of this disclosure, including: an end effector 12 and a control system, wherein the control system controls the end effector 12 to perform assembly operations; the end effector 12 includes: a pneumatic parallel gripper 10, a positioning pin 7 and a camera.

[0066] The pneumatic parallel gripper 10 has a left gripper 9 and a right gripper 6 positioned opposite each other at its end. The control system controls the opening and closing of the pneumatic parallel gripper 10. Specifically, the pneumatic parallel gripper 10 is electrically connected to the control system, which is based on a PLC controller. The PLC controller controls the opening and closing of the pneumatic parallel gripper 10 by controlling whether the solenoid valve of the pneumatic parallel gripper 10 is energized. For example, when the solenoid valve is energized, the air path switches, thus opening the pneumatic parallel gripper 10; when the solenoid valve is de-energized, the air path switches, thus closing the pneumatic parallel gripper 10. The left gripper 9 and right gripper 6 are positioned opposite each other at the end of the pneumatic parallel gripper 10. The left gripper 9 and right gripper 6 move in tandem with the opening and closing of the pneumatic parallel gripper 10 to complete assembly actions such as the left gripper 9 and right gripper 6 engaging with or disengaging from the mobile phone assembly carrier 13.

[0067] In this embodiment, the positioning pin 7 is connected to the fixing plates located on both sides of the pneumatic parallel gripper 10 via the carrier contact baffle 5. The carrier contact baffle 5 is located between the left gripper 9 and the right gripper 6 and is vertically arranged on one side of the fixing plate. That is, fixing plates, such as the left fixing plate 4 and the right fixing plate 11, are arranged on both sides of the pneumatic parallel gripper 10. The left fixing plate 4 and the right fixing plate 11 are arranged opposite to each other, and the left gripper 9 and the right gripper 6 are respectively located between the left fixing plate 4 and the right fixing plate 11. In this embodiment, the carrier contact baffle 5 is arranged on one side of the left fixing plate 4 and the right fixing plate 11 and is perpendicular to the left fixing plate 4 and the right fixing plate 11, as shown in Figure 6. The carrier contact baffle 5 is located between the left gripper 9 and the right gripper 6, and the extension direction of the carrier contact baffle 5 is perpendicular to the line connecting the left gripper 9 and the right gripper 6. The positioning pin 7 is arranged on the carrier contact baffle 5 and is used to fit with the positioning hole of the mobile phone assembly carrier 13.

[0068] As shown in Figures 5 and 6, the left fixing plate 4 and the right fixing plate 11 are fixed to both sides of the pneumatic parallel gripper 10 by bolts. The carrier contact baffle 5 is fixed to one end of the left fixing plate 4 and the right fixing plate 11. Two positioning pins 7 are set on the carrier contact baffle 5, facing each other, with one end of each pin set on the carrier contact baffle 5 and perpendicular to it. It should be noted that the position and number of positioning pins 7 are based on the positioning holes of the mobile phone assembly carrier 13 and can be adjusted accordingly.

[0069] In this embodiment, the camera is mounted on the fixed plate to determine the moving position of the positioning pin 7. For example, the camera is an industrial camera 3, which is connected to the left fixed plate 4 via a camera bracket 2. The working principle of the end effector 12 in this embodiment is as follows: the end effector 12 is connected to the robotic arm. Driven by the robotic arm, the end effector 12 moves to the front of the mobile phone assembly loading platform 14, as shown in Figure 8. The industrial camera 3 on the end effector 12 detects the specific position of the mobile phone assembly carrier 13, as shown in Figure 7. The control system moves the end effector 12 so that the positioning pin 7 aligns with the positioning hole of the mobile phone assembly carrier 13. Then, the pneumatic parallel gripper 10 is opened, and both the left gripper 9 and the right gripper 6 are engaged in the mobile phone assembly carrier 13. Finally, the pneumatic parallel gripper 10 is closed, allowing the end effector 12 to grip the mobile phone assembly carrier 13 on the mobile phone assembly loading platform 14, as shown in Figure 10.

[0070] As shown in Figure 11, the end effector 12 moves the mobile phone assembly carrier 13 above the mobile phone assembly operation platform 15, as shown in Figure 9. The positioning hole of the mobile phone assembly carrier 13 is aligned with the positioning pin of the mobile phone assembly operation platform 15. The mobile phone assembly carrier 13 is placed on the mobile phone assembly operation platform 15. Then, the control system controls the pneumatic parallel gripper 10 to open and exit the mobile phone assembly carrier 13. After the pneumatic parallel gripper 10 exits the mobile phone assembly carrier 13, the control system closes the pneumatic parallel gripper 10. Finally, the robotic arm drives the end effector 12 back to its original position.

[0071] In some embodiments, an end connector 1 is also included, which is disposed on the side of the fixed plate away from the carrier contact baffle 5, for connecting an external robotic arm.

[0072] The end effector 12 also includes an end connector 1, which is fixed to the end of the pneumatic parallel gripper 10 away from the left gripper 9 and the right gripper 6, and is connected to the left and right fixed plates 11 by bolts as shown in Figure 4. The end of the end connector 1 away from the pneumatic parallel gripper 10 is connected to the robotic arm, and the end effector 12 can be moved under the drive of the robotic arm.

[0073] In some embodiments, rubber pads 8 are respectively provided on the left gripper 9 and the right gripper 6, located on the inner sides of the left gripper 9 and the right gripper 6 respectively, as shown in Figure 4.

[0074] Rubber pads 8 are provided on the left jaw 9 and the right jaw 6 respectively. The rubber pads 8 are made of hard or soft rubber and are located on the inner side of the left jaw 9 and the right jaw 6 respectively. When the left jaw 9 and the right jaw 6 are engaged in the mobile phone assembly carrier 13 to clamp the mobile phone assembly carrier 13, the setting of the rubber pads 8 can improve the clamping stability and increase the clamping force, making it less likely to loosen or shift. In addition, the rubber pads 8 can also reduce the local deformation and surface damage of the mobile phone assembly carrier 13 and improve the processing safety.

[0075] The reinforcement learning-based mobile phone assembly and feeding device according to the embodiments of this disclosure solves several problems existing in the current mobile phone assembly and feeding process, especially how to improve the accuracy and efficiency of the feeding operation, and increase the flexibility, adaptability and intelligence of the feeding, so as to achieve more efficient, more accurate and more intelligent feeding operation, improve feeding efficiency, enhance system flexibility, simplify operation process and improve product quality.

[0076] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A mobile phone assembly and feeding method based on reinforcement learning, characterized in that, include: The robotic arm uses a sensing module to collect status data during the loading process; wherein, the status data includes environmental status information and its own motion status; The state data is trained using a preset reward function and a reinforcement learning algorithm to train a reinforcement learning feeding strategy model, and an optimal action strategy is generated based on the trained reinforcement learning feeding strategy model. The optimal motion strategy is analyzed using a PLC system, and the motors and related actuators of the robotic arm are controlled to perform control operations according to the analyzed instructions.

2. The method according to claim 1, characterized in that, The method further includes: Continuously monitor the real-time status information of the robotic arm during task execution; The real-time status information is fed back to the material feeding strategy model so that the model can adjust and optimize the action strategy based on the real-time status information.

3. The method according to claim 1, characterized in that, The sensing module includes a vision system, a state sensor, and a position sensor; the action state includes at least the coordinate position and speed of the robotic arm, and the type and position of the object being grasped; the reinforcement learning algorithm includes multiple methods such as deep Q-network, policy gradient, and proximal policy optimization.

4. The method according to claim 3, characterized in that, The position sensor and the status sensor are respectively installed at the end of the robotic arm and at key positions in the assembly station; wherein, the position sensor is used to collect the position and orientation of the end effector of the mobile phone equipment loading device, the target position and orientation of the mobile phone assembly loading platform, and the position and orientation of the mobile phone assembly operation platform; the status sensor is used to collect collision detection information and control cabinet status information.

5. The method according to claim 1, characterized in that, The reward function:

6. A mobile phone assembly and feeding system based on reinforcement learning, characterized in that, include: The data acquisition module is used to collect the status data of the robotic arm during the material loading task; wherein, the status data includes environmental status information and its own action status; The data processing module is used to preprocess the status data sent by the data acquisition module to obtain preprocessed data; The model training module is used to train a reinforcement learning feeding strategy model by learning the state data according to a preset reward function and using a reinforcement learning algorithm, and to generate the optimal action strategy based on the trained reinforcement learning feeding strategy model. The loading operation module is used to use the PLC system to parse the optimal motion strategy, so as to control the motor and related actuators of the robotic arm to perform control operations according to the parsed instructions.

7. A mobile phone assembly and feeding device based on reinforcement learning, wherein, include: An end effector and a control system, wherein the control system controls the end effector to perform assembly operations; The end effector includes: A pneumatic parallel gripper has a left gripper and a right gripper positioned opposite each other at its ends; the control system controls the pneumatic parallel gripper to open or close. A positioning pin, which connects to a fixed plate located on both sides of the pneumatic parallel gripper after being transferred by a carrier contact baffle; the carrier contact baffle is located between the left and right grippers and is perpendicularly disposed on one side of the fixed plate; and A camera, which is mounted on the fixed plate, is used to determine the movement position of the positioning pin.

8. The mobile phone assembly and feeding device according to claim 7, wherein, It also includes an end connector disposed on the side of the fixed plate away from the vehicle contact baffle for use with an external robotic arm.

9. The mobile phone assembly and feeding device according to claim 7 or 8, wherein, Rubber pads are provided on the left and right grippers respectively, located on the inner sides of the left and right grippers.

10. The mobile phone assembly and feeding device according to claim 9, wherein, The rubber pad is either hard rubber or soft rubber.

11. The mobile phone assembly and feeding device according to claim 7, wherein, The positioning pins are two in number and arranged opposite to each other. The positioning pins are arranged perpendicularly to the vehicle contact baffle, with one end of each pin being disposed on the vehicle contact baffle.

Citation Information

Patent Citations

  • Mechanical arm control method based on guided DQN control

    CN111152227A

  • Mechanical arm autonomous grabbing method based on deep reinforcement learning and dynamic movement primitives

    CN111618847A

  • Multi-mechanical-arm collaborative assembly method and system based on deep reinforcement learning

    CN111881772A

  • Robot micro-assembly grabbing system based on deep reinforcement learning

    CN112847368A

  • Mechanical arm control method and system based on deep reinforcement learning

    CN114012735A