Multi-modal assembly process recognition system and method based on deep learning
Through a multi-modal assembly process recognition system based on deep learning, real-time video acquisition and motion recognition are carried out using cameras and computer equipment to automatically compare assembly actions, thus solving the problem of easy errors in manual assembly, realizing full-time inspection of the parts assembly process, and reducing product defects and rejection rates.
Patent Information
- Application Number
- CN202211295547.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-10-21
AI Technical Summary
In the product production and assembly process, the manual assembly process is prone to operational errors due to fatigue, resulting in low accuracy of independent inspection and high product rejection rate.
A multi-modal assembly process recognition system based on deep learning is adopted. It uses cameras, computer equipment, position sensors, limit switches and release switches to identify assembly actions through real-time video acquisition and deep learning technology, and automatically compares preset action combinations to ensure the accuracy of the assembly process.
It realizes full-time inspection of the parts assembly process, reduces the probability of product defects and unqualified situations, and improves the accuracy and efficiency of the assembly process.
Smart Images

Figure CN115761876B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of information technology and artificial intelligence technology, and in particular to a multi-modal assembly process recognition method based on deep learning. Background Art
[0002] Manual assembly of parts is a common method of production and assembly. Because manual assembly of parts on a production line is a highly repetitive process, it is often necessary to inspect whether the assembly process has been completed.
[0003] In the related art, workers usually conduct self-inspection to determine whether there is a problem with the assembly process of the parts based on the assembly results of the parts. If there is a problem with the assembly process of the parts, the finished parts are traced and reassembled.
[0004] However, due to the repetitive operation steps involved in the manual assembly process, workers are prone to operating errors due to fatigue, resulting in low accuracy of self-inspection, which in turn leads to product defects and a high rejection rate. Summary of the Invention
[0005] This application is about a multi-modal assembly process identification method and system based on deep learning, which can provide full-time inspection of the parts assembly process and reduce the probability of product defects and non-conformities. The technical solution is as follows:
[0006] On the one hand, a multimodal assembly process recognition system based on deep learning is provided, which includes a camera, a computer device, a position sensor, a limit switch, and a release switch;
[0007] The camera, the position sensor, the limit switch, and the release switch are respectively connected to the computer device for communication;
[0008] The position sensor is configured to send a detection signal to the computer device in response to the part to be assembled being located at the assembly position;
[0009] The limit switch is used to limit the movement of the production line where the workpiece to be assembled is located in response to the workpiece to be assembled being located at the assembly position;
[0010] The computer device is configured to receive the signal to be detected; and send a video acquisition signal to the camera based on the signal to be detected;
[0011] The camera is configured to receive the video acquisition signal; perform video acquisition of the position of the parts to be assembled based on the video acquisition signal to obtain an assembly video; and send the assembly video to the computer device in real time;
[0012] The computer device is configured to receive the assembly video; perform action recognition on the assembly video based on deep learning technology to obtain an assembly action combination, wherein the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions, and the assembly action combination is used to process the to-be-assembled parts into assembled parts; compare the assembly action combination with a preset assembly action combination; and in response to the assembly action combination being consistent with the preset assembly action combination, send a release signal to the release switch;
[0013] The release switch is used to receive the release signal; and start the production line where the assembled part is located based on the release signal.
[0014] In another aspect, a multi-modal assembly process recognition method based on deep learning is provided. The method is applied to a computer device within the multi-modal assembly process recognition system based on deep learning as described above. The method comprises:
[0015] receiving a signal to be detected, where the signal to be detected is a signal generated by the position sensor;
[0016] Sending a video acquisition signal to the camera based on the signal to be detected;
[0017] Receiving the assembly video sent by the camera;
[0018] performing action recognition on the assembly video based on deep learning technology to obtain an assembly action combination, wherein the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions, and the assembly combination is used to process the to-be-assembled parts into assembled parts;
[0019] Comparing the assembly action combination with a preset assembly action combination;
[0020] In response to the assembly action combination being consistent with the preset assembly action combination, a release signal is sent to a release switch.
[0021] The beneficial effects brought about by the technical solution of this application include at least:
[0022] During the parts assembly process, cameras record the location of the parts to be assembled and the processing of the parts in real time, and send the video content to the computer equipment. Based on deep learning artificial intelligence technology, the process content is identified to obtain a set of assembly actions for the parts to be assembled. The assembly actions are compared with the preset actions to determine whether the installation process is correct. When the process is correct, the computer equipment will indicate that the pass switch is working, contacting the restrictions on the part position and production line status, and indicating that the part assembly work is completed. During the parts assembly process, with the help of the camera's full-time motion monitoring and the computer equipment's intelligent recognition, the parts assembly process will not have process errors, and the parts assembly process can be inspected at all times, reducing the probability of product defects and non-conformities. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A structural schematic diagram of a multi-modal assembly process recognition system based on deep learning provided by an exemplary embodiment of the present application is shown.
[0025] Figure 2 A schematic diagram of a framework of another multi-modal assembly process recognition process based on deep learning provided by an exemplary embodiment of the present application is shown.
[0026] Figure 3 A structural schematic diagram of a multi-modal assembly process recognition system based on deep learning provided by an exemplary embodiment of the present application is shown.
[0027] Figure 4 A flowchart of a multi-modal assembly process recognition method based on deep learning provided by an exemplary embodiment of the present application is shown.
[0028] Figure 5 A flowchart of another multi-modal assembly process identification method based on deep learning provided by an exemplary embodiment of the present application is shown.
[0029] Figure 6 A schematic diagram showing content displayed by a display device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0031] Figure 1 A schematic diagram of a multi-modal assembly process recognition system based on deep learning is shown in FIG. Figure 1 The system includes a camera 110, a computer device 120, a position sensor 130, a limit switch 140, and a release switch 150. The camera, the position sensor, the limit switch, and the release switch are respectively connected to the computer device for communication.
[0032] In the embodiments of the present application, position sensors, limit switches, and release switches are all implemented as components on the parts assembly line. The position sensors and limit switches work together to limit the position of the parts to be assembled when they are moved to the assembly position, and provide signal feedback, allowing the computer equipment to be informed of the start of the assembly process and to perform the assembly process identification process. After the assembly process identification is completed and the assembly process is correct, the release switch is activated, allowing the assembled parts to be sent to subsequent parts, such as parts defect detection or storage.
[0033] In the embodiments of the present application, a camera and a computer device work together to complete the assembly process identification process. The camera is implemented as a device with image acquisition and communication functions. It can perform video capture based on communication signals and transmit the assembly video in real time through a communication connection with the computer device.
[0034] In the embodiment of the present application, the computer device is implemented as a device with data processing, receiving, sending and processing functions. Optionally, the computer device is implemented as a personal computer, or an edge computing device based on artificial intelligence. Optionally, the computer device has a framework dedicated to video processing and production, please refer to Figure 2 , which integrates five functions, namely, stream pulling 201, pre-processing 202, target detection 203, post-processing 204, and storage and display 205, to perform the assembly process identification process required by this application. It should be noted that in the embodiment of this application, the storage function of the computer device can be a local storage function or a local storage function.
[0035] Therefore, in the above process, the position sensor is used to send a signal to be detected to the computer device in response to the part to be assembled being located at the assembly position;
[0036] A limit switch, configured to limit the movement of the production line where the workpiece to be assembled is located in response to the workpiece to be assembled being located at the assembly position;
[0037] A computer device is used to receive a signal to be detected; and send a video acquisition signal to a camera based on the signal to be detected;
[0038] The camera is used to receive a video acquisition signal; based on the video acquisition signal, the position of the assembly part is captured to obtain an assembly video; and the assembly video is sent to a computer device in real time;
[0039] A computer device is configured to receive an assembly video; perform action recognition on the assembly video based on deep learning technology to obtain an assembly action combination, wherein the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions, and the assembly action combination is used to process the to-be-assembled parts into assembled parts; compare the assembly action combination with a preset assembly action combination; and send a release signal to a release switch in response to the assembly action combination being consistent with the preset assembly action combination;
[0040] The release switch is used to receive a release signal and start the production line where the assembled parts are located based on the release signal.
[0041] In summary, the system provided by the embodiment of the present application, during the process of parts assembly, uses a camera to record the position of the parts to be assembled and the processing process of the parts in real time, and sends the video content to a computer device. Based on deep learning artificial intelligence technology, the process content is identified to obtain a set of assembly actions for the parts to be assembled, and the assembly actions are compared with the preset actions to determine whether the installation of the process is correct. When the process is correct, the computer device indicates that the pass switch is working, contacts the restrictions on the part position and the production line status, and indicates that the assembly of the parts is completed. During the process of parts assembly, with the help of the camera's full-process action monitoring and the computer device's intelligent recognition, there will be no process errors in the parts assembly process, and full-time inspection is provided for the parts assembly process, reducing the probability of product defects and unqualified conditions.
[0042] Combined with the actual needs of the production line, Figure 3 A schematic diagram of a multi-modal assembly process recognition system based on deep learning is shown in FIG. Figure 3 The system includes a camera 310, a computer device 320, a position sensor 330, a limit switch 340, a release switch 350 and a display 360.
[0043] The display device is connected to the computer device for communication and displays the assembly identification related content in real time according to the needs of the staff and the assembly process. This application will explain the interaction process between the display device and the computer device in subsequent embodiments.
[0044] Figure 4A flowchart of a multi-modal assembly process identification method based on deep learning provided by an exemplary embodiment of the present application is shown. This method is described by taking the application of the method to a computer device within a multi-modal assembly process identification system based on deep learning as shown in any of the above embodiments as an example. The method includes:
[0045] Step 401: Receive a signal to be detected.
[0046] In the embodiment of the present application, the signal to be detected is a signal generated by a position sensor.
[0047] As a preliminary step in this process, the part to be assembled is transported to the computer equipment on the production line. When it moves to the position corresponding to the position sensor, the position sensor responds to the part's position and sends a detection signal to the computer equipment. The limit switch is activated, restricting the part's movement on the production line. Optionally, after the part stops moving, the worker can process the part at the workstation.
[0048] Step 402: Send a video acquisition signal to the camera based on the signal to be detected.
[0049] This process involves the computer controlling the camera to capture video signals after determining that the part assembly process has begun. After the video capture signal is sent, the camera receives the video capture signal and, based on the video capture signal, captures the assembly position using a preset shooting angle and shooting mode.
[0050] Step 403: Receive the assembly video sent by the camera.
[0051] In the embodiment of the present application, the camera sends the assembly video in real time, so the computer device synchronously receives the assembly video in real time.
[0052] Step 404: Perform action recognition on the assembly video based on deep learning technology to obtain an assembly action combination.
[0053] In an embodiment of the present application, a computer device processes video content based on deep learning technology to identify the assembly process of workers in real time.
[0054] Optionally, the computer device includes an action library that stores the contact states and positions of each assembly action corresponding to each assembly operation. The contact states indicate how the worker contacts the part to be assembled, and the positions indicate the relative position of the worker's hand and the part to be assembled during the corresponding assembly process. By recording samples and learning from these samples based on features, the computer device can identify the worker's assembly actions.
[0055] It should be noted that the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions. The assembly combination is used to process the to-be-assembled parts into assembled parts.
[0056] Step 405 : Compare the assembly action combination with the preset assembly action combination.
[0057] In an embodiment of the present application, a computer device stores preset assembly actions corresponding to the parts to be assembled. In one example, a named component is the Mprop component, which has four corresponding assembly actions: "Inject Oil," "Press," "Tighten Screws," and "Install Green Accessories." The preset assembly action combination records the four actions and the execution sequence relationship between them.
[0058] Step 406: In response to the assembly action combination being consistent with the preset assembly action combination, a release signal is sent to the release switch.
[0059] In the embodiment of the present application, the release signal indicates that an error has occurred in the assembly process, and the assembled parts can enter the subsequent processing process, such as the transit warehousing process, or the defect detection process.
[0060] In summary, the method provided by the embodiment of the present application, during the process of parts assembly, uses a camera to record the position of the parts to be assembled and the processing process of the parts in real time, and sends the video content to a computer device. Based on deep learning artificial intelligence technology, the process content is identified to obtain a set of assembly actions for the parts to be assembled, and the assembly actions are compared with the preset actions to determine whether the installation of the process is correct. When the process is correct, the computer device indicates that the pass switch is working, contacts the restrictions on the part position and production line status, and indicates that the assembly of the parts is completed. During the process of parts assembly, with the help of the camera's full-process action monitoring and the computer device's intelligent recognition, there will be no process errors in the parts assembly process, and full-time inspection is provided for the parts assembly process, reducing the probability of product defects and unqualified conditions.
[0061] Next, the interaction process between the computer device and the display device is described in conjunction with the situation where a display device exists in the system.
[0062] Figure 5 A schematic diagram of a multi-modal assembly process identification method based on deep learning is shown in an exemplary embodiment of the present application. Figure 3 Taking the computer device and display device in the multi-modal assembly process recognition system based on deep learning as an example, the method includes:
[0063] Step 501: A computer device receives an assembly video.
[0064] This process is the process of the computer device obtaining the assembly video from the camera, which will not be described in detail here.
[0065] Step 502: The computer device pre-processes the assembly video.
[0066] This process is a preprocessing process for the video. After preprocessing, the computer equipment can call the model to perform the process of process recognition for the assembly video.
[0067] In step 503 , in response to the pre-processed assembly video, the computer device performs real-time action recognition on the pre-processed assembly video using an action recognition model to obtain an assembly action combination.
[0068] In the embodiment of the present application, the action recognition model is a Yolov5 model, and the action recognition model corresponds to an action recognition model training sample set. The action recognition model set includes at least two sample actions annotated with sample action recognition results, and the sample actions are annotated with contact state features and position features. It should be noted that the action recognition model has sample learning and deep learning capabilities. In this case, the action recognition model includes an image input layer, a feature extraction layer, and a loss function output layer to adapt to the recognition requirements of the assembly process.
[0069] Step 504 : The computer device sends a display instruction to the display device in response to receiving the assembly video.
[0070] Optionally, the display instruction includes an assembly video, that is, the computer device directly sends information requiring video display to the display device.
[0071] It should be noted that, in some embodiments of the present application, the identification and display of the assembly process are simultaneous processes, that is, while the video signal is being acquired, the assembly action is identified in the manner described in steps 502 to 503 .
[0072] Step 505: The display device receives a display instruction.
[0073] Step 506: The display device displays the assembly video based on WEB technology.
[0074] In the embodiment of the present application, the display device uses World Wide Web (WEB) technology to display the assembly video.
[0075] In the embodiments of the present application, the display device can be implemented as a dedicated display terminal, such as a display screen, or a mobile or fixed terminal with integrated display function, such as a mobile phone, a tablet computer, and a personal computer. The present application does not limit the specific implementation form of the display device.
[0076] It should be noted that in some embodiments, the display device needs to display the execution status of the assembly action, that is, the computer device's recognition result of the assembly process in real time. In this case, the display instruction also includes assembly action combination data. After receiving the display instruction, the display device will superimpose and display at least two preset assembly action identifiers corresponding to the preset assembly action combination on the assembly video according to the assembly sequence and assembly content in the assembly action combination. Please refer to Figure 6 A preset assembly action combination area 620 is superimposed on the assembly video display area 610. This area includes at least two preset assembly action identifiers 621. In this embodiment, the preset assembly action identifiers 621 are implemented as text identifiers to visually display the contents of the preset assembly action. Optionally, the display interface also includes a current process display area 630 for visually displaying the currently executing assembly process.
[0077] In step 507 , in response to the assembly action being consistent with the preset assembly action, the computer device sends an action consistency display signal to the display device.
[0078] Step 508: The display device receives the action consistent display signal.
[0079] Step 509 : The display device displays the preset assembly action identifier in a first highlight mode based on the action consistent display signal.
[0080] In step 510 , in response to the assembly action being inconsistent with the preset assembly action, the computer device sends an action inconsistent display signal to the display device.
[0081] It should be noted that there is a possibility that steps 510 to 513 are not executed, and steps 510 to 513 are executed only when the assembly action recognition result indicates that the assembly process is missing or erroneous.
[0082] Step 511: The display device receives an action inconsistency display signal.
[0083] Step 512: The display device displays the preset assembly action identifier in a second highlight mode based on the action inconsistency display signal.
[0084] In an embodiment of the present application, as a supplement, the display device displays the preset assembly action identifier being detected, that is, the preset assembly action identifier corresponding to the current process, in a third highlight mode; and displays the undetected preset assembly action identifier in a normal mode.
[0085] It should be noted that in the embodiments of the present application, the first highlight mode, the second highlight mode, and the third highlight mode are all font or text box display modes used to clearly prompt the user. In one example, the first display mode changes the color of the text content in the preset assembly action identifier to green, the second display mode displays the text content in the preset assembly action identifier in red, and the third display mode displays the text content in the preset assembly action identifier in yellow.
[0086] In an optional embodiment, the computer device performs active tool recognition on the assembly video to obtain an active tool recognition result; and sends the active tool recognition result to a display device;
[0087] The display device receives the active tool identification result and displays the active tool in a box selection based on the active tool identification result. For example, if the computer device identifies the active tool as a screwdriver, the display interface will prompt the viewer with a box selection that the active tool in the current process is a screwdriver.
[0088] In step 513 , the display device displays the preset assembly action identifier in a second highlight mode to provide an alarm validation prompt.
[0089] In the embodiment of the present application, the computer device will send a release signal only when the assembly action is completely consistent with the preset assembly action. In other cases, the display device will issue an alarm and display content in the interface to remind the staff of the incorrect assembly process.
[0090] In step 514 , the computer device sends a release signal to the release switch in response to the assembly action combination being consistent with the preset assembly action combination.
[0091] When the computer equipment detects that the entire assembly process is correct, it releases the assembled parts and allows them to enter the subsequent assembly process.
[0092] In summary, the method provided by the embodiment of the present application, during the process of parts assembly, uses a camera to record the position of the parts to be assembled and the processing process of the parts in real time, and sends the video content to a computer device. Based on deep learning artificial intelligence technology, the process content is identified to obtain a set of assembly actions for the parts to be assembled, and the assembly actions are compared with the preset actions to determine whether the installation of the process is correct. When the process is correct, the computer device indicates that the pass switch is working, contacts the restrictions on the part position and production line status, and indicates that the assembly of the parts is completed. During the process of parts assembly, with the help of the camera's full-process action monitoring and the computer device's intelligent recognition, there will be no process errors in the parts assembly process, and full-time inspection is provided for the parts assembly process, reducing the probability of product defects and unqualified conditions.
[0093] The method provided in the embodiments of this application uses a display device to visualize the assembly recognition process of computer equipment, while simultaneously playing the assembly screen and detection content in real time. In actual application scenarios, the entire assembly recognition process of two parts is monitored, further reducing the probability of product defects and non-conformance.
[0094] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A multi-modal assembly process recognition system based on deep learning, characterized in that: The system includes a camera, a computer device, a position sensor, a limit switch, and a release switch; The camera, the position sensor, the limit switch, and the release switch are respectively connected to the computer device for communication; The position sensor is configured to send a detection signal to the computer device in response to the part to be assembled being located at the assembly position; The limit switch is used to limit the movement of the production line where the part to be assembled is located in response to the part to be assembled being located at the assembly position; The computer device is configured to receive the signal to be detected; Sending a video acquisition signal to the camera based on the signal to be detected; The camera is used to receive the video acquisition signal; Performing video capture of the position of the parts to be assembled based on the video capture signal to obtain an assembly video; sending the assembly video to the computer device in real time; The computer device is used to receive the assembly video; performing action recognition on the assembly video based on deep learning technology to obtain an assembly action combination, wherein the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions, and the assembly action combination is used to process the to-be-assembled parts into assembled parts; comparing the assembly action combination with a preset assembly action combination; and sending a release signal to the release switch in response to the assembly action combination being consistent with the preset assembly action combination; The release switch is configured to receive the release signal; and start the production line where the assembled part is located based on the release signal; The system further includes a display device; The computer device is further configured to, while the video acquisition signal is being transmitted, sequentially compare the assembly actions in the assembly action combination with the preset assembly actions in the preset assembly action combination; and in response to the assembly action being consistent with the preset assembly action, send an action consistency display signal to the display device; In response to the assembling action being inconsistent with the preset assembling action, sending an action inconsistent display signal to the display device; The display device is used to receive an assembly action consistent display signal or an assembly action inconsistent display signal; Based on the assembly action consistent display signal, the first highlight mode is used to Assembly action identification Based on the assembly action inconsistency display signal, the preset assembly action identifier is displayed in a second highlight mode.
2. The system according to claim 1, wherein: The computer device is also used to preprocess the assembly video; in response to the assembly video being preprocessed, real-time action recognition is performed on the preprocessed assembly video through an action recognition model to obtain the assembly action combination, the action recognition model is a Yolov5 model, and the action recognition model corresponds to an action recognition model training sample set, and the action recognition model set includes at least two sample actions marked with sample action recognition results, and the sample actions are marked with contact state features and position features.
3. The system according to claim 1, wherein: The action recognition model includes an input image processing layer, a feature extraction layer, a loss function and an output layer. The feature extraction layer is used to extract contact state features and position features.
4. The system according to claim 1, wherein: The computer device is further configured to, in response to receiving the assembly video, send a display instruction to the display device, wherein the display instruction includes the assembly video; The display device is used to receive the display instruction; and display the assembly video based on the World Wide Web (WEB) technology according to the display instruction.
5. The system according to claim 4, characterized in that The display instruction also includes preset assembly action combination data corresponding to the preset assembly action combination; The display device is further configured to display at least two preset assembly action identifiers corresponding to the preset assembly action combination according to the preset assembly action combination data.
6. The system according to claim 1, wherein: The display device is further configured to display undetected preset assembly action identifiers in a normal display mode and the preset assembly action identifiers being detected in a third highlight mode.
7. The system according to claim 1, wherein: The computer device is further configured to perform active tool recognition on the assembly video to obtain an active tool recognition result; and send the active tool recognition result to the display device; The display device is configured to receive the active tool identification result; The active tool is displayed in a frame selection form based on the active tool identification result.
8. The system according to claim 4, wherein: The display device is further configured with a sound module; The display device is further configured to display the preset assembly action identifier in a second highlight mode to provide an alarm effectiveness prompt.
9. A multi-modal assembly process recognition method based on deep learning, characterized in that: The method is applied to a computer device in a multi-modal assembly process recognition system based on deep learning according to any one of claims 1 to 8, and the method comprises: receiving a signal to be detected, where the signal to be detected is a signal generated by the position sensor; Sending a video acquisition signal to the camera based on the signal to be detected; Receiving the assembly video sent by the camera; performing action recognition on the assembly video based on deep learning technology to obtain an assembly action combination, wherein the assembly action combination includes at least two assembly actions and an assembly action sequence corresponding to the assembly actions, and the assembly action combination is used to process the to-be-assembled parts into assembled parts; Comparing the assembly action combination with a preset assembly action combination; In response to the assembly action combination being consistent with the preset assembly action combination, a release signal is sent to a release switch.
Citation Information
Patent Citations
Working procedure detection device based on deep learning, and working procedure detection method thereof
CN108491759A
Assembly monitoring method and device based on deep learning and readable storage medium
CN109816049A