Method and system for monitoring an assembly process of at least one component to be assembled into a device
By using artificial intelligence evaluation circuits based on artificial convolutional networks and mobile vision sensors, the accuracy and efficiency issues in the assembly process of electrical connectors in existing technologies have been solved. This enables efficient and low-cost monitoring of the assembly process, ensuring correct equipment assembly and improving production efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FORD GLOBAL TECH LLC
- Filing Date
- 2025-12-26
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies for monitoring the assembly of electrical connectors or other parts suffer from insufficient accuracy, require complex pattern detection, have high requirements for camera focusing, and cannot detect incomplete insertion configurations, leading to incorrect equipment configurations and affecting manufacturing efficiency.
An artificial intelligence evaluation circuit based on artificial convolutional networks is adopted. A video stream is recorded through a movable visual sensor, the installation configuration of the component is analyzed, and a notification is output in real time to ensure that the component meets the specified installation configuration.
It improves the accuracy and efficiency of the assembly process, reduces computational costs, decreases false positive events, ensures correct equipment assembly, reduces post-manufacturing analysis and maintenance workload, and improves overall production efficiency.
Smart Images

Figure CN122368876A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to a method and system for monitoring the assembly process of at least one component to be assembled into a device. Background Technology
[0002] The number of electrical connectors is steadily increasing in the design and manufacture of new equipment, such as new vehicles. Existing vision systems lack sufficient precision or require complex patterns to detect the correct assembly configuration of such electrical connectors or other parts (such as bolts) to be assembled within the corresponding equipment. In particular, existing vision-based approaches cannot detect configurations where the corresponding connector or part is in principle located at or inside the intended receptacle but is not fully inserted (e.g., not properly secured (also known as "not fully in place")). Another drawback of existing vision systems is the requirement for mandatory, precise camera focusing. This is why such vision systems are prone to error when the relative position between the camera and the relevant parts is variable. Furthermore, current systems are also unusable if the item to be inspected is externally invisible / hidden and the vision sensor is fixedly mounted in the external space of the equipment (while the relevant parts will be assembled inside the equipment). In particular, electrical connectors or other parts (such as bolts) are often located in positions that are difficult to access and monitor from the external space.
[0003] However, improperly installed connectors or parts of this type can lead to misconfiguration of the equipment on which they are installed. Therefore, even though the equipment includes improperly installed parts, the corresponding equipment (such as a vehicle) may be released from the manufacturing facility. This results in additional control operations and maintenance tasks downstream of the actual manufacturing process, leading to reduced manufacturing efficiency.
[0004] US 2024 / 192145 A1 discloses a system including a station information system, a portable vision system, and a quality monitoring system. The station information system includes a station computing device configured to provide notifications related to manufacturing operations performed on a part. The portable vision system includes a quality inspection module configured to include a station task module configured to perform quality inspection tasks based on images. The quality monitoring system includes a quality monitoring computing device configured to request the portable vision system to perform quality inspection tasks based on trigger messages from the station information system and to provide the station information system with task data messages related to the quality inspection tasks performed by the portable vision system. The station computing device is configured to provide notifications via a user interface device based on the task data messages from the quality monitoring system.
[0005] US 2022 / 0066435 A1 discloses an apparatus for determining errors occurring during the execution of an industrial process based on the degree of particle scattering present in a video stream captured by a camera.
[0006] KR 2022 0076792 A discloses an apparatus for an operator to identify abnormally tight connectors in an automotive production line in real time based on real-time determination, which utilizes an audio detection mechanism to record the tightening sound of the connectors.
[0007] KR 10-1665644 B1 discloses a vehicle wiring harness connector terminal inspection system and method that determines the terminal status of one or more connectors based on the analysis of acquired terminal images of one or more connectors obtained from an acquisition unit. The images are acquired by a camera attached to a robotic structure (such as a robotic arm).
[0008] US 2022 / 0382262 A1 discloses an apparatus for detecting multiple defects on a vehicle surface based on the analysis of multiple captured images associated with the vehicle obtained by an image capture device as the vehicle passes through a vehicle assembly line. The camera is fixedly mounted to the assembly line.
[0009] KR 10-2112809 B1 discloses an apparatus for determining whether a vehicle connector component is correctly seated in a coupling unit based on an artificial neural network and further based on one or more images. The artificial neural network is configured to analyze the state information of the vehicle's novel connector components for defects. The one or more images are used to analyze the housing of the vehicle connector and determine whether multiple connectors are defective. Images of the connector are acquired during the connector manufacturing process and subsequently used to assess whether the connector is properly assembled.
[0010] US 2022 / 0136872 A1 discloses an inspection apparatus that includes a portable vision data system, a user interface system, a wireless communication system, and a controller. The controller may include an AI framework for performing inspections, such as visual inspections of parts being assembled.
[0011] CN 215897828U discloses a camera with edge computing capabilities. The device is configured to include a touchscreen, a camera module, and an integrated motherboard with an AI framework to receive images or videos and transmit them for identification processing via a neural network model. The identification results are sent to the touchscreen.
[0012] Therefore, while some of these existing technological approaches utilize fixed-position vision sensors, others employ portable vision sensors, and still others even utilize AI-based architectures, the operating costs for these evaluation architectures are high. In other words, the operational efficiency of these evaluation architectures is low. Therefore, there is a need for a more efficient method and system than existing technological approaches to monitor the assembly process of at least one component to be assembled into a device. Summary of the Invention
[0013] The subject matter of the independent claims satisfies the corresponding needs. Further embodiments are indicated in the dependent claims and the following description, wherein each embodiment may represent aspects of this disclosure independently or in combination.
[0014] The following provides an overview of certain embodiments disclosed herein. It should be understood that these aspects are presented merely to provide a brief overview of these embodiments and are not intended to limit the scope of this disclosure. This disclosure may cover various aspects that may not be set forth below. Some aspects are explained from a methodological perspective, while others are explained from a systemological perspective. However, corresponding aspects will be transferred from a methodological perspective to a systemological perspective, and vice versa.
[0015] According to one aspect, embodiments of this disclosure relate to a method for monitoring the assembly process of at least one component to be assembled into a device. The method includes at least the following steps: - Use at least one vision sensor to record at least one video stream of at least one component to be assembled.
[0016] - For a specified mounting configuration of at least one component to be assembled, the evaluation circuit utilizes artificial intelligence based on at least one artificial convolutional network to analyze at least one recorded video stream.
[0017] - Based on the analysis performed by the evaluation circuit, a notification is output via a human-machine interface. This notification depends at least on the actual installation configuration of at least one component to be assembled and the specified installation configuration of at least one component to be assembled.
[0018] This method is based on the finding that artificial convolutional networks can be used to reduce the computational cost required to determine whether the actual installation configuration of a component to be assembled conforms to a specified installation configuration. The key element is the convolutional network, which is capable of extracting predetermined features from at least one video stream, thus processing the video stream and determining whether the specified installation configuration is met more efficiently compared to existing approaches utilizing other evaluation architectures or common artificial neural networks. Therefore, the assembly process can be evaluated with improved accuracy while reducing computational costs. Consequently, the rate of false positives can be reduced compared to existing approaches—false positives being components that are actually correctly installed but are marked as incorrectly installed, or vice versa. This increases the rate of evaluated equipment for proper assembly, thereby reducing or even minimizing the rate of defective equipment leaving the manufacturing site. Therefore, the workload of post-manufacturing analysis and post-manufacturing repair can be reduced. Therefore, the overall production efficiency of the equipment in which components are assembled is improved compared to existing approaches. Advantageously, appropriate notifications are output via a human-machine interface (HMI) to inform the user or operator accordingly.
[0019] Furthermore, the method does not depend on specific parameters of the recorded video stream. For example, evaluation circuits utilizing artificial intelligence based on artificial convolutional networks can even be designed to evaluate incompletely recorded video data, such as video data where the distance and focus between the vision sensor and the component to be assembled are inconsistent. Moreover, the method does not depend on a fixedly mounted vision sensor. Therefore, vision sensors capable of recording vehicle flow can be employed, even if the components and parts to be assembled are not visible to the external space of the device in which they are assembled.
[0020] According to another aspect, embodiments of this disclosure relate to a system for monitoring the assembly process of at least one component to be assembled into a device. The system includes: at least one visual sensor, evaluation circuitry utilizing artificial intelligence based on at least one artificial convolutional network, and a human-machine interface. The at least one visual sensor is configured to record at least one video stream of the at least one component to be assembled. The evaluation circuitry utilizing artificial intelligence based on at least one artificial convolutional network is configured to analyze the recorded at least one video stream for a specified installation configuration of the at least one component to be assembled. The system is configured to output a notification via the human-machine interface based on the evaluation circuitry. This notification depends at least on the actual installation configuration of the at least one component to be assembled and the specified installation configuration of the at least one component to be assembled.
[0021] The advantages achieved by the method described in this article are also implemented by this system in a corresponding manner.
[0022] Optionally, the analysis of at least one recorded video stream performed by the evaluation circuitry is performed in real time and / or at least partially simultaneously with the recording of at least one video stream performed by at least one visual sensor. In other words, the steps of recording at least one video stream and analyzing the recorded at least one video stream can be performed in parallel with each other and at least partially overlap in time. Therefore, notification can be determined in real time and output via a human-machine interface based on the analysis. Thus, the user will immediately receive notification of the corresponding results of the analysis.
[0023] Preferably, the step of outputting the notification can also be performed in a timely manner, at least partially overlapping with the recording of at least one video stream. In other words, the output of the notification can also be performed in real time.
[0024] In an alternative approach, the analysis of at least one recorded video stream performed by the evaluation circuitry is performed only after the recording of at least one video stream by at least one vision sensor. In this case, the video stream is recorded first, and then the analysis is performed. Therefore, a notification is output only after the video stream has been recorded. For example, this approach might be useful for training scenarios where video streams recorded earlier than the desired time point are analyzed at the desired time point, and then a corresponding notification is output.
[0025] Preferably, the step of outputting a notification can also be performed only after the analysis. In this case, the analysis is completed first, and the corresponding notification is output based on the results.
[0026] In some examples, the devices into which components are assembled may relate to vehicles such as cars or trucks. As new cars and trucks continue to evolve, onboard electronics become increasingly complex, utilizing additional connectors. Furthermore, specific bolts or mounting components can be used to secure individual parts to each other. These connectors and parts to be assembled are typically installed according to an intended installation configuration (e.g., according to a specific mounting location). However, in some cases, improper installation configurations may occur due to the possibility that the snap-fit mechanism may not be properly satisfied. For example, snap-fit tabs may not properly engage with their corresponding snap-fit brackets. Therefore, the components to be assembled may loosen again during the use of the device into which the components are assembled. The methods described above can advantageously be used to assess whether the specified installation configuration is properly satisfied during the assembly process of the components to be assembled.
[0027] Preferably, the components to be assembled can involve at least connectors, bolts, springs, parts, hoses, or cables. In fact, the components to be assembled can involve any component (fixedly or reversibly) (preferably reversibly mounted) installed inside or at the underlying device. In this regard, assembling components into the device is not limited to installing them within the internal space of the underlying device. More precisely, the installation procedure of attaching the components to the outer surface of the device is also within the scope of this method. For example, some parts are mounted on the chassis of a vehicle. Of course, the assembly of these components in that location can also be evaluated using the methods described above.
[0028] Alternatively, the video stream may also include individual, independent subsequent images indicating the assembly process of the parts to be assembled. Therefore, the video stream is not limited to video data, but may also include a collection of subsequently captured, independent images recorded in light of the assembly process of the parts to be assembled.
[0029] A specified installation configuration can be considered as the installation configuration of a component to be assembled, which ultimately determines the assembly procedure. For example, a specified installation configuration may relate to a specific installation location, installation orientation, or coupling with another component or part. Therefore, the specified installation configuration describes how the corresponding component will be assembled into the relevant equipment. Obviously, after the component to be assembled is assembled into the equipment, the actual installation configuration may deviate from the specified installation configuration if an error occurs during the assembly process. For example, an operator might exemplarily connect a connector to a bracket according to an incorrect orientation. In different examples, the connector to be assembled may not be fully inserted into the bracket, resulting in the snap-fit mechanism not fully closing. Several additional improper installation configurations can be envisioned that could cause the actual installation configuration to deviate from the specified installation configuration. According to existing approaches, such deviations are identified in a post-manufacturing evaluation process; however, this requires considerable effort. The method described herein enables the identification of such improper installation procedures on-site (also known as in-plant) during the assembly process. The output notification allows the operator to immediately correct the improper installation configuration so that the actual installation configuration of the component to be assembled is the same as the specified installation configuration. Therefore, production efficiency is greatly improved because the post-manufacturing evaluation process can be omitted.
[0030] In some embodiments, the vision sensor is a movable vision sensor, a wearable vision sensor, or a handheld vision sensor. For example, a movable vision sensor can be attached to an operator (who can wear it) who performs the assembly process of a component (such as a main body cam) to be assembled into the device. Because the vision sensor is movable, a video stream of the assembly process can be appropriately recorded even if the component to be assembled into the device is installed within its internal space (and therefore may not be visible to the external space). Therefore, reliable visual data can be recorded, enabling the assessment of whether the installation procedure was performed appropriately. If the vision sensor is wearable by the operator, the operator can still perform the assembly process manually. For example, the vision sensor can be head-mounted or chest-mounted to be able to record an appropriate video stream.
[0031] According to one aspect, each vision sensor may include a communication device configured to transfer data representing the recorded video stream to an external device, such as evaluation circuitry. Specifically, the data representing the recorded video stream can be communicated in real time. This means that the time delay caused by the communication process is very short and negligible. Alternatively, the data representing the recorded video stream can also be communicated to an external computer, such as an edge PC including the evaluation circuitry.
[0032] Preferably, the communication device of the vision sensor is capable of wireless communication. Wireless communication protocols such as Bluetooth, NFC (Near Field Communication), Wi-Fi, RFID (Radio Frequency Identification), or any other suitable wireless communication technology can be employed in this regard.
[0033] In alternatives, the communication protocol can also involve wired communication standards, such as Ethernet or industrial bus protocols.
[0034] In some cases, the recorded video stream is analyzed in real time. This means that evaluation circuits, particularly those utilizing artificial intelligence based on artificial convolutional networks, are configured to evaluate the corresponding video streams in real time. Therefore, notifications can be output in real time, for example, to the user or operator performing the assembly process. Consequently, improperly executed installation procedures can be corrected on-site, ensuring that the actual installation configuration meets the specified requirements.
[0035] In some embodiments, the method further includes the following initial steps: - Receive vehicle identification number.
[0036] - Based on the received vehicle identification number, all parts to be assembled are read from the data storage device.
[0037] Video streams of all components to be assembled are recorded using at least one vision sensor attached to the operator. The recorded video streams are analyzed by an evaluation circuit utilizing artificial intelligence based on at least one artificial convolutional network for a specified installation configuration of all components to be assembled. Based on the analysis of all components to be assembled by the evaluation circuit, a notification is output via a human-machine interface.
[0038] In other words, the steps of receiving the vehicle identification number and reading out all the parts to be assembled are performed before the step of recording the video stream. This means that the method can be extended to evaluate (especially simultaneously evaluate) multiple parts to be assembled into the equipment. In this regard, the allocation of vehicle identification numbers and corresponding parts to be assembled into the equipment is utilized, as it specifies which parts are of interest. For example, the vehicle identification number can be received when the relevant body (chassis) enters a specific assembly line employing this method and system. Therefore, the data storage device can specify which parts will be assembled along the assembly line into the equipment (here, the vehicle) with the corresponding vehicle identification number and evaluated for improper installation. In particular, the control device coupled to the evaluation circuit or the evaluation circuit itself can receive the corresponding identification number and read out the parts from the data storage device. Therefore, the assembly procedure to be performed along a specific assembly line can be specified within the data storage device, which can then be used to perform the method described earlier herein. In some cases, the vehicle identification number can be associated with the chassis number.
[0039] In some examples, the data storage device may also include different sets of vehicle identification numbers and corresponding components to be assembled into different assembly lines or manufacturing sites. This provides the possibility of utilizing distributed data storage devices, such as server devices, accessible by corresponding control devices or evaluation circuitry.
[0040] In some implementations, the data storage device includes at least one web-enabled database of components to be assembled and associated vehicle identification numbers. Additional components to be assembled and associated vehicle identification numbers can be added to the database via a user interface. Therefore, the database can be effectively expanded, and new information on additional components to be assembled can be added. For example, the database may also include information about different manufacturing sites or assembly lines. In other words, the database can contain different groups of vehicle identification numbers and components to be assembled, depending on the different manufacturing sites or assembly lines. Thus, a unified database accessible via the Internet can be established, enabling efficient management. It also provides the possibility of adjusting the method based on new components to be assembled into the equipment if the equipment design is updated or modified.
[0041] User interfaces can, for example, involve networked interfaces accessible via internet communication protocols. For instance, a user interface can be related to a website.
[0042] In some cases, the evaluation circuit can also be configured to request a vehicle identification number from the manufacturing site or assembly line. In this regard, the evaluation circuit may include, or be coupled to, a communication device that transmits the corresponding request to control equipment at the manufacturing site or assembly line. Once a new vehicle identification number is received, the method described above is performed for that newly received vehicle identification number.
[0043] According to one embodiment, once the vehicle identification number is received and the corresponding component to be assembled is read from the data storage device, the remaining steps of the method are executed simultaneously with the start of the assembly process. The progress of the method can then be optionally (e.g., via user notification) provided to the operator assembling the corresponding component into the device. Thus, the operator is informed of the current status of the method.
[0044] Based on some examples, the method further includes the following steps: - If the actual installation configuration of at least one component to be assembled does not correspond to the specified installation configuration of at least one component to be assembled, the assembly process of at least one component to be assembled into the device shall be interrupted.
[0045] This provides the possibility of interrupting the assembly process when the actual installation configuration differs from the specified installation configuration. In some examples, the assembly line can be interrupted to interrupt the assembly process. For example, the assembly line may include a conveyor or different transport mechanism for transporting the equipment to which the components are assembled from from the inlet to the outlet. This conveyor or corresponding transport mechanism can be interrupted until the incorrect installation configuration of the components is adjusted to correspond to the specified installation configuration. Therefore, the deviation between the actual and specified installation configurations is corrected on-site, and optionally in real-time, without the additional need for post-manufacturing evaluation procedures. This also provides the advantageous effect of informing the operator assembling the components into the equipment of the improper assembly procedure and making them aware of the measures required to meet the specified installation configuration.
[0046] Preferably, for all components to be assembled within the corresponding equipment, the evaluation circuit may need to determine at the end of the equipment (optionally a vehicle) assembly process that the actual installation configuration corresponds to a specified installation configuration in order to release the equipment from the manufacturing site or assembly line. If any deviation is determined between the actual installation configuration and the specified installation configuration of the components to be assembled, the assembly process can be interrupted and / or the release of the equipment from the manufacturing site or assembly line can be interrupted based on control signals issued by the evaluation circuit, as previously described. Therefore, the method ensures that any event involving the release of equipment with improperly assembled components is prevented.
[0047] Common artificial neural networks consist of neurons. A typical artificial neural network comprises several layers, each containing multiple neurons. Each neuron in a particular layer of a related artificial neural network is typically coupled to all neurons in previous and subsequent layers (so-called fully connected neurons). Each neuron is assigned a weight distribution, which can be viewed as a probability map—that is, how the input signal received by the neuron from a particular previous neuron is modified given the forwarding of the corresponding signal to a particular subsequent neuron. In some nomenclature, the weight distribution can be represented as a vector specifying the weight distribution of a single neuron's connections.
[0048] Compared to conventional neurons in common artificial neural networks, in one embodiment, an artificial convolutional network may include at least one learnable kernel. The kernel differs from a common neuron in that it is not coupled to all neurons in previous and subsequent layers. The convolutional layers and the network's kernels are only partially connected. Therefore, a specific kernel in a convolutional layer of the network is only coupled to a portion of the neurons or kernels in previous and subsequent layers. Thus, computational costs are reduced by implementing an artificial convolutional network with kernels. However, the precision of the processing employed by an artificial convolutional network remains high. This is because the kernel is able to extract specific features of the processed signal, allowing modifications induced by the kernel to be precisely tailored to the extracted features. Therefore, each kernel is assigned a feature map (sometimes called an activation map, typically represented as a vector). The feature map describes the relationship between the received input signal and the provided output signal. In other words, the feature map describes how the input signal received by the kernel is processed by the kernel during convolution. Therefore, a kernel can be considered as part of an artificial neural network comprising a set of learnable weights and biases that are applied to the received input signal during the convolution operation.
[0049] Because the kernel of an artificial convolutional network is configured to extract specific features from the processed signal (in this case, from processed video data such as video streams or images), the requirement to force the recorded video stream to be in focus across all its parts can be omitted. Even if the relevant recorded video stream is not recorded corresponding to the focus of the video sensor, the artificial convolutional network can recognize the shortcomings based on the extracted features. Therefore, based on the artificial convolutional network, the extracted features can enable accurate evaluation of whether a specified installation configuration is met, even though the relative distance between the vision sensor and the component to be assembled varies in the recorded video stream.
[0050] Preferably, the evaluation circuitry for artificial intelligence based on at least one artificial convolutional network is machine learning-based. Therefore, the learnable kernel represents an architecture, and feature maps describing the relevant kernel functions can be adjusted for this architecture. Specifically, the feature maps can be adjusted adaptively based on the training procedure or during normal operation. Thus, the artificial convolutional network utilizes machine learning.
[0051] According to one aspect, an artificial convolutional network can be a part of an artificial neural network. While an artificial neural network can include classical neurons (as mentioned earlier, classical neurons can be fully connected neurons), an artificial convolutional network can also construct part (or all) of an artificial neural network. Therefore, the artificial network can be tailored as needed. For example, the input and / or output layers of the artificial network can include fully connected layers with fully connected neurons instead of kernels. In this respect, the input layer receives the input signal provided to the artificial network, namely the recorded video stream and optional information about the parts to be assembled. The output layer provides the output signal of the artificial network, namely, an evaluation of whether the actual installation configuration of the parts to be assembled corresponds to the specified installation configuration. Therefore, the input procedure for inputting the corresponding signals into the artificial network and the extraction procedure for outputting the corresponding signals from the artificial network can be simplified and designed according to known criteria.
[0052] In some embodiments, an artificial network may include at least two hidden layers. In this regard, the hidden layers may be hypothetical layers arranged between the input layer and the output layer. By having at least two subsequent hidden layers, the artificial network is based on deep learning.
[0053] Deep learning enables artificial networks to employ complex processing routines by executing multiple subsequent independent processing procedures. These subsequent independent procedures collectively establish a common processing routine. This common processing routine leads to co-processing, which, for independent subroutines, is based on a broader foundation of potential impact, thereby improving the overall quality of the results.
[0054] Deep learning can be supervised, semi-supervised, or unsupervised. In particular, unsupervised deep learning routines offer the possibility of further improving the operational efficiency required to optimize processing procedures.
[0055] According to one aspect, an evaluation circuit utilizing artificial intelligence based on at least one artificial convolutional network includes at least one object detection algorithm. The object detection algorithm may particularly relate to the YOLOv5 algorithm. The evaluation circuit utilizing artificial intelligence based on at least one artificial convolutional network analyzes a recorded video stream based on the object detection algorithm to identify and monitor at least one component to be assembled within the recorded video stream. The object detection algorithm is able to monitor the component to be assembled within the video stream provided to the evaluation circuit, thereby allowing for accurate conclusions regarding whether the actual assembly configuration corresponds to a specified assembly configuration. In particular, the YOLOv5 algorithm has been shown to accurately identify the corresponding features of objects within video data. Furthermore, object detection algorithms (such as the YOLOv5 algorithm) can be capable of analyzing the recorded video stream in real time, for example, with negligible time delays.
[0056] In some cases, object detection algorithms can be designed to be trained using supervised machine learning. In other words, an object detection algorithm can be fed a set of training images to fine-tune itself for effectively monitoring parts within the provided video data. In this regard, it may be necessary to label the parts that the object detection algorithm will monitor within at least a portion of the training data provided to the algorithm, for example, by bounding boxes. Typically, the training data can involve a video stream or multiple independent images. For example, the parts to be monitored can be labeled in the initial images of a set of training images or the initial images of a video stream. Based on the specificity of the labeled objects, the algorithm can then be able to identify and monitor the corresponding objects throughout the entire training data.
[0057] In some embodiments, the evaluation circuit utilizing artificial intelligence based on at least one artificial convolutional network can be trained via a training interface based on a training dataset comprising at least one training video stream of at least one component to be assembled. The at least one component to be assembled is labeled within at least a portion of the training video stream. Information regarding the specified installation configuration of the at least one component to be assembled within at least the training video stream is assigned to the training video stream. Here, the training video stream is not limited to an actual video stream representing a recorded frame (film). Instead, the training video stream can also be a set of independent training images capable of monitoring the assembly process of components to be assembled into a device. Based on the training data, information about which component will be monitored within the training video stream is provided to the artificial convolutional network. This is done when the corresponding component is labeled. For example, frames can be drawn around the corresponding component within a portion of the training video stream or in an initial image.
[0058] Artificial convolutional networks can monitor the components to be assembled during the assembly process and assess whether the actual installation configuration corresponds to the specified installation configuration received by the network. For control purposes, information about the actual installation configuration of the assembly process for the training video stream can also be provided to the artificial convolutional network.
[0059] Based on the comparison of autonomous evaluation and control information on whether the actual installation configuration corresponds to the specified installation configuration, the intrinsic parameters of the artificial convolutional network can be adaptively adjusted to improve the accuracy of the evaluation procedures for subsequent training events and routine operations.
[0060] In other words, according to one example, a training video stream is fed to an artificial convolutional network (ACNN). Within the training video stream (e.g., within its initial image), the parts to be assembled are labeled. The ACNN then monitors the parts being assembled during the training video stream. Specifically, the ACNN evaluates whether the actual assembly configuration matches a specified assembly configuration, and the network has been pre-provided with information related to the specified assembly configuration. Ultimately, the ACNN determines whether the autonomous actual assembly configuration of the parts to be assembled corresponds to the specified assembly configuration. This result can be compared with control information indicating whether a correspondence actually exists between the actual assembly configuration and the specified assembly configuration of the parts to be assembled for the assembly procedure represented by the training video stream.
[0061] If the results determined by the artificial convolutional network deviate from the control information, the intrinsic parameters of the artificial convolutional network (such as the feature maps of the convolutional network kernel) can be adaptively adjusted by the network itself. In this respect, the network employs machine learning or deep learning. This training procedure can be performed under unsupervised, semi-supervised, or supervised conditions. Supervised training procedures can converge faster, thereby improving training efficiency.
[0062] Alternatively, the training interface can be viewed as an interface through which the artificial convolutional network can be trained by providing training video streams. For example, the training interface can also be established by a webpage through which different users can access the artificial convolutional network to provide additional training datasets. Thus, the training interface enables the provision of new training datasets, such as for additional components for which an installation configuration is to be evaluated. Furthermore, the artificial convolutional network can also be used for additional assembly procedures, such as assembly procedures performed in different manufacturing sites or assembly lines.
[0063] According to one aspect, the training interface segments the training video stream into individual, independent training images and provides these individual training images to an evaluation circuit utilizing artificial intelligence. For example, a 30-second video stream can be divided into approximately one thousand subsequent training images, corresponding to 33 images per second. This high density allows the artificial convolutional network to accurately monitor the parts to be assembled during the assembly process. Therefore, the accuracy of the evaluation process regarding whether the actual assembly configuration corresponds to the specified assembly configuration is improved.
[0064] During the training process, i.e., while providing the training dataset and tuning the intrinsic parameters of the artificial convolutional network, data augmentation and fine-tuning algorithms can also be used. For example, the data involved in the training dataset can be post-processed through data augmentation to adapt to the needs of the artificial convolutional network. Furthermore, fine-tuning algorithms that combine supervised and unsupervised machine learning techniques can improve training efficiency. In some cases, the training operator can adjust specific parameters during the training process to accelerate convergence, thereby reducing the time required for training.
[0065] According to another aspect, this disclosure also relates to a data processing apparatus comprising means for performing a computer-implemented method as described above, the method for monitoring the assembly process of at least one component to be assembled into the apparatus as described above. The advantages achieved by the above-described method are also correspondingly achieved by the data processing apparatus.
[0066] According to another aspect, this disclosure also relates to a computer program product comprising instructions that, when executed by a computer, cause the computer to perform the computer-implemented method as described above for monitoring the assembly process of at least one component to be assembled into a device. The advantages achieved by the aforementioned computer-implemented method are also correspondingly achieved by the computer program product.
[0067] According to another aspect, this disclosure also relates to a computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform a computer-implemented method as described above for monitoring the assembly process of at least one component to be assembled into a device. The advantages achieved in view of the above-described computer-implemented method are also correspondingly achieved by the computer-readable storage medium. Attached Figure Description
[0068] The foregoing aspects and further advantages of the claimed subject matter will be more readily understood and better appreciated when taken in conjunction with the accompanying drawings and by referring to the following detailed description. In the accompanying drawings, - Figure 1 This is a schematic diagram of a system according to an embodiment for monitoring the assembly process of at least one component to be assembled into a device. - Figures 2A to 2C This is a schematic diagram illustrating an exemplary scenario of the actual in-process installation configuration and the specified installation configuration within the system, and... - Figure 3 This is a schematic diagram of a method for monitoring the assembly process of at least one component to be assembled into a device, according to an embodiment. Detailed Implementation
[0069] The detailed description set forth below with reference to the accompanying drawings (where like numbers refer to like elements) is intended as a description of various embodiments of the disclosed subject matter and is not intended to represent the only embodiment. Each embodiment described in this disclosure is provided by way of example or illustration only and should not be construed as superior to or advantageous to other embodiments. The illustrative examples provided herein are not intended to be exhaustive or to limit the claimed subject matter to the precise forms disclosed. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the described embodiments. Therefore, the described embodiments are not limited to the embodiments shown but should be accorded the widest scope consistent with the principles and features disclosed herein.
[0070] All features disclosed below with respect to exemplary embodiments and / or drawings may be combined with features of aspects of this disclosure (including features of preferred embodiments thereof) individually or in any sub-combination, provided that the resulting combination of features is reasonable to those skilled in the art.
[0071] For the purposes of this disclosure, the phrase "at least one of A, B, and C" refers, for example, to (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C), including all further possible combinations when more than three elements are listed. In other words, the term "at least one of A and B" generally means "A and / or B," i.e., "A" alone, "B" alone, or "A and B."
[0072] Figure 1 This is a schematic diagram of a system 10 for monitoring the assembly process of at least one component 12 to be assembled into device 14, according to an embodiment.
[0073] Here, the device 14 into which component 12 will be assembled is a vehicle. The component 12 to be assembled can be a bolt, connector, hose, or different parts. The assembly procedure is performed by operator 18 at assembly line 16.
[0074] Operator 18 is equipped with vision sensors 20. The first vision sensor 20 is head-mounted, while the second vision sensor 20 is chest-mounted. In other words, the vision sensors 20 can be worn by operator 18.
[0075] Each vision sensor 20 is configured to record a video stream of the assembly process of component 12 into device 14. Since the vision sensors 20 are worn by operator 18, video streams can be recorded for several procedures even if the component 12 to be assembled into device 14 is installed inside the device 14 (which may not be easily identifiable from outside the device 14). Therefore, operator 18 wearing the vision sensors 20 is able to record video streams that can be used to precisely monitor the details of the assembly process.
[0076] Typically, component 12 is assembled into device 14 according to specified installation configuration 22. However, due to errors occurring during the assembly process, the actual installation configuration 24 may deviate from specified installation configuration 22 after component 12 is initially assembled into device 14. In this regard, Figures 2A to 2C This is a schematic diagram of an exemplary scenario of actual installation configuration 24 and specified installation configuration 22 that may occur when using system 10.
[0077] Typically, component 12 to be assembled into device 14 will be coupled to another component, such as bracket 26 of device 14. The difference between the actual installation configuration 24 and the specified installation configuration 22 can, in most cases, be characterized by displacement or incorrect orientation according to the indicated Cartesian coordinate system.
[0078] according to Figure 2A In the scenario shown, the actual installation configuration 24 differs from the designated installation configuration 22 in that component 12 is not assembled such that it is fully inserted into the bracket 26, which, according to the designated installation configuration 22, component 12 should be inserted into the bracket 26. Therefore, the actual installation configuration 24 of component 12 differs from the designated installation configuration 22 in terms of displacement along the x-axis.
[0079] according to Figure 2B In the exemplary scenario shown, component 12 is fully inserted into bracket 26, but the actual installation configuration 24 corresponds to an incorrect orientation relative to the specified installation configuration 22. This means that component 12 is incorrectly inserted into bracket 26.
[0080] Figure 2C The third exemplary scenario shown indicates that incorrect coupling may also exist. Although component 12, to be assembled into device 14, is correctly inserted into bracket 26 in principle, the specified installation configuration 22 indicates that component 12 is inserted into the wrong bracket 26.
[0081] Therefore, due to errors that occur during the assembly process, the actual installation configuration 24 may differ from the specified installation configuration 22.
[0082] The vision sensor 20 is configured to record a video stream of the assembly process of the component 12 to be assembled into the device 14 in a manner that can be distinguished from the designated installation configuration 22 in the actual installation configuration 24.
[0083] According to this embodiment, device 14 (here, a vehicle) includes control device 28 coupled to communication device 30.
[0084] The assembly process takes place at assembly line 16 in manufacturing site 32. Assembly line 16 includes a conveyor mechanism 34 for moving device 14 along a given trajectory. For this purpose, motor unit 36 is coupled to conveyor mechanism 34.
[0085] To analyze the video stream recorded by the vision sensor 20, the system 10 includes an evaluation circuit 38 that utilizes artificial intelligence based on at least one artificial convolutional network 40. The evaluation circuit 38 may, for example, be part of an edge PC assigned to the assembly line 16.
[0086] According to this embodiment, the evaluation circuit 38 is configured to communicate with the control device 28 of the device 14 to which the component 12 will be assembled. For this purpose, the evaluation circuit 38 is coupled to a communication device 42, which is configured to communicate with the communication device 30 of the device 14. For example, through communication between the communication device 42 and the communication device 30 of the device 14, information about the device 14 (such as a vehicle identification number) can be exchanged, allowing the system 10 to identify which device 14 is currently in a process along the assembly line 16.
[0087] Preferably, the communication between communication device 42 and communication device 30 of device 14 is wireless, for example, based on Wi-Fi.
[0088] Once the evaluation circuit 38 learns information about a specific identifier of the equipment 14 that is in the process of moving along the assembly line 16, the evaluation circuit 38 can also utilize the data storage device 44, which includes the database 46.
[0089] Database 46 stores specific datasets of vehicle identification numbers or other ID types, which allow for exclusive identification of the equipment 14 being processed. Furthermore, database 46 also includes information on components 12 to be assembled into specific equipment 14. In this regard, the corresponding component 12 to be assembled into specific equipment 14 is assigned a corresponding vehicle identification number or corresponding ID type. Additionally, multiple different datasets may be included for different manufacturing sites 32 or different assembly lines 16.
[0090] According to this embodiment, the data storage device 44 and the database 46 are networked, enabling them to be accessed through a user interface 48. For example, the user interface 48 can be established via a webpage, through which access to the database 46 can be granted. The user interface 48 allows specific datasets in the database 46 to be added or modified as needed. This allows for the redefinition or modification of components 12 to be assembled into device 14, given a specific device 14, which can be identified within the database 46 by a corresponding vehicle identification number or a corresponding ID type.
[0091] To analyze the video stream recorded by the vision sensor 20, an artificial convolutional network 40 of the evaluation circuit 38 is applied. According to this embodiment, the artificial convolutional network 40 is part of a more general artificial neural network 50. In this respect, the artificial neural network 50 includes several layers 52, such as an input layer 54, an output layer 56, and a hidden layer 58 between the input layer 54 and the output layer 56.
[0092] Optionally, the analysis of the recorded video stream can be performed in real time by the evaluation circuit 38.
[0093] In an alternative approach, analysis of the recorded video stream can only be performed after the stream has been recorded. In this case, the recorded video stream can be stored intermediately in data storage device 44.
[0094] Each layer 52 of the artificial neural network 50 includes at least one artificial neuron 60, or, in the case of a convolutional layer 62, a kernel 64.
[0095] In principle, the artificial neural network 50 receives the input signal provided to the input layer 54. Then, as will be described in more detail below, the input signal is modified based on the artificial neurons 60 and the kernel 64. The output layer 56 provides the output signal for further processing.
[0096] While the artificial neurons 60 are typically fully connected neurons because they are coupled to each neuron 60 of the previous layer 52 and all neurons 60 of the subsequent layer 52, the kernel 64 is only coupled to a portion of the neurons 60 or kernel 64 of the previous and subsequent layers 52.
[0097] Each artificial neuron 60 is assigned a weight distribution, which can be viewed as a probability map, i.e., how to modify the input signal received by the corresponding artificial neuron 60 from a specific previous neuron 60, given that a corresponding signal is forwarded to a specific subsequent neuron 60. In contrast to the artificial neurons 60, the kernels 64 of the convolutional layers 62 and the convolutional network 40 are only partially connected. This is because the kernels 64 of the artificial convolutional network 40 extract specific features of the processed signal, such that the modifications induced by the artificial kernels 64 are precisely tailored to the extracted features. Therefore, the artificial convolutional network 40 with learnable kernels 64 allows for customization to suit the intended purpose of the evaluation circuit 38.
[0098] In principle, the convolutional network 40 with learnable kernel 64 is configured to evaluate the video stream recorded by the vision sensor 20 in order to determine the actual installation configuration 24 of the component 12 to be assembled into the device 14, given the corresponding specified installation configuration 22. In other words, based on the communication between the communication device 42 and the communication device 30 of the device 14, the evaluation circuit 38 can receive the corresponding vehicle identification number of the device 14. Using this information, the evaluation circuit 38 can access the database 46 of the data storage device 44 and read out which components 12 will be assembled into the device 14.
[0099] Furthermore, the communication device 42 can also be configured to receive video streams recorded from the vision sensor 20, which can communicate with the corresponding video stream via a wireless communication protocol such as Wi-Fi, Bluetooth, NFC, RFID (Radio Frequency Identification), or any other suitable wireless communication technology. Therefore, the artificial convolutional network 40 can evaluate whether the components 12 to be assembled into the device 14 are assembled therein, such that their actual installation configuration 24 corresponds to the specified installation configuration 22 specified in the database 46. This evaluation procedure is performed in real time by the evaluation circuit 38, so the time delay is negligible given the assembly process. As an output signal, the evaluation circuit 38 indicates whether the actual installation configuration 24 corresponds to the specified installation configuration 22. Subsequently, the evaluation circuit 38 can use the communication device 42 to initiate the output of a notification via the human-machine interface 66 to inform the user or operator 18 of the system 10 of the established correspondence between the actual installation configuration 24 and the specified installation configuration 22.
[0100] For example, a corresponding notification output via human-machine interface 66 can indicate that the actual installation configuration 24 does not correspond to the specified installation configuration 22.
[0101] If the evaluation circuit 38 determines that the actual installation configuration 24 does not correspond to the specified installation configuration 22, in some embodiments, the evaluation circuit 38 may also be configured to interrupt the assembly process of the device 14. Specifically, the evaluation circuit 38 may be configured to prevent the device 14 from leaving the assembly line 16 if the determined actual installation configuration 24 does not correspond to the specified installation configuration 22. To this end, the evaluation circuit 38 may output a corresponding signal to the motor unit 36, causing the conveyor belt 34 to stop. Therefore, the movement of the device 14 is interrupted until the actual installation configuration 24 is corrected by the operator 18. Thus, errors occurring during the assembly process can be identified as existing in real time and can be corrected on-site by the operator 18. Therefore, post-manufacturing evaluation outside the assembly line 16 can be prevented or at least reduced.
[0102] In order to train the artificial neural network 50, especially the artificial convolutional network 40, the artificial neural network 50 is coupled to the training interface 68.
[0103] In some embodiments, the training interface 68 may be networked. A training dataset 70 can be provided via the training interface 68. The training dataset 70 may include a training video stream of the assembly process of the component 12 to be assembled into the device 14, and corresponding information about the specified installation configuration of the respective component 12. In some cases, the training interface 68 may be configured to split the training video stream into separate training images, which allows the relevant components to be assembled according to the training dataset 70 to be labeled, for example, by bounding boxes. Based on the labeling, the user of the training dataset 70 provides sufficient information to the evaluation circuitry 38, i.e., which component will be evaluated within the training dataset 70. Alternatively, the component 12 to be assembled may also be labeled by the user within a portion of the training video stream (e.g., within its initial image).
[0104] The evaluation circuit 38, based on artificial intelligence using an artificial convolutional network 40, is machine learning-based. This means that the evaluation circuit 38 is configured such that the weight distributions of the artificial neurons 60 and the learnable artificial kernel 64 can adaptively adjust their weight distributions and feature maps with respect to how the corresponding input signals are modified. This makes the training procedure using the training dataset 70 applied by the training interface 68 efficient. For example, for the first training run using the first training dataset 70, the evaluation circuit 38 can determine that the actual installation configuration 24 corresponds to the specified installation configuration 22, even though they are actually different from each other; this can be indicated as control information within the training dataset 70. To improve the accuracy of the relevant determination procedure performed by the evaluation circuit 38, the weight distributions of the artificial neurons 60 and the feature maps of the artificial kernel 64 can be adjusted for subsequent training runs. Therefore, for subsequent training runs, the accuracy of the determination procedure regarding whether the actual installation configuration 24 corresponds to the specified installation configuration 22 can be improved. In this respect, training can be performed in an unsupervised, semi-supervised, or supervised manner. Typically, for supervised training procedures, the time required to reach convergence such that the relevant determination procedure of the evaluation circuit 38 represents an appropriate correspondence of the differences is minimized.
[0105] To monitor and identify components 12 to be assembled into device 14, the AI-based evaluation circuit 38 may employ an object detection algorithm 72. In this embodiment, the object detection algorithm 72 may include the YOLOv5 algorithm. Since the components 12 to be assembled into device 14 are labeled (e.g., by bounding boxes) within the training dataset 70, the object detection algorithm 72 is able to monitor the corresponding components 12 during the training dataset 70 and the associated assembly process. Therefore, the training procedure also results in the training of the object detection algorithm 72, which can then be used during the normal operation of the evaluation circuit 38. This allows the evaluation circuit 38 to monitor the corresponding components 12 to be assembled into device 14 within the video stream recorded by the vision sensor 20 during normal operation, without needing to label the components 12 within the recorded video stream. In other words, based on the training procedure, the object detection algorithm 72 can autonomously monitor the components 12 within the corresponding video stream.
[0106] Figure 3 This is a schematic diagram of a method 80 for monitoring the assembly process of at least one component 12 to be assembled into device 14 according to an embodiment. Optional steps are shown in dashed lines.
[0107] According to optional step S1 of method 80, a vehicle identification number is received. Specifically, the vehicle identification number can be received by evaluation circuit 38. For this purpose, evaluation circuit 38 can utilize communication device 42, which can request the vehicle identification number from device 14. Alternatively, device 14 can send the vehicle identification number when it enters assembly line 16.
[0108] Following optional step S1, method 80 may further include optional step S2, according to which all parts 12 to be assembled are read from data storage device 44 based on the vehicle identification number received (particularly received via evaluation circuit 38). In other words, once it is known which exact device 14 is being processed along assembly line 16, the database 46 of data storage device 44 can be used to evaluate which parts 12 will be assembled into device 14. Since database 46 includes datasets (in which the corresponding parts 12 to be assembled are assigned to the corresponding vehicle identification numbers), this provides the evaluation circuit 38 with the possibility of identifying the parts 12 to be evaluated.
[0109] According to step S3, a video stream of the component 12 to be assembled is recorded using the vision sensor 20. If, prior to step S3, the evaluation circuit 38 identifies multiple components 12 to be assembled into the device 14, a video stream is recorded for all components 12 to be assembled. Because the vision sensors 20 are movable (as they can be worn by the operator 18), an appropriate video stream can be recorded even if the component 12 is assembled in a volume inside the device 14 that is not visible to the external space.
[0110] In this embodiment (solid line), step S4 is performed after step S3. In step S4 of method 80, for a specified mounting configuration 22 of the component 12 to be assembled, an evaluation circuit 38 utilizing artificial intelligence based on at least one artificial convolutional network 40 analyzes the recorded video stream. In this regard, the evaluation circuit 38 utilizes the convolutional network 40 and its kernel 64 to determine the actual mounting configuration 24 of the component 12 assembled into the device 14 and compares the actual mounting configuration 24 with the specified mounting configuration 22. Information about the specified mounting configuration 22 is obtained by reading the corresponding information from the data storage device 44 in step S2.
[0111] Clearly, step S4 is performed on all components 12 to be assembled into device 14. This means that evaluation circuit 38 analyzes the video stream recorded by vision sensor 20 in view of all components 12 assembled into device 14. For the analysis of the assembly process, evaluation circuit 38 utilizes object detection algorithm 72. Thus, the corresponding component 12 can be reliably monitored within the recorded video stream. Therefore, at least the recorded video stream and the specified installation configuration 22 are input signals to artificial neural network 50. The output signal of artificial neural network 50 indicates whether the actual installation configuration 24 corresponds to the specified installation configuration 22.
[0112] exist Figure 3In the alternative indicated by the dashed line, steps S3 and S4 are executed with at least partial overlap in time. In this case, the analysis performed by the evaluation circuit is done in real time. Therefore, immediate results of the analysis are readily available, enabling rapid notification.
[0113] As a result of the analysis performed in step S4, according to the illustrated embodiment (solid lines), method 80 includes a subsequent step S5, in which a notification is output via human-machine interface 66. The notification depends at least on the actual installation configuration 24 of the component 12 to be assembled and the designated installation configuration 22 of the component to be assembled. Specifically, the notification may indicate whether the actual installation configuration 24 corresponds to the designated installation configuration 22. Furthermore, the notification may also indicate any discrepancies between the actual installation configuration 24 and the designated installation configuration 22 as determined by evaluation circuitry 38. Thus, the notification indicates whether component 12 has been properly assembled into device 14. If not, the notification output via human-machine interface 66 informs operator 18 of the actions required to make the actual installation configuration 24 correspond to the designated installation configuration 22.
[0114] Because the evaluation circuit 38 uses an artificial convolutional network 40, the computational cost is significantly reduced compared to existing approaches using common artificial neural networks 50. This is because the artificial convolutional network 40 is implemented using a learnable kernel 64, which processes the corresponding input signal through the corresponding feature map. In particular, the computational cost is reduced because the kernel 64 is only connected to subsequent and previous artificial neurons 60 or a portion of the kernel 64. Therefore, fewer computational procedures are necessarily performed, thereby improving the evaluation efficiency of the assembly process of component 12.
[0115] Based on notifications provided via human-machine interface 66, operator 18 is directly informed if the assembly process of component 12 is performed appropriately. Therefore, it can be ensured that no equipment 14 leaves assembly line 16 without all components 12 being properly assembled into it. Thus, post-manufacturing evaluation outside of assembly line 16 can be prevented.
[0116] Obviously, if multiple components 12 will be assembled into the device 14, then in step S5, based on the corresponding correspondence between the actual installation configuration 24 and the specified installation configuration 22, a notification for each component 12 is output via the human-machine interface 66.
[0117] In an alternative embodiment (dashed line), step S5 may also be performed in time at least partially overlapping with steps S3 and / or S4. Therefore, the notification is output more quickly.
[0118] Optionally, method 80 may further include step S6, wherein evaluation circuit 38 may be configured to interrupt the assembly process of component 12 to be assembled into device 14 if the actual mounting configuration 24 of component 12 to be assembled does not correspond to the specified mounting configuration 22 of component 12 to be assembled. In this regard, evaluation circuit 38 may output a corresponding signal that causes motor unit 36 to stop the movement of conveyor mechanism 34, so that device 14 stops in place without leaving assembly line 16. Thus, it is possible to prevent any device 14 from leaving assembly line 16 without properly assembling all components 12 into the device 14. Of course, alternative mechanisms for interrupting the movement of device 14 are also conceivable. Essentially, it is possible to prevent device 14 from leaving manufacturing site 32 without properly assembling all components 12 into the corresponding device 14.
[0119] Method 80 may further include an optional step S7, according to which an evaluation circuit 38 based on artificial intelligence using at least one artificial convolutional network 40 is trained via a training interface 68 on a training dataset 70, which includes at least one training video stream of the parts 12 to be assembled. Specifically, during the training procedure, the machine learning capabilities of the evaluation circuit 38 can be utilized. Therefore, feature maps from the adaptively adjusted weight distribution of the artificial neurons 60 and the artificial kernel 64 can be evaluated based on whether the provided training dataset 70 is appropriately composed of an artificial neural network 50 including the artificial convolutional network 40.
[0120] To this end, the parts 12 to be assembled are marked in at least a portion of the training video stream of the training dataset 70. Furthermore, information regarding the specified installation configuration 22 of the parts 12 to be assembled within at least the training video stream is assigned to the training video stream. Additionally, the training dataset 70 may also include information regarding the correspondence between the actual installation configuration 24 and the specified installation configuration 22 related to the assembly process in the corresponding training dataset 70.
[0121] According to optional step S8, training interface 68 segments the training video stream into individual independent training images and provides these individual independent training images to evaluation circuit 38 utilizing artificial intelligence. Therefore, since evaluating individual images may be easier than evaluating the actual video stream, computational costs are further reduced. Furthermore, segmentation into individual images provides the possibility of labeling the components 12 involved in the assembly process of training dataset 70 within specific images of training dataset 70 (e.g., within its initial images).
[0122] Evaluation circuit 38, vision sensor 20, or a separate device (e.g., a control device coupled to evaluation circuit 38) may also be configured to split the video stream recorded by vision sensor 20 into separate images, such that the evaluation performed in step S4 of method 80 can be performed in a manner similar to a training procedure.
[0123] Some embodiments disclosed herein, particularly corresponding modules (one or more) and / or units (one or more), utilize circuitry (e.g., one or more circuits) to implement the standards, protocols, methods, or techniques disclosed herein, operatively coupling two or more components to generate information, process information, analyze information, generate signals, encode / decode signals, convert signals, transmit and / or receive signals, control other devices, etc. Any type of circuitry can be used.
[0124] In one embodiment, the circuit includes, among other things, one or more computing devices, such as a processor (e.g., a microprocessor), a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a system-on-a-chip (SoC), etc., or any combination thereof, and may include discrete digital or analog circuit elements or electronic devices, or combinations thereof. In one embodiment, the circuit includes hardware circuit implementations (e.g., implementations in analog circuits, implementations in digital circuits, etc., and combinations thereof).
[0125] This application may reference quantities and numbers. Unless otherwise stated, these quantities and numbers should not be considered limiting, but rather examples of possible quantities or numbers associated with this application. Also in this respect, this application may use the term "multiple" to refer to quantities or numbers. In this respect, the term "multiple" means any number more than one, such as two, three, four, five, etc. The terms "approximately," "about," "close to," etc., indicate plus or minus 5% of the stated value.
[0126] Although this disclosure has described and illustrated one or more embodiments, equivalent changes and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. Furthermore, while a particular feature of this disclosure may be disclosed only for one of several embodiments, such feature may be combined with one or more other features of other embodiments, which may be desirable and advantageous for any given or particular application.
Claims
1. A method (80) for monitoring the assembly process of at least one component (12) to be assembled into a device (14), wherein the method (80) comprises the following steps: - Use at least one visual sensor (20) to record at least one video stream of the at least one component (12) to be assembled. - For the specified installation configuration (22) of the at least one component (12) to be assembled, an evaluation circuit (38) utilizing artificial intelligence based on at least one artificial convolutional network (40) analyzes at least one recorded video stream, and - Based on the analysis performed by the evaluation circuit (38), a notification is output via the human-machine interface (66), wherein the notification depends at least on the actual installation configuration (24) of the at least one component (12) to be assembled and the specified installation configuration (22) of the at least one component (12) to be assembled.
2. The method (80) according to claim 1, wherein the method (80) further comprises the following initial steps: -Receive vehicle identification number, and -Based on the received vehicle identification number, all parts to be assembled (12) are read from the data storage device (44). The at least one vision sensor (20) attached to the operator (18) is used to record a video stream of all the parts (12) to be assembled. The specified installation configuration (22) for all components (12) to be assembled is determined by the evaluation circuit (38) utilizing the artificial intelligence based on at least one artificial convolutional network (40), which analyzes the recorded video stream. The analysis performed on all components (12) to be assembled based on the evaluation circuit (38) outputs a notification via the human-machine interface (66).
3. The method (80) of claim 2, wherein the data storage device (44) includes at least one networked database (46) of the components (12) to be assembled and associated vehicle identification numbers, and wherein additional components (12) to be assembled and associated vehicle identification numbers can be added to the database (46) via a user interface (48).
4. The method (80) according to any one of the preceding claims, wherein the method (80) further comprises the following steps: If the actual installation configuration (24) of the at least one component (12) to be assembled does not correspond to the designated installation configuration (22) of the at least one component (12) to be assembled, the assembly process of the at least one component (12) to be assembled into the device (14) is interrupted.
5. The method (80) according to any one of the preceding claims, wherein the artificial convolutional network (40) has at least one kernel (64).
6. The method (80) according to any one of the preceding claims, wherein the artificial convolutional network (40) is part of at least one artificial neural network (50).
7. The method (80) according to any one of the preceding claims, wherein the evaluation circuit (38) utilizing the artificial intelligence based on the at least one artificial convolutional network (40) is based on machine learning.
8. The method (80) according to any one of the preceding claims, wherein the evaluation circuit (38) utilizing the artificial intelligence based on the at least one artificial convolutional network (40) includes at least one object detection algorithm (72), particularly the YOLOv5 algorithm, wherein the evaluation circuit (38) utilizing the artificial intelligence based on the at least one artificial convolutional network (40) analyzes the recorded video stream based on the object detection algorithm (72) to identify and monitor the at least one component (12) to be assembled within the recorded video stream.
9. The method (80) according to any one of the preceding claims, wherein the at least one visual sensor (20) is wearable by an operator (18) who assembles the at least one component (12) to be assembled into the device (14).
10. The method (80) according to any one of the preceding claims, wherein the evaluation circuit (38) utilizing the artificial intelligence based on at least one artificial convolutional network (40) is capable of being trained via a training interface (68) based on a training dataset (70), the training dataset (70) comprising at least one training video stream of at least one component (12) to be assembled, wherein the at least one component (12) to be assembled is labeled within at least a portion of the training video stream, and wherein information of the specified installation configuration (22) of the at least one component (12) to be assembled within at least the training video stream is assigned to the training video stream.
11. The method (80) of claim 10, wherein the training interface (68) segments the training video stream into individual independent training images and provides the individual independent training images to the evaluation circuit (38) utilizing the artificial intelligence.
12. A system (10) for monitoring the assembly process of at least one component (12) to be assembled into a device (14), the system (10) comprising at least one vision sensor (20), an evaluation circuit (38) utilizing artificial intelligence based on at least one artificial convolutional network (40), and a human-machine interface (66). The at least one visual sensor (20) is configured to record at least one video stream of the at least one component (12) to be assembled. The evaluation circuit (38) utilizing the artificial intelligence based on the at least one artificial convolutional network (40) is configured to analyze at least one recorded video stream for a specified mounting configuration (22) of the at least one component (12) to be assembled, and The system (10) is configured to output a notification via the human-machine interface (66) based on the evaluation circuit (38), wherein the notification depends at least on the actual installation configuration (24) of the at least one component (12) to be assembled and the specified installation configuration (22) of the at least one component (12) to be assembled.
Citation Information
Patent Citations
Camera based on edge calculation
CN215897828U
System and method for inspecting wiring harness connector terminal of automobile
KR101665644B1
Automotive connector automated production system and method
KR102112809B1
Real-time anomaly detection for industrial processes
US20220066435A1
Stand-alone inspection apparatus for use in a manufacturing facility
US20220136872A1