Image Multi-View Semantic Change Detection System and Method for Assembly Sequence Monitoring
By designing a multi-view semantic change detection system for image, and adopting dense connection and self-attention mechanisms, the problems of missing and misinstallation caused by single occlusion and texture in mechanical assembly are solved, and intelligent monitoring and accurate identification of the assembly process are achieved.
Patent Information
- Application Number
- CN202210667801.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-14
AI Technical Summary
The existing image change detection technology is difficult to effectively apply in the field of mechanical assembly, especially in the case of severe occlusion and single part color and texture information, the lack of semantic information, resulting in misinstallation and misinstallation problems.
A multi-view semantic change detection system for assembly sequence monitoring is designed, including feature extraction module, attention module, step identification module and measurement module. It adopts a densely connected feature fusion mechanism and a self-attention mechanism that integrates context features to enhance the computer visual feature representation ability, and the step identification module recognizes the assembly stage of changing components.
It realizes intelligent monitoring of the mechanical assembly process, can identify changing areas and assembly stages, improves detection accuracy and efficiency, and alleviates the problems of occlusion and texture singleness.
Smart Images

Figure CN115115819B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and intelligent manufacturing, and particularly to an image multi-view semantic change detection system and method for assembly sequence monitoring. Background Art
[0002] In traditional manufacturing assembly processes, manual and discrete operations are mainly used, which are characterized by many assembly operation links and complex operation processes. With the accelerating replacement cycle of mechanical products, the highly customized production mode has led to increased product complexity, shortened development cycles, and a large number of variants. These factors inevitably affect the production of mechanical products, resulting in problems such as missing parts and misassembly during the product assembly process. Therefore, detecting whether the position information of newly assembled parts in each assembly step is accurate from multiple perspectives helps improve the production efficiency and product quality of mechanical products, accelerate the automation and intelligence of mechanical assembly, and has important research value for the intelligent monitoring of the assembly process of mechanical products.
[0003] Image change detection technology aims to process and analyze images of the same area at different time periods to obtain the changed areas on the images, and has important application values in environmental monitoring, urban planning, disaster monitoring, etc. In recent years, deep learning technology has achieved excellent results in computer vision tasks. The image change detection network methods based on deep learning are mainly divided into two types: supervised change detection network methods and unsupervised change detection network methods. Supervised change detection networks are mainly trained through training samples to obtain an optimal model, and then this optimal model is used to map new data samples to corresponding output results. Since unsupervised change detection networks do not have labeled data, most of these methods directly classify data based on the similarity between data samples to obtain the changed areas.
[0004] Currently, image change detection technology mainly monitors targets with the same perspective such as satellite images and aerial images, but is rarely applied to the mechanical assembly field, and the detection results lack semantic information. This is mainly because compared with satellite images, the parts of mechanical assemblies have characteristics such as serious occlusion, single part color and texture information, making it difficult to perform change detection on the assembly process, and there is also a lack of corresponding datasets. Summary of the Invention
[0005] The purpose of the present invention is to provide an image multi-view semantic change detection system and method for assembly sequence monitoring to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] An image multi-view semantic change detection system for assembly sequence monitoring, comprising: a feature extraction module, an attention module, and a metric module, further comprising: a step recognition module;
[0008] The feature extraction module extracts dual-temporal image feature information of different perspectives input to the detection system respectively;
[0009] The attention module performs weighted processing on the extracted dual-temporal image feature information, and the weighted dual-temporal image feature information is input to the step recognition module and the metric module respectively;
[0010] The step recognition module detects the category of the changed target object, identifies the current assembly stage of the changed component parts, and monitors the assembly sequence;
[0011] The metric module judges the changed area of the image, assigns a value to the changed area according to the target category obtained by the step recognition module, so as to obtain a semantic change image.
[0012] Preferably, the step recognition module has a convolutional neural network that uses the Transformers method to process global feature information.
[0013] Preferably, the feature extraction module has a densely connected feature fusion mechanism. The feature extraction module connects the node outputs in the shallow sub-decoder to the deep sub-decoder nodes. When the feature fusion mechanism works, it sequentially transmits the fine-grained features in the encoder to the deep decoder, and finally outputs multiple groups of feature maps with the same size.
[0014] Preferably, the attention module has a self-attention Cot mechanism that fuses context feature information. The steps of the self-attention Cot mechanism are as follows:
[0015] First, context encoding is performed on the input value through a 3×3 convolution to mine the static context feature information between adjacent keys, thereby generating a static context key;
[0016] Then, according to the mutual relationship between the query and the static context key, two consecutive 1×1 convolutions are used to perform dynamic attention matrix learning under the guidance of the static context key. The learned attention matrix is used to aggregate all input values, thereby realizing the representation of dynamic context feature information;
[0017] Finally, the static context feature information and the dynamic context feature information are fused and output.
[0018] Preferably, the metric module first adds up multiple groups of feature maps output by the feature extraction module, then uses the self-attention Cot mechanism to weight the four groups of feature maps, and at the same time splices the four groups of feature maps, and uses the self-attention Cot mechanism to weight and process again to obtain the extracted features. The extracted features are used to automatically select and focus on more effective information between different groups to generate the image change region.
[0019] A detection method for an image multi-view semantic change detection system based on assembly sequence monitoring is characterized by including the following stages: a dataset establishment stage, a training stage, and a testing stage;
[0020] The dataset establishment stage generates training samples for the image multi-view semantic change detection system for assembly sequence monitoring to learn;
[0021] In the training stage, the feature extraction module learns the assembly image feature information of the training samples, and after being processed by the attention module, the step recognition module, and the metric module, outputs the semantic change image of the training samples, judges whether this semantic change image meets the training requirements, and finally saves the optimal model after multiple trainings;
[0022] In the testing stage, the feature extraction module extracts features from the newly input assembly image and obtains the semantic change image according to the optimal model.
[0023] Preferably, the steps of the dataset establishment stage are:
[0024] First, establish an assembly 3D model with the same size as the assembly in the mechanical and real scenes, divide the assembly model into 3D models of multiple assembly steps, then import the 3D models of each assembly step in turn and color-label each part, set the origin of the coordinate system at the same time and export it as a set format file, then import this file and generate a composite image, collect images from different angles, and finally extract the corresponding color labels in the images, and change the color values in the color labels to be used as the change semantic features.
[0025] Preferably, the steps of the training stage are:
[0026] S1: Input the image at the previous moment from different perspectives as the reference image T1 and the image at the later moment as the image to be detected T2 into the feature extraction module respectively;
[0027] S2: The feature extraction module extracts the feature information of the above-mentioned dual-time images respectively. This module uses the dense connection jump fusion mechanism to increase the weight value of the shallow information of the fine-grained features, so that the network has rich feature information;
[0028] S3: The attention module performs weighted processing on the feature information of the above dual-temporal images, making full use of the contextual feature information between adjacent keys to guide the learning of the dynamic attention matrix, thereby further enhancing the computer vision feature representation ability;
[0029] S4: Input the feature information after weighted processing into the step recognition module and the metric module respectively. The step recognition module determines the current assembly stage, and the metric module obtains the changed area according to the feature information, and assigns a value to the changed area according to the current assembly stage to obtain a semantic change image;
[0030] S5: Continuously iterate steps S1 to S4 using the training sample images in the dataset until the set number of training times is reached, and save the optimal model during the training process. Compared with the prior art, the present invention has the following beneficial effects:
[0031] 1. The image multi-view semantic change detection system for assembly sequence monitoring described in the present invention adds a step recognition module compared with other change detection systems. It can not only detect the changed area of the assembly image, but also identify the current assembly stage of the changed component, overcoming the difficulties of serious occlusion, single part color and texture information of mechanical assembly parts under satellite image monitoring, and facilitating the monitoring of the mechanical assembly sequence.
[0032] 2. The image multi-view semantic change detection system for assembly sequence monitoring described in the present invention enhances the computer vision feature representation ability through a dense connection feature fusion mechanism adopted in the feature extraction module and a self-attention Cot mechanism that fuses contextual features adopted in the attention module, so as to realize the intelligent monitoring of the mechanical product assembly process.
[0033] 3. The image multi-view semantic change detection system for assembly sequence monitoring described in the present invention fuses feature information through a dense connection feature fusion mechanism and a tight skip connection between the encoder and the decoder, which can effectively reduce the loss of shallow feature information of the neural network, maintain high-resolution and fine-grained feature representation, and effectively alleviate problems such as poor processing of edge pixels in the detection result and missed detection of small targets.
[0034] 4. The step recognition module adopted by the image multi-view semantic change detection system for assembly sequence monitoring described in the present invention can effectively encode local information and global information in a tensor, combining the advantages of convolutional neural networks with low sensitivity to spatial induction deviation and data augmentation and the advantages of Transformers such as adaptive weighting of input vectors and global processing, which helps to learn better feature information with fewer parameters and simple training samples.
[0035] 5. The multi-view semantic change detection method for image-based assembly sequence monitoring according to the present invention enhances the weight value of the shallow information of fine-grained features through a densely connected feature fusion mechanism in the training stage, enabling the network to have rich feature information. In addition, the self-attention Cot mechanism that fuses context feature information in the training stage can make full use of the context feature information between adjacent positions in the input information to guide the learning of the dynamic attention matrix, thereby further enhancing the computer vision feature representation ability and improving the monitoring performance of the network architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 FIG. is a schematic diagram of an image multi-view semantic change detection system and method for assembly sequence monitoring provided by the present invention.
[0037] Figure 2 FIG. is a densely connected feature extraction model of an image multi-view semantic change detection system and method for assembly sequence monitoring provided by the present invention.
[0038] Figure 3 FIG. is a self-attention model that fuses context feature information of an image multi-view semantic change detection system and method for assembly sequence monitoring provided by the present invention.
[0039] Figure 4 FIG. is an assembly step recognition model of an image multi-view semantic change detection system and method for assembly sequence monitoring provided by the present invention.
[0040] Figure 5 FIG. is a training flowchart of an image multi-view semantic change detection system and method for assembly sequence monitoring provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] The present invention aims to propose a mechanical assembly sequence monitoring method, which can detect the changes in the assembly process to achieve the monitoring of missing parts, misassembly, assembly steps, etc. To this end, the specific embodiments of the present invention provide an assembly image multi-view semantic change detection system for assembly sequence monitoring; a densely connected feature extraction model; a self-attention model that fuses context feature information; and a training flowchart for the assembly image multi-view semantic change detection for assembly sequence monitoring.
[0043] Refer to Figure 1, an assembly image multi - perspective semantic change detection system for assembly sequence monitoring shown in the present invention includes four components: a feature extraction module, an attention module, a step recognition module, and a metric module. This method takes dual - time images from different perspectives as input, and the dual - time images are images of the same area obtained at different times through satellite remote sensing technology. The feature extraction module extracts the feature information of the dual - time images respectively, and the attention module performs weighted processing on the extracted feature information of the dual - time images to further enhance the computer vision feature representation ability; the weighted feature information is respectively input into the step recognition module and the metric module to judge the current assembly step and the changed area of the image respectively, and the changed area on the image is assigned according to the assembly step, so as to obtain a semantic change image. Different from other change detection systems, this network adds a step recognition module, which can identify the types of parts in the changed area. The following is a further specific introduction to each module:
[0044] (1) Feature extraction module:
[0045] The structure of the feature extraction module is as Figure 2 shown. The present invention innovatively designs a dense - connection feature fusion mechanism, which fuses feature information through close skip connections between the encoder and the decoder, can effectively reduce the loss of shallow - layer feature information of the neural network, maintain high - resolution and fine - granularity feature representation, and effectively alleviate problems such as poor processing of edge pixels in the detection results and missed detection of small targets. This module connects the node outputs in the shallow - layer sub - decoder to the nodes in the deep - layer sub - decoder. For example, after the first downsampling, the obtained and outputs are feature - concatenated to obtain the fused feature The fused feature is respectively connected to the upsampled , and Then, upsampling is performed again for feature fusion. Let represent the output of node , and the formula of is defined as follows:
[0046] (1)
[0047] where the function represents the convolutional block operation, the function represents the 2 2 max - pooling operation for downsampling, and the function represents the upsampling using transposed convolution. represents the connection in the channel dimension, aiming to fuse feature information. When , the encoder downsamples and extracts features; when When it is time, the dense skip connection mechanism starts to work, sequentially transmitting the fine-grained features in the encoder to the deep decoder, and finally outputting four sets of feature maps with the same size. This module can maintain the fine-grained feature representation, effectively alleviating problems such as poor processing of edge pixels in the detection results and missed detection of small targets.
[0048] (2) Attention module:
[0049] As shown in Figure 3 the figure, the present invention designs a self-attention Cot (Contextual Transformer) mechanism that fuses contextual feature information. Transformer is a deep learning self-attention neural network. The self-attention Cot mechanism combines the self-attention mechanism in Transformer with convolution operations to capture static and dynamic contextual information in the image.
[0050] The self-attention mechanism includes three key factors derived from the recommendation system: query, key, and value. Query and key are feature vectors for calculating weights, and value is a vector representing the input features. Its basic principle is: given a query, calculate the correlation between the query and the key, and then find the most suitable value according to the correlation between the query and the key.
[0051] The Cot mechanism integrates the mining of context and the learning of self-attention into a unified framework. It fully explores the neighboring context information to improve the learning of self-attention in an efficient way, thereby enhancing the expressive ability of the output features. In this structure, the encoding of the key uses convolution operations for encoding, so that the context information between neighbors can be obtained. Then, the global context information is obtained through two consecutive convolutions. Finally, the output result is obtained through the fusion of the neighboring context information and the global context information.
[0052] Compared with the traditional self-attention mechanism that only uses isolated query-key to calculate the attention matrix and fails to fully utilize the rich context feature information between keys, this module can make full use of the context feature information between adjacent positions in the input information to guide the learning of the dynamic attention matrix, thereby further enhancing the computer vision feature representation ability and then improving the monitoring performance of the network architecture. The self-attention Cot mechanism first performs context encoding on the input value through a 3 ×3 convolution to mine the static context feature information between adjacent keys, thereby generating the static context key key; then, according to the mutual relationship between the query and the static context key key, under the guidance of the static context key, two consecutive 1 1. Convolution is used to perform dynamic attention matrix learning; the learned attention matrix is used to aggregate all input values, thereby realizing the representation of dynamic context feature information; finally, the static context feature information and the dynamic context feature information are fused and output.
[0053] Assume the input information is a feature map ∈ R H×W×C , where is the height, is the width, is the number of channels. The self-attention Cot mechanism first uses k k group convolutions on adjacent keys of the feature map in the spatial dimension, performs weighted processing on the context association of each key, and obtains the context key 1 ∈ R H×W×C , 1 reflects the static context feature information between adjacent keys. Take 1 as the static context feature information of the input feature map . Then, using the context key 1 and the query concatenated as the condition, use two consecutive 1 1 convolutions to perform attention matrix learning. The attention matrix is defined as follows:
[0054] (2)
[0055] where, represents the convolution operation with the Relu activation function, while represents the convolution operation without an activation function. Finally, according to the attention matrix , calculate the attention feature map 2 by aggregating all values:
[0056] (3)
[0057] Given that the attention feature map 2 captures the dynamic interaction feature information between input information, define 2 as the dynamic context feature information. Finally, fuse and output the static context feature information 1 and the dynamic context feature information 2 .
[0058] (4)
[0059] The self-attention Cot mechanism can capture the above two kinds of spatial context feature information between the input keys simultaneously, that is, the static context feature information obtained by 3 ×3 convolution and the dynamic context feature information obtained based on context self-attention, thereby enhancing the visual representation ability.
[0060] (3) Step recognition module:
[0061] As shown in Figure 4 , the step recognition module is innovatively designed based on the binary change detection in the mechanical assembly process in the present invention. This module can detect the category of the changed target object, and then identify the current assembly stage of the changed component parts, so as to realize the monitoring of the assembly sequence. This module has a lightweight Mobile Vit network. MobileVit uses the Transformers method to process global feature information, that is, uses Transformers to extract image feature information as convolution. The step recognition module effectively encodes local information and global information in a tensor, combines the advantages of convolutional neural networks (such as having low sensitivity to spatial induction deviation and data augmentation) and Transformers (such as input adaptive weighting and global processing), which helps to learn better feature information with fewer parameters and simple training samples. Figure 4 in (convolution ) represents the standard convolution, MV2 refers to the MobileNetv2 network, and ↓2 means to perform downsampling processing.
[0062] (4) Metric module: The metric module can effectively automatically select and focus on more effective information between different groups through the extracted features to generate the image change region. This module first adds the four groups of feature maps output by the feature extraction module, then uses the self-attention Cot mechanism to weight the four groups of feature maps, and at the same time splices the four groups of feature maps and uses the self-attention Cot mechanism to weight and process again. The specific process is as follows:
[0063] (5)
[0064] (6)
[0065] (7)
[0066] (8)
[0067] Among them Denotes the concatenation of feature maps, function Denotes the operation of concatenating the feature map repeated n times along the channel dimension, Denotes element-wise multiplication, and finally a 1 1 convolution is used to obtain the changed region :
[0068] (9)
[0069] Among them Denotes a 1 1 convolutional layer generates The changed region ("a" is set to 2 here, representing change and non-change).
[0070] In addition, in image change detection, the unchanged sample data is often more than the changed sample data. To weaken the impact of the information imbalance of the changed sample data, the present invention adopts a hybrid loss function (weighted cross-entropy loss and the combination of losses) to optimize the network learning process, and the specific definition is as follows:
[0071] (10)
[0072] To describe the weighted cross-entropy loss , the changed region is regarded as a set of points, which is expressed as:
[0073] (11)
[0074] Among them represents a value in and represent the height and width of. The weighted cross-entropy loss is defined as:
[0075] (12)
[0076] where the value of a is 1 or 0, representing change and non-change, and at the same time the changed region participates in the calculation of loss:
[0077] (13)
[0078] Among them represents the true change label, and finally the changed region is assigned a value according to the target category obtained by the step recognition module to obtain the final semantic change image.
[0079] The specific process of using the above-mentioned modules to perform multi-view semantic change detection on a mechanical assembly includes: the dataset establishment stage, the training stage, and the testing stage. In the dataset establishment stage, a certain number of training samples are generated for the network to learn; in the training stage, the feature extraction module learns the assembly image feature information of the training samples, and after being processed by the attention module, the step recognition module, and the metric module, it outputs the semantic change image of the training samples, determines whether this semantic change image meets the training requirements, and finally saves the optimal model after multiple trainings; in the testing stage, the features of the newly input assembly image are directly extracted, and the semantic change image of the assembly process is obtained according to the optimal model saved in the training stage. The specific processes of the three stages are as follows:
[0080] Dataset establishment stage:
[0081] To establish a multi-view semantic change detection dataset for a mechanical assembly, first, a 3D model of the mechanical assembly is established through SolidWorks according to the dimensions of the assembly in the real scene. The assembly model is divided according to a certain assembly step. Then, the 3D model of each assembly step is imported into 3D Max software in turn to mark the color of each part. At the same time, the origin of the coordinate system is set and exported as an ive format file. Then, this file is imported and composite images are generated. Images are collected from different angles. Finally, the corresponding color labels in the images are extracted, and the color values in the color labels are reset as the feature of the change semantic label. The dataset of the present invention includes images of each assembly node from different perspectives and corresponding semantic change label images.
[0082] Training stage:
[0083] Reference Figure 4 , the specific training process of a multi-view semantic change detection method for assembly images for assembly sequence monitoring according to the present invention is as follows:
[0084] S1: Input the images T1 (reference image) at the previous moment and the images T2 (images to be detected) at the next moment from different perspectives into the feature extraction module respectively.
[0085] S2: The feature extraction module extracts the feature information of the above-mentioned dual-time images respectively. This module uses a dense connection and skip fusion mechanism to enhance the weight value of the shallow information of the fine-grained features, so that the network has rich feature information.
[0086] S3: The attention module performs weighted processing on the feature information of the above-mentioned dual-time images, makes full use of the context feature information between adjacent keys to guide the learning of the dynamic attention matrix, and thus further enhances the computer vision feature representation ability.
[0087] S4: Input the weighted feature information into the step recognition module and the metric module respectively. The step recognition module determines the current assembly stage, and the metric module obtains the changed area according to the feature information, and assigns the changed area according to the current assembly stage to obtain the semantic change image.
[0088] S5: Continuously iterate steps S1 to S4 using the training sample images in the dataset until the set number of training times is reached, and save the optimal model during the training process.
[0089] Testing stage:
[0090] During testing, input two new dual-time images during the assembly process from different perspectives, and directly output the semantic change image during the assembly process using the optimal model saved in the training stage.
[0091] To verify the effectiveness of the proposed method for multi-view semantic change detection of assembly images for assembly sequence monitoring, the existing change detection methods Das Net (Chen J, Yuan Z, Peng J, et al. DASNet: Dual attentive fully convolutional siamese networks for change detection in high-resolution satellite images[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2020, 14: 1194-1206), Change Star (Zheng Z, Ma A, Zhang L, et al. Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing Imagery[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 15193-15202.), Sscd Net (Sakurada K, Shibuya M, Wang W. Weakly supervised silhouette-based semantic scene change detection[C] / / 2020 IEEE International conference on robotics and automation (ICRA). IEEE, 2020: 6861-6867.) and SiamUnet (Fang S, Li K, Shao J, et al. SNUNet-CD: A densely connected siamese network for change detection of VHR images[J]. IEEE Geoscience and Remote Sensing Letters, 2021, 19: 1-5.) are compared with the network of the present invention. The semantic change detection dataset established in the above step S1 is used for the dataset, and the evaluation metrics are accuracy (Pr), recall (Re) and mean value (F1).The test results are shown in Table 1 as follows:
[0092] Table 1
[0093]
[0094] It can be seen from Table 1 that the F1 index of the method proposed in the present invention reaches 96.27%, and the detection performance is better than that of the comparative change detection method.
[0095] Advantages of the present invention:
[0096] (1) To realize the intelligent monitoring of the mechanical product assembly process, the present invention proposes a multi-view semantic change detection method for assembly images oriented to assembly sequence monitoring, designs a dense connection feature fusion mechanism and an attention mechanism for fusing context features, and enhances the computer vision feature representation ability.
[0097] (2) The present invention adds a step recognition module on the basis of the change detection system, which can not only detect the change area of the assembly image, but also identify the current assembly stage of the changed parts, and is applicable to mechanical assembly sequence monitoring.
[0098] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image multi-view semantic change detection system for assembly sequence monitoring, comprising: A feature extraction module, an attention module, and a metric module, characterized by further comprising: a step recognition module; The feature extraction module extracts dual-temporal image feature information from different perspectives of the input detection network respectively; The attention module performs weighted processing on the extracted dual-temporal image feature information, and the weighted dual-temporal image feature information is input into the step recognition module and the metric module respectively; The step recognition module detects the category of the changed target object, identifies the current assembly stage of the changed component parts, and monitors the assembly sequence; the step recognition module has a convolutional neural network that uses the Transformers method to process global feature information; The metric module judges the changed area of the image, assigns a value to the changed area according to the target category obtained by the step recognition module, so as to obtain a semantic change image; first, multiple groups of feature maps output by the feature extraction module are added, and then the self-attention Cot mechanism is used to perform weighted processing on the four groups of feature maps. At the same time, the four groups of feature maps are concatenated, and the self-attention Cot mechanism is used again to perform weighted processing on the concatenated feature maps to obtain the extracted features. The extracted features are used to automatically select and focus on more effective information between different groups to generate the image changed area. The specific process is as follows: Among them, represents the feature map concatenation, and the function represents the connection operation of repeating the feature map n times in the channel dimension. represents the element-wise product, and finally, a 1 ×1 convolution is used to obtain the change region : Among them represents a 1 ×1 convolutional layer, generating changing regions of .
2. The image multi-view semantic change detection system for assembly sequence monitoring according to claim 1, wherein The feature extraction module has a densely connected feature fusion mechanism. The feature extraction module connects the node outputs in the shallow sub-decoder to the nodes in the deep sub-decoder. When the feature fusion mechanism works, the fine-grained features in the encoder are sequentially transmitted to the deep decoder, and finally multiple groups of feature maps with the same size are output.
3. The image multi-view semantic change detection system for assembly sequence monitoring according to claim 1, characterized in that The attention module has a self-attention Cot mechanism that fuses context feature information. The steps of the self-attention Cot mechanism are as follows: First, context encoding is performed on the input value through a 3×3 convolution to mine the static context feature information between adjacent keys, thereby generating a static context key; Then, according to the mutual relationship between the query and the static context key, two consecutive 1×1 convolutions are used to perform dynamic attention matrix learning under the guidance of the static context key. The learned attention matrix is used to aggregate all input values, thereby realizing the representation of dynamic context feature information; Finally, the static context feature information and the dynamic context feature information are fused and output.
4. The detection method of the image multi-view semantic change detection system for assembly sequence monitoring according to any one of claims 1 to 3, characterized in that, Including the following stages: the dataset establishment stage, the training stage, and the testing stage; The dataset establishment stage generates training samples for the multi-perspective semantic change detection network for assembly sequence monitoring to learn; In the training stage, the feature extraction module learns the assembly image feature information of the training samples, and after being processed by the attention module, the step recognition module, and the metric module, outputs the semantic change image of the training samples, judges whether this semantic change image meets the training requirements, and finally saves the optimal model after multiple trainings; In the testing stage, the feature extraction module extracts features from the newly input assembly image and obtains the semantic change image according to the optimal model.
5. The method for detecting semantic changes in multiple perspectives of an image for assembly sequence monitoring according to claim 4, wherein The steps of the dataset establishment stage are as follows: First, establish a 3D model of the assembly that is consistent with the dimensions of the assembly in the mechanical and real scenarios. Divide the assembly model into 3D models of multiple assembly steps. Then, import the 3D models of each assembly step in sequence and color-label each part. At the same time, set the origin of the coordinate system and export it as a file in a set format. Then, import this file and generate a composite image, collect images from different angles, and finally extract the corresponding color tags in the images. Modify the color values in the color tags as the changing semantic features.
6. The method for multi-view semantic change detection of images for assembly sequence monitoring according to claim 4, characterized in that The steps in the training stage are as follows: S1: Input the image at the previous moment from different perspectives as the reference image T1 and the image at the subsequent moment as the image to be detected T2 into the feature extraction module respectively. S2: The feature extraction module extracts the feature information of the above-mentioned dual-time images respectively. This module uses a dense connection and jump fusion mechanism to enhance the weight value of the shallow information of fine-grained features, so that the network has rich feature information. S3: The attention module performs weighted processing on the feature information of the above-mentioned dual-time images, makes full use of the context feature information between adjacent keys to guide the learning of the dynamic attention matrix, and thus further enhances the computer vision feature representation ability. S4: Input the weighted feature information into the step recognition module and the metric module respectively. The step recognition module judges the current assembly stage, and the metric module obtains the changed area according to the feature information, and assigns a value to the changed area according to the current assembly stage to obtain a semantic change image. S5: Continuously iterate steps S1 to S4 using the training sample images in the dataset until the set number of training times is reached, and save the optimal model during the training process.
Citation Information
Patent Citations
Assembly change detection method and device based on attention mechanism and medium
CN113269237A
Workpiece defect detection method and device fusing multi-attention mechanism
CN113822885A