An assembly sequence planning method

By using an improved masked region convolutional neural network and ontology theory, the problem of high computational complexity in assembly sequence planning is solved, achieving efficient and accurate assembly sequence generation and supporting automated assembly sequence planning.

CN120278473BActive Publication Date: 2025-10-24LIAONING UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-10-24
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Existing assembly sequence planning methods have high computational complexity in complex product assembly and are difficult to effectively reduce assembly time and cost.

Method used

An improved masked region convolutional neural network is used for image segmentation of assembly parts, and assembly rule reasoning is performed in combination with ontology theory to generate assembly sequence planning information.

Benefits of technology

It improves the accuracy and efficiency of assembly sequence planning, reduces computational complexity, ensures clear assembly relationships between parts, and supports automated assembly sequence planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278473B_ABST
    Figure CN120278473B_ABST
Patent Text Reader

Abstract

The application discloses an assembly sequence planning method and relates to the technical field of computers. The method acquires an initial mask region convolutional neural network; adopts a UNet3+ network structure to improve a full convolutional network in a mask branch, inserts a convolution block attention module into each bottleneck residual block layer of a residual network 50, and obtains the mask region convolutional neural network; trains the mask region convolutional neural network, and obtains a target mask region convolutional neural network; extracts features of multiple part images of a target assembly body to be assembled through the target mask region convolutional neural network, and obtains two-dimensional mask images of each part image; acquires an assembly rule of the target assembly body; constructs an assembly information ontology, and performs instantiation; according to the instantiated assembly information ontology, executes the assembly rule, and performs reasoning to obtain assembly sequence planning information. The method can reduce the calculation complexity of assembly sequence planning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to an assembly sequence planning method. BACKGROUND

[0002] Assembly sequence planning usually plays an important role in the manufacturing of complex products, such as automobiles, ships, shield machines and other large engineering machinery. According to research statistics, more than 40% of the cost of product manufacturing is generated by product assembly, and the assembly time accounts for about 20%-50% of the total production time. Assembly sequence planning has a direct impact on the quality, cost and feasibility of assembly process. Intelligent assembly sequence planning is an important technology to realize intelligent assembly of complex products and enhance the market competitiveness of enterprises.

[0003] Most of the research on intelligent assembly sequence planning focuses on assembly sequence planning (ASP). In the prior art, assembly sequence planning is usually performed by a knowledge-based method, a cut-set-based method, a multi-color set-based method or a heuristic algorithm-based method. However, the knowledge-based method is difficult to acquire and express knowledge, and requires process planning personnel to have a high level, so it is difficult to ensure that the matching between the actual assembly model and the standard knowledge base does not have errors in application; the cut-set-based method, the multi-color set-based method and the heuristic algorithm-based method all belong to pure mathematical methods, and as the number of parts increases, the solution space of these methods increases sharply, resulting in high computational complexity of assembly sequence planning. SUMMARY

[0004] Therefore, it is necessary to provide an assembly sequence planning method to solve the above technical problems. The method can reduce the computational complexity of assembly sequence planning.

[0005] The present application adopts the following technical solutions:

[0006] The present application provides an assembly sequence planning method, comprising:

[0007] An initial mask region convolutional neural network is obtained; the initial mask region convolutional neural network comprises a backbone network and a mask branch; the backbone network is constructed by a residual network 50 and a pyramid network module;

[0008] The full convolutional network in the mask branch is improved using a UNet3+ network structure, and a convolution block attention module is inserted after each bottleneck residual block layer in the residual network 50 to obtain a mask region convolutional neural network;

[0009] An assembly body part image dataset is obtained, and the assembly body part image dataset is preprocessed to obtain a preprocessed assembly body part image dataset;

[0010] train the mask region convolutional neural network through the preprocessed assembly part image data set, to obtain a target mask region convolutional neural network;

[0011] obtain a plurality of part images of a target assembly to be assembled, and perform feature extraction on the plurality of part images of the target assembly through the target mask region convolutional neural network to obtain a two-dimensional mask image of each part image; each pixel point in the two-dimensional mask image represents a part type;

[0012] obtain an assembly rule of the target assembly; the assembly rule is determined according to quality factors, precision factors, size factors, connection relationship factors, position factors and a preset part assembly sequence relationship of each part;

[0013] construct an assembly information ontology, and instantiate the assembly information ontology through the two-dimensional mask images of all part images; the instantiated assembly information ontology includes object attributes of parts and data attributes of parts; the object attributes represent connection relationships between parts in the assembly; and the data attributes represent quality, size, precision and position relationships of the parts;

[0014] According to the instantiated assembly information ontology, the assembly rule is executed, and reasoning is performed to obtain assembly sequence planning information.

[0015] Preferably, the mask region convolutional neural network comprises a region proposal network, an anchor box mapping module, an interest region alignment module and a fixed feature map size module; the two-dimensional mask image of each part image is obtained by performing feature extraction on the plurality of part images of the target assembly through the target mask region convolutional neural network, and specifically includes:

[0016] For any part image of the target assembly, the part image is input into a backbone network, and is processed through a residual network 50, a convolution block attention module and a pyramid network module connected in sequence to obtain a processing result of the backbone network;

[0017] The processing result of the backbone network is input into the region proposal network for processing to obtain an output of the region proposal network;

[0018] The output of the region proposal network and the processing result of the backbone network are input into the anchor box mapping module to obtain an output of the anchor box mapping module;

[0019] The output of the anchor box mapping module is input into the interest region alignment module to obtain a processing result of the interest region alignment module;

[0020] The processing result of the interest region alignment module is sampled to a preset size through the fixed feature map size module to obtain a sampling result;

[0021] The sampling result is input into a UNet3+ network structure for up-sampling and down-sampling operations to obtain a mask feature map; the mask feature map has a preset size and multiple channels; each channel of the mask feature map represents different feature information.

[0022] The mask feature map is classified to generate a two-dimensional mask image of the part image.

[0023] Preferably, the data attributes include weight levels, precision levels, size, position relationships, start, end, assembly relationship quantities.

[0024] Preferably, the object attributes include parts, assemblies, assembled, mounted on, connected through, assembled after, start installation, end installation, and connection.

[0025] Preferably, the assembly part image dataset is preprocessed to obtain a preprocessed assembly part image dataset, specifically including:

[0026] Each part image in the assembly part image dataset is subjected to diversity processing; the diversity processing includes angle transformation, texture change, and size adjustment.

[0027] The part image subjected to the diversity processing is subjected to sharpening processing by Laplacian sharpening.

[0028] The part image subjected to the sharpening processing is subjected to median filtering to obtain a preprocessed image.

[0029] All the preprocessed images are combined into a preprocessed assembly part image dataset.

[0030] The application provides an assembly sequence planning device, comprising:

[0031] A first acquisition module is configured to acquire an initial mask region convolutional neural network; the initial mask region convolutional neural network comprises a backbone network and a mask branch; the backbone network is constructed by a residual network 50 and a pyramid network module.

[0032] An improvement module is configured to improve a full convolutional network in the mask branch by using a UNet3+ network structure, and insert a convolution block attention module into each bottleneck residual block layer of the residual network 50 to obtain a mask region convolutional neural network.

[0033] A second acquisition module is configured to acquire an assembly part image dataset and pre-process the assembly part image dataset to obtain a preprocessed assembly part image dataset.

[0034] The training module is configured to train the mask region convolutional neural network through the preprocessed assembly part image dataset, and obtain a target mask region convolutional neural network;

[0035] The extraction module is configured to acquire a plurality of part images of a target assembly to be assembled, and perform feature extraction on the plurality of part images of the target assembly by using the target mask region convolutional neural network, so as to obtain a two-dimensional mask image of each part image; each pixel point in the two-dimensional mask image represents a part type;

[0036] The third acquisition module is configured to acquire an assembly rule of the target assembly; the assembly rule is determined according to a quality factor, a precision factor, a size factor, a connection relationship factor, a position factor and a preset part assembly sequence relationship of each part;

[0037] The construction module is configured to construct an assembly information ontology, and instantiate the assembly information ontology by using the two-dimensional mask images of all the parts; the instantiated assembly information ontology includes an object attribute of the part and a data attribute of the part; the object attribute represents a connection relationship between the parts in the assembly; and the data attribute represents a quality, a size, a precision and a position relationship of the part;

[0038] The reasoning module is configured to execute the assembly rule according to the instantiated assembly information ontology, and perform reasoning to obtain assembly sequence planning information.

[0039] The application provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the assembly sequence planning method.

[0040] The application provides a computer device, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to implement the assembly sequence planning method.

[0041] The above at least one technical solution adopted by the application can achieve the following beneficial effects:

[0042] The initial mask region convolutional neural network is obtained, and the initial mask region convolutional neural network is improved to obtain the mask region convolutional neural network, so that the accuracy of segmentation is higher; the assembly part image dataset is obtained, and the assembly part image dataset is preprocessed to obtain the preprocessed assembly part image dataset, the preprocessing can effectively remove noise, enhance image contrast, unify image size and format, improve the diversity of parts, ensure the richness of the part dataset, enhance the feature information of the parts, and ensure the segmentation effect of the mask region convolutional neural network; the mask region convolutional neural network is trained through the preprocessed assembly part image dataset to obtain a target mask region convolutional neural network; a plurality of part images of a target assembly body to be assembled are obtained, and the target mask region convolutional neural network is used for feature extraction on the plurality of part images of the target assembly body to obtain a two-dimensional mask image of each part image; the feature extraction improves the segmentation accuracy of the part features, and the parts are accurately segmented from the product to obtain detailed part segmentation information; an assembly rule of the target assembly body is obtained, and the assembly rule provides clear part-level semantic description for generation of the assembly sequence; an assembly information ontology is constructed, and the assembly information ontology is instantiated through the two-dimensional mask image; according to the instantiated assembly information ontology, the assembly rule is executed, and reasoning is performed to obtain assembly sequence planning information, which can systematically organize assembly knowledge, clearly define the assembly relationship and constraint conditions between parts, and reason out the assembly sequence of the product by using assembly prior knowledge, thereby providing decision support for automatic assembly sequence planning. The method can reduce the computational complexity of the assembly sequence planning. BRIEF DESCRIPTION OF DRAWINGS

[0043] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application, and do not limit the application in any way. In the drawings:

[0044] Figure 1 A flowchart of an assembly sequence planning method provided by the application is shown in the figure;

[0045] Figure 2 A CBAM attention mechanism module embedding and image size change diagram provided by the application is shown in the figure;

[0046] Figure 3 A UNet3+ network structure diagram provided by the application is shown in the figure;

[0047] Figure 4 An assembly body and part three-dimensional model dataset display diagram provided by the application is shown in the figure;

[0048] Figure 5 A data enhancement and noise processing diagram provided by the application is shown in the figure;

[0049] Figure 6 The target mask area convolutional neural network model structure improvement schematic diagram provided by the present application is shown in the following figure:

[0050] Figure 7 The hierarchical structure schematic diagram of the assembly provided by the present application is shown in the following figure:

[0051] Figure 8 The ontology information model attribute definition schematic diagram provided by the present application is shown in the following figure:

[0052] Figure 9 The fine labeling schematic diagram of the reducer parts provided by the present application is shown in the following figure:

[0053] Figure 10 The mAP50 and overall loss function change schematic diagram provided by the present application is shown in the following figure:

[0054] Figure 11 The prediction results and model evaluation index schematic diagram of the three network models provided by the present application are shown in the following figure:

[0055] Figure 12 The example node schematic diagram of the assembly provided by the present application is shown in the following figure:

[0056] Figure 13 The exploded view of the segmented reducer and parts provided by the present application is shown in the following figure:

[0057] Figure 14 The ontology instance and rule reasoning result schematic diagram provided by the present application is shown in the following figure:

[0058] Figure 15 The feasible assembly sequence scheme schematic diagram of the rule regulation reasoning provided by the present application is shown in the following figure:

[0059] Figure 16 The assembly sequence planning method framework provided by the present application is shown in the following figure:

[0060] Figure 17 The assembly sequence planning device schematic diagram provided by the present application is shown in the following figure:

[0061] Figure 18 The computer equipment schematic diagram for realizing the assembly sequence planning method provided by the present application is shown in the following figure. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described clearly and completely below in combination with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0063] A device such as a desktop, a server, a notebook, etc. capable of executing the solution of the present application. For the convenience of explanation, the following will be explained taking a server as the execution subject.

[0064] Researchers apply genetic algorithm to assembly sequence planning problem, propose to use fitness function to construct feasible solution space, then use genetic algorithm to search in feasible solution space, for task sequence planning affecting complex assembly system efficiency and stability, researchers propose an adaptive quantum genetic algorithm based on artificial potential field and objective function gradient. Researchers balance assembly sequence and assembly line, calculate the sum of assembly task time, change assembly direction time and change tool time, then propose the optimal assembly sequence to be used when the given generation cycle is adopted. Researchers propose to introduce a uniform factor in particle swarm optimization algorithm to balance the influence of cognitive condition, efficiency is obviously higher than particle swarm algorithm, has better solving effectiveness and consistency. The above researches face some difficulties:

[0065] (1) The number of assembly parts is increasing, the number of questions and answers will rise sharply, which is very difficult for manual operation.

[0066] (2) Evolutionary algorithm consumes a lot of computing resources and time when dealing with large-scale problems, has slow convergence speed, difficult parameter selection, is easily affected by premature problem and depends on problem characteristics.

[0067] Deep learning is a powerful feature extraction technique that uses deep neural networks to automatically learn complex feature representations of data. Deep learning has achieved great success in computer vision and other fields. The introduction of the Fully Convolutional Networks (FCN) architecture has made semantic segmentation more effective in utilizing deep learning techniques. To address the problem of coarse semantic segmentation maps with FCN, researchers have proposed integrating context information into the fully convolutional neural network to improve semantic segmentation performance. Researchers have proposed a semantic image segmentation model based on deep convolutional networks and fully connected conditional random fields (Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs, DeepLabv1) by combining deep convolutional neural networks with fully connected conditional random fields. Based on DeepLabv1, researchers have proposed a semantic image segmentation model based on deep convolutional networks, atrous convolution, and fully connected conditional random fields (Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, DeepLabv2). Researchers have proposed a new loss function that uses information theory divergence measures and geometric regularization constraints to train deep clustering networks, achieving competitive performance on synthetic benchmarks and real-world datasets. Not only has the overall semantic segmentation accuracy been improved, but the performance of small object segmentation has also been improved. Researchers have proposed a more structured encoder-decoder architecture called the U-Net network. The U-Net network was proposed to address the problem of medical image segmentation and can achieve good segmentation performance using only a small amount of image data. It is a lightweight network. Researchers have proposed a conceptually simple, flexible, and general mask region-based convolutional neural network (Mask Region-based Convolutional Neural Network, Mask R-cnn) that is an extension of the Faster Region-based Convolutional Neural Networks (Faster Region-based Convolutional Neural Networks, Faster R-CNN). It adds a branch for predicting object masks, enabling instance segmentation.

[0068] Deep learning-based semantic segmentation methods have been successfully applied to various fields, but there are several problems in the field of mechanical assemblies:

[0069] (1) Data set is the basis of the whole deep learning research, but there is a lack of public data set in the field of mechanical assembly, which makes the related research less.

[0070] (2) Although the research of semantic segmentation based on deep learning has achieved a lot of results, due to the complex structure of mechanical assembly, the lack of color information of mechanical parts and the single texture, etc., these research results are rarely successfully applied to the field of semantic segmentation of mechanical assembly.

[0071] Ontology unifies different knowledge from different sources, realizes knowledge integration and sharing, and ontology can describe concepts of different granularity, support logical reasoning. Therefore, ontology has been applied to many fields, and there have been some related researches in the field of design and manufacturing that use ontology technology to realize engineering semantic representation, knowledge reasoning and knowledge sharing, etc. Researchers designed a product information modeling framework for product lifecycle management based on ontology method. In order to overcome the problems of discrete assembly process, diverse assembly resources, and complex data in the process of assembly task execution, researchers proposed a modeling method based on ontology for twin workshop, constructed the ontology model of assembly resources and assembly process, and through ontology instantiation, the assembly resource instance and assembly process instance can dynamically participate in the modeling of twin workshop. Researchers developed an ontology-based knowledge decision support system, integrated the information extracted from the Computer Aided Design (CAD) model into the ontology, and used a rule-based reasoning system to represent and reason the new knowledge of variant products. Researchers proposed a knowledge support product design system for the aerospace industry, which uses ontology-based method for semantic knowledge management. In order to strengthen the management and reuse of historical process design knowledge, researchers carried out research on modeling of design knowledge, proposed a fuzzy interval description logic theory based on ontology, and developed an ontology-based design knowledge expression model. Researchers constructed an assembly design ontology and an ontology-based assembly design framework to realize collaborative product development. Researchers proposed a product service configuration method based on ontology modeling. Researchers proposed an automatic decision-making method based on ontology and Case-Based Reasoning (CBR) for disassembling mechanical products. Researchers used ontology method for tolerance synthesis and main model similarity calculation, etc.

[0072] In view of the above assembly method research and the shortcomings of image segmentation network and the advantages of ontology technology, this research proposes an improved Mask R-CNN instance segmentation algorithm to accurately segment the two-dimensional image of the assembled product and obtain the product part information, and then reason out the assembly sequence of the assembled product through ontology modeling.

[0073] For the application of assembly sequence planning of complex engineering, the acquisition of geometric information is limited by many factors. Fortunately, based on deep learning, the semantic segmentation can identify the part category by semantic segmentation of the assembly body, which is similar to the process of observing the CAD model of complex products by process engineers. Based on the existing rich knowledge and experience and the reuse concept, the experience and knowledge learned in the product assembly process have the potential for reuse. Combined with the advantages of ontology theory in semantic expression, intelligent reasoning and knowledge reuse, ontology theory is introduced into the intelligent reasoning of answer set programming (ASP). The ontology model for assembly sequence planning integrates various types of information through a unified and structured format, can realize concept description of different levels of detail, and support logical reasoning function. In order to ensure that the relevant information of assembly sequence planning can be accurately understood by computer and effectively reduce the uncertainty in the process of ASP, the deep learning and ontology theory in the field of artificial intelligence are introduced into ASP at the same time, and an assembly sequence planning method is proposed.

[0074] The method proposed in the application uses semantic segmentation to segment the assembly body parts in the image to obtain assembly part information, uses ontology theory to establish a perfect assembly information model, writes semantic rules to complete reasoning on the assembly sequence, and obtains an efficient and feasible assembly sequence. Through semantic segmentation, the part category can be effectively identified, and the experience and knowledge reuse of assembly sequence planning can be used to more easily and accurately generate a feasible assembly sequence.

[0075] The technical solutions provided by the embodiments of the application are described in detail below with reference to the drawings.

[0076] Figure 1 The flowchart of the assembly sequence planning method in the application specifically includes the following steps:

[0077] S101: An initial mask region convolutional neural network is obtained; the initial mask region convolutional neural network includes a backbone network and a mask branch; the backbone network is constructed through a residual network 50 and a pyramid network module.

[0078] Specifically, the basic structure of the initial mask region convolutional neural network includes a backbone network (Backbone Network, Backbone), a pyramid network module (Feature Pyramid Network, FPN), a region proposal network (Region Proposal Network, RPN), a region of interest alignment module (Region of Interest Align, ROIAlign), and a prediction head including a bounding box regression, a class classification, and a mask generation. Through the mask branch, the initial MaskR-CNN can extract specific shape and position information of each part target in the assembly model.

[0079] S102: The full convolutional network in the mask branch is improved by using a UNet3+ network structure, and a convolution block attention module is inserted after each bottleneck residual block layer in the residual network 50, to obtain a mask region convolutional neural network.

[0080] Specifically, in order to improve the performance of the initial mask region convolutional neural network in the part instance segmentation algorithm in the assembly model in terms of feature extraction capability, segmentation accuracy, and computing performance, the initial Mask R-CNN is improved to obtain the mask region convolutional neural network. The full convolutional network in the mask branch is improved by using a UNet3+ network structure. The feature extraction network has strong ability to accurately extract rich feature information such as edges, textures, and shapes from the assembly CAD image. In order to improve the accuracy and efficiency of feature extraction in the initial mask region convolutional neural network, the feature pyramid network (Feature Pyramid Networks, FPN) is introduced, which realizes efficient extraction of assembly image features by constructing multi-scale feature maps.

[0081] The FPN realizes efficient extraction of assembly image features by constructing multi-scale feature maps. When constructing the network model, the Resnet50+FPN structure is selected as the backbone network, which not only has strong feature extraction capability, but also guarantees the stability and robustness of the model. In the training process, the pre-trained model is used, and only the last five layers are fine-tuned, while the weight parameters of other layers remain unchanged. Such a training strategy not only improves the training efficiency of the model, but also guarantees the performance of the model. Finally, the mask region convolutional neural network successfully outputs feature maps with a size of 28x28 and a channel number of 256, which provide strong support for the subsequent part instance segmentation task.

[0082] Specifically, using attention mechanism can help the model focus better on the local area of the input, thereby improving the performance of feature extraction. Therefore, the present application introduces a Convolutional Block Attention Module (CBAM), that is, CBAM, which combines channel attention and spatial attention to capture the correlation between features in different dimensions and highlight the most important parts of the feature map, further improving the learning ability of the model for target features. As shown in Figure 2 The CBAM attention mechanism module is inserted after each BottleNeck layer of the Resnet50. Figure 2 The CBAM attention mechanism module provided by the present application is embedded and the image size change schematic diagram is shown in Figure 2 The input data is an image with a size of 448x448, which is processed by the modules connected in sequence as shown in Figure 2 The size of the feature map is 14x14 as shown in Figure 2

[0083] The backbone network extraction assembly and the part image feature information are transmitted to the region proposal network, three kinds of anchors (128, 256, 512) and three kinds of horizontal and vertical ratios (0.5, 1.0, 2.0) are defined, nine different anchor points are generated to determine the position and shape of the candidate box, and then the candidate box is extracted through the RPN network layer to obtain the accurate position information of the segmentation target. ROI Align traverses each candidate region, keeps the floating-point boundary from being quantized, and greatly improves the accuracy of segmentation. Finally, a feature map with a size of 14x14 is output.

[0084] In order to improve the segmentation quality of the assembly parts, the feature extraction is refined and the deep and shallow feature fusion is considered. The improved UNet3+ is used instead of FCN as the head of the initial mask region convolutional neural network mask branch. Full-scale Skip Connections (FSC) are introduced, which combines low-level semantics and high-level semantics from full-scale feature maps and has fewer parameters. Full-scale skip connections are introduced, which combine low-level and high-level semantics from all scale feature maps and the network architecture of UNet3+, Figure 3 The UNet3+ network structure provided by the present application is shown in Figure 3 The UNet3+ is composed of an encoder (left) and a decoder (right) as shown in the figure, in the encoder, each layer is responsible for capturing and refining feature information of the image at different scales. These feature maps not only contain rich low-level details, but also integrate high-level semantic understanding. Each layer of the encoder ​Full-scale skip connection represented by a dashed line and the corresponding layer of the decoder are connected. The network is allowed to utilize the high-resolution features reserved in the encoder during the decoding process, thereby improving the efficiency and accuracy of feature fusion. For example, the feature maps generated by the encoder 、 、 、 and and the feature maps generated by the decoder 4 are connected using a dashed line. The calculation formula of the corresponding layer of the decoder is shown in equation (1):

[0085] (1)

[0086] wherein, is the corresponding feature map of the decoder, represents the down-sampling layer in the encoding direction, represents the number of encoders, (.) represents the convolution operation, and the function (.) represents the feature aggregation mechanism realized by convolution, batch normalization, and ReLu activation function, and the function (.) down-sampling operation, and the function (.) represents the up-sampling operation, represents the concatenation and fusion of the channel dimension, represents the feature map of each layer of the encoder, k is a constant determined according to i , ranging from [1, N], represent the feature map of the k encoder and the k decoder, respectively, represents the scale of the first feature map to the feature map, represents the scale of the feature map to the feature map.

[0087] In the feature pyramid network, the loss function is usually defined as the sum of the classification loss and the bounding box regression loss, and the loss function is shown in equation (2):

[0088] (2)

[0089] wherein, L is the loss, 、 represent the classification loss and the bounding box regression loss, respectively.

[0090] The loss function of the masked region convolutional neural network is defined as the sum of classification loss, bounding box regression loss, and segmentation loss. 、 Represent classification loss and bounding box regression loss respectively, The segmentation loss corresponding to the mask segmentation module uses the multi-class cross-entropy loss (MCEL). This loss function aims to quantify the inconsistency between the probability distribution predicted by the model and the true probability distribution, thereby measuring the performance of the classification task.

[0091] The loss function of Mask R-CNN is calculated as shown in formula (3):

[0092] (3);

[0093] in, is the segmentation loss of the mask segmentation module, 、 represent the classification loss and bounding box regression loss respectively.

[0094] The classification loss is calculated as shown in formula (4):

[0095] (4);

[0096] in, is the classification loss, i As anchor point, Anchor The true label (0 or 1), Anchor The predicted probability of .

[0097] The bounding box regression loss adopts the smooth L1 loss function, which introduces a smooth term on the basis of L1 loss, thereby enhancing the stability of model training. The calculation method of the bounding box regression loss is shown in formula (5):

[0098] (5);

[0099] in, is the bounding box regression loss, is the number of anchor points involved in regression, is the coordinate transformation parameter of the predicted value, is the coordinate transformation parameter of the real value, is the smoothing function L1, Anchor The true label.

[0100] In the Mask R-CNN network, the loss function is used as an index to measure the performance of each branch of the network model. The loss function of the classification branch aims to quantify the accuracy of the model's prediction of target categories, and a multi-class cross-entropy loss function is usually used for calculation. The mask segmentation module uses the average binary cross-entropy loss function. This function is used to evaluate the model's binary classification prediction (foreground and background) for each pixel.

[0101] The calculation method of the segmentation loss of the mask segmentation module is shown in formula (6):

[0102] (6);

[0103] wherein, is the segmentation loss of the mask segmentation module, is the spatial dimension of the mask output, is the prediction result at the position, is the true label at the position, is the position coordinate.

[0104] By minimizing the loss function, the mask region convolutional neural network can continuously improve its performance in target detection and instance segmentation tasks.

[0105] S103: Obtain an assembly part image dataset, and pre-process the assembly part image dataset to obtain a pre-processed assembly part image dataset.

[0106] In an example embodiment, the assembly part image dataset is pre-processed to obtain a pre-processed assembly part image dataset, specifically including: performing diversity processing on each part image in the assembly part image dataset; the diversity processing includes angle transformation, texture change, and size adjustment; performing sharpening processing on the part image after diversity processing by Laplacian sharpening; performing median filtering on the sharpened part image to obtain a pre-processed image; and combining all pre-processed images into a pre-processed assembly part image dataset.

[0107] Specifically, due to the particularity of the industrial field, there are fewer related datasets available for reference and use, so the assembly part image dataset is constructed. By transforming the angle, changing the texture, and adjusting the size of the part library three-dimensional model, the diversity of the image is increased, and the generalization ability of the improved instance segmentation model for complex assembly situations is improved, such as Figure 4The three-dimensional model data set display diagram of the assembly and parts provided by the application is shown, which ensures that the part and assembly model image sources are diverse and high in clarity, involving different angles, lighting conditions and background environments of the parts. Since mechanical products are usually silver-gray, the pixel values of the part model image are linearly converted from the original range (usually 0-255) to the standardized range (such as 0-1) to eliminate the brightness difference between different images and improve the model training efficiency and accuracy.

[0108] In order to improve the richness of the assembly part image data set and improve the generalization ability of the improved instance segmentation algorithm to the actual assembly scene, different data enhancement methods are used to improve the sample size of the data set. First, the image is processed by Laplacian sharpening, and the image before processing is shown in Figure a of Figure 5 The second derivative of the part image is calculated using the Laplace operator to detect the details and textures in the part image, and the clarity and contrast of the edge details of the part image are enhanced, as shown in Figure b of Figure 5 .

[0109] The Laplace operator is a second-order differential operator, which represents the second-order derivative operation on the pixel value in the and directions of the image. The calculation method of the Laplace operator is shown in formula (7):

[0110] (7);

[0111] wherein, is the Laplace operator, is a scalar function, is a scalar function f The second-order partial derivative of the coordinate x , represents the rate of change of the curvature or rate of change of the function f in the x direction, is the function f The second-order partial derivative of the coordinate y , represents the rate of change of the curvature or rate of change of the function f in the y direction, x and y are spatial coordinates used to describe the position of the scalar function f in space.

[0112] The image after edge detection is shown in formula (8):

[0113] (8);

[0114] wherein, is the image after edge detection, is the original image, i.e. the image without edge detection processing, is the coordinate of a pixel in the image, is a constant used to adjust the degree of influence of the Laplacian operator on the original image, is the original image The Laplacian transform of the original image

[0115] In the assembly image instance segmentation process, the presence of noise may cause the performance of the instance segmentation algorithm to decline. Some isolated noise is removed by median filtering while preserving most of the edge information of the part image. Figure 5 The c figure in the image is the image with added salt and pepper noise. By median filtering to remove noise, the enhanced contour information of the Laplacian operator is effectively preserved, Figure 5 The d figure in the image is the image processed using median filtering. For a two-dimensional image, median filtering is shown in equation (9):

[0116] (9);

[0117] wherein, represents the image processed by median filtering, is a window, usually set to 3x3, 5x5, is the original image the pixel value at the coordinate , represents the median operation, that is, the median value is found from all pixel values in the window W and is taken as the pixel value of the image after filtering at the coordinate .

[0118] S104: Train the mask region convolutional neural network using the preprocessed assembly part image dataset to obtain a target mask region convolutional neural network.

[0119] Train the mask region convolutional neural network using the preprocessed assembly part image dataset to obtain a target mask region convolutional neural network.

[0120] S105: Obtain a plurality of part images of a target assembly to be assembled, and perform feature extraction on the plurality of part images of the target assembly by the target mask region convolutional neural network to obtain a two-dimensional mask image for each part image; each pixel point in the two-dimensional mask image represents a part type.

[0121] In an exemplary embodiment, the mask region convolutional neural network includes a region proposal network, an anchor box mapping module, an interest region alignment module, and a fixed feature map size module; a target mask region convolutional neural network is used to extract features of multiple part images of a target assembly to obtain a two-dimensional mask image of each part image, specifically including: for any part image of the target assembly, the part image is input into the backbone network, and processed by the residual network 50, the convolution block attention module and the pyramid network module connected in series to obtain the processing result of the backbone network; the processing result of the backbone network is input into the region proposal network for processing to obtain the output of the region proposal network; the region proposal network is used to extract features of multiple part images of a target assembly to obtain a two-dimensional mask image of each part image. The output of the network and the processing result of the backbone network are input into the anchor box mapping module to obtain the output of the anchor box mapping module; the output of the anchor box mapping module is input into the region of interest alignment module to obtain the processing result of the region of interest alignment module; the processing result of the region of interest alignment module is sampled to a preset size through the fixed feature map size module to obtain a sampling result; the sampling result is input into the UNet3+ network structure for up and down sampling operations to obtain a mask feature map; the size of the mask feature map is a preset size and has multiple channels; each channel of the mask feature map represents different feature information; the mask feature map is classified to generate a two-dimensional mask image of the part image.

[0122] Specifically, the preset size is set according to specific engineering practices. For example, in the present invention, the preset size is 14×14 and the number of channels is 256.

[0123] Specifically, if Figure 6 The figure shows the improved structure of the target mask region convolutional neural network provided by the present invention. The workflow of the target mask region convolutional neural network model is as follows: First, the feature map output by ROI Align is upsampled to a size of 14×14 by bilinear interpolation; then, it is input into the UNet3+ network for upsampling and downsampling operations to extract more feature details; secondly, after being processed by the UNet3+ network, the obtained mask feature map is 14×14 in size and has 256 channels, which represent different feature information; finally, these feature maps are classified to generate the final two-dimensional mask image, in which each pixel in the two-dimensional mask image represents the probability that the position belongs to a certain part category. For example, Figure 6 The red pixels in the image represent the end caps, and the green pixels represent the gears.

[0124] Specifically, the feature extraction network, with its powerful capabilities, can accurately extract rich feature information such as edges, textures, and shapes from assembly CAD images. To further improve the accuracy and efficiency of feature extraction in the Mask R-CNN model, this paper introduces a pyramid network (FPN) module. FPN constructs multi-scale feature maps to efficiently extract features from assembly images. When constructing the network model, the present invention selected a Resnet50 + FPN structure as the backbone network. This combination not only has powerful feature extraction capabilities but also ensures model stability and robustness. During training, the present invention utilized a pre-trained model and fine-tuned only the last five layers, while keeping the weight parameters of the remaining layers unchanged. This training strategy not only improves model training efficiency but also ensures model performance. Ultimately, the network successfully outputs feature maps with a size of 28×28 and 256 channels, which provide strong support for subsequent part instance segmentation tasks.

[0125] Specifically, using an attention mechanism can help the model better focus on local regions of the input, thereby improving feature extraction performance. This paper introduces a Convolutional Block Attention Module (CBAM), also known as a CBAM attention mechanism module, which combines channel attention and spatial attention to capture the correlation between features in different dimensions, highlighting the most important parts of the feature map, and further improving the model's ability to learn target features.

[0126] S106: Acquire assembly rules of the target assembly; the assembly rules are determined based on the quality factors, precision factors, size factors, connection relationship factors, position factors and the preset assembly sequence relationship of each part.

[0127] Assembly knowledge representation is the premise and foundation of assembly sequence planning, and assembly features are the carriers of assembly-related information about parts and components. Among domain knowledge description methods, ontology models offer numerous advantages in expressing conceptual hierarchies and semantics, sharing and reusing knowledge, and reasoning about knowledge. Therefore, to meet the needs of assembly sequence planning, basic classes, attributes, and declarations are defined to describe assembly knowledge and enable reasoning about assembly knowledge.

[0128] In the field of assembly sequence planning, the rich engineering semantic knowledge contained in product constitutes the basis and prerequisite of the process. The semantic knowledge that influences the assembly sequence planning mainly includes the following aspects: the part hierarchy of product, the feature attribute of part, and the physical attribute of part, etc. In view of this, the product semantic knowledge system is composed of three levels of spatial objects, namely the product layer, the component layer and the part layer. In addition, the system also covers the attributes of assembly parts, the relationship between assembly parts, the types of assembly constraints, and two types of core assembly relationships based on hierarchy and constraints.

[0129] Specifically, the definition of the formal description of product, firstly, the application defines the product as a set, denoted as , wherein represents n components constituting the product, represents m parts constituting the product. The component as a concept name represents a specific product entity. In addition, the application introduces the component name as an attribute for identifying the name of the product. In order to clearly describe the assembly hierarchy relationship between the product and the assembly, the application defines the following predicates: as assembly, assembled, as part and connected. These predicates collectively express that the product is a complex system composed of multiple assemblies and parts through specific assembly relationships.

[0130] Specifically, the formal description of component, similarly, the application defines the component as a set, denoted as , wherein represents e sub-components constituting the component, and represents f parts constituting the component. Here, the component as a concept name specifically refers to a certain component part of the product. The component name attribute is used to identify the name of the assembly. In order to describe the assembly hierarchy relationship between the assembly and the part, the application also uses the four predicates of as assembly, assembled, as part and connected. The four predicates collectively reveal that the assembly is a complex structure composed of multiple sub-components and parts.

[0131] Specifically, the definition of the formal description of part and its attributes, the application defines the part as a part set with physical attributes and geometric attributes, denoted as , wherein represents k parts constituting the assembly or product. The part as a concept name represents a basic constituent unit of the assembly or product, and the part name attribute is used to identify the name of the part.

[0132] In order to distinguish different attributes of the part, the application introduces the following two concepts, respectively:

[0133] hasPartNonGeometricAttributes and hasPartGeometricAttributes. hasPartNonGeometricAttributes represents the non-geometric attributes of the part, such as mass, precision, size, and other information that affects part reasoning; while hasPartGeometricAttributes is a predicate name that indicates that the part has geometric structure attributes. The definition of these attributes helps to more accurately describe and analyze the role and function of the part in the product and assembly.

[0134] The assembly rules determined based on the assembly semantic model include:

[0135] Rule 1: The part with lighter weight is assembled after the part with heavier weight, and the heaviest part is installed first as the reference part.

[0136] Rule 2: The part with higher precision is assembled after the part with lower precision. The box cover is assembled before the fastener.

[0137] Rule 3: The part with smaller size is assembled after the part with larger size.

[0138] Rule 4: The part with fewer assembly relationships is assembled after the part with more assembly relationships.

[0139] Rule 5: Based on the product center, the part farther from the product center is assembled after the part closer to the product center.

[0140] Rule 6: In the assembly sequence, the gear is installed after the key.

[0141] S107: Construct an assembly information ontology, and instantiate the assembly information ontology through two-dimensional mask images of all part images; the instantiated assembly information ontology includes object attributes of the parts and data attributes of the parts; the object attributes represent the connection relationships between the parts in the assembly body; and the data attributes represent the mass, size, precision, and position relationship of the parts.

[0142] The construction process of the assembly information ontology specifically includes: constructing an ontology model with the application domain being the assembly sequence planning domain; obtaining related terms of the ontology model, collecting concepts related to the assembly sequence planning domain, and supplementing the definition of ontology knowledge; defining the hierarchical relationship between classes in the ontology model; the class is a unary relationship; defining object attributes and data attributes in the ontology model; limiting the definition domain and value domain of the object attributes and data attributes according to the requirements of the assembly sequence planning; creating instances according to the assembly sequence planning to obtain the assembly information ontology.

[0143] The hierarchical relationship between classes is that the component is a superclass; the component is used to define the common attributes of the component's subclasses; the subclasses include fastening components, box components, driving components, and sealing components.

[0144] The data attributes include weight level, precision level, size, position relationship, start, end, assembly relationship number; the object attributes include part, assembly, assembled, mounted on, connected through, assembled after, start installation, end installation, and connection.

[0145] Specifically, with the development of ontology concepts, the most widely used method for building ontology applications is the seven-step method proposed by Stanford University:

[0146] 1. Determine the application field of the ontology, which is the assembly sequence planning field.

[0147] 2. Consider reusing existing ontologies, and need to rebuild a dedicated ontology model.

[0148] 3. List the important terms in the ontology, collect concept definitions related to assembly information, and supplement ontology knowledge definitions.

[0149] 4. Define the hierarchical relationship between classes, as shown in the hierarchical structure diagram of the assembly provided by the present application. According to the representation model, the term representing the one-to-one relationship is defined as a class, and the hierarchical relationship between classes is defined. The meaning of the hierarchical relationship between all classes in the assembly information ontology is as follows: the component is a superclass, used to define the common attributes of its subclasses, and the subclasses include fastening components, box components, driving components, and sealing components. Figure 7

[0150] 5. Define the attributes, and the terms representing the two-to-one relationship can be defined as attributes, as shown in the attribute definition diagram of the ontology information model provided by the present application. The specific value domain of the defined attribute in the assembly is defined in Table 1. The meaning is as follows: Figure 8

[0151] Object attributes: assembly before indicates that a component needs to be assembled before another component. For example, the key needs to be assembled before the gear. Connected through indicates the connection between two components. Connected through indicates the connection relationship between two parts, especially when connecting with a bolt set. Mounted on describes the relationship between a part mounted on another part. For example, the key is mounted on the shaft. Assembled after indicates that a component needs to be assembled after another component. For example, the gear needs to be assembled after the shaft.

[0152] Data attributes: weight level indicates the mass of the part; position relationship indicates the relative position relationship of the part; precision level indicates the precision of the part; size indicates the size of the part; assembly relationship number indicates the number of matching relationships of the part.​​

[0153] 6. Define the limit of attribute, limit the domain and value of attribute according to the requirement of assembly sequence planning.

[0154] 7. Create instance, create instance for given product assembly sequence planning according to the requirement of actual application. For example, the combination of parts needs to be instantiated according to the given assembly sequence planning.

[0155] Table 1

[0156]

[0157] S108: According to the object attribute and data attribute of the parts in the instantiated assembly information ontology, execute the assembly rule and perform reasoning to obtain the assembly sequence planning information.

[0158] According to the object attribute and data attribute of the parts in the instantiated assembly information ontology, execute the assembly rule and perform reasoning to obtain the assembly sequence planning information. Take the reducer as an example, the specific process is as follows:

[0159] 1. Build the assembly information ontology: complete the construction of the assembly information ontology, lay the foundation for subsequent work.

[0160] 2. Add part instance: add each part instance of the reducer product to the constructed ontology.

[0161] 3. Formulate semantic rules: use the assembly semantic model to construct semantic rules for reasoning the assembly sequence of parts in the Protege software.

[0162] 4. Execute the reasoning process: use the reasoning machine Pellet (Incremental) of Protege to perform reasoning according to the set semantic rules. In this process, the binary relationship between parts can be obtained. For example, referring to part of the binary relationship of the rules in Figure 14: (gear 1 assembly sequence is later than box 1), (gear 1 assembly sequence is later than box cover 1), (box cover 1 assembly sequence is later than box 1).

[0163] 5. Arrange the assembly sequence: integrate the binary relationship of parts obtained by reasoning, and finally determine the feasible assembly sequence after semantic analysis.

[0164] Specifically, the assembly information model is described using the Web Ontology Language (OWL), which has strong expressive power and logicality. The assembly rules are expressed using the Semantic Web Rule Language (SWRL), which enables reasoning between assembly information. The assembly sequence planning information is obtained through reasoning of part special dependency relationships and part attribute relationships.

[0165] Part special dependency relationship rule: For example, the key needs to be assembled before the gear.

[0166] In complex mechanical assembly processes, it is a crucial task to set the weight, precision, size, and relative position relationship between parts reasonably, which directly affects the performance, reliability, and production efficiency of the final product.

[0167] First, in terms of weight, larger parts tend to have greater inertia. Therefore, following the strategy of "assembling larger parts first" can effectively reduce the difficulty and complexity of subsequent assembly steps. This helps to avoid assembly precision decline or assembly failure due to excessive weight in the later assembly stage.

[0168] Second, precision is one of the important indicators of part manufacturing quality. In the assembly process, the principle of "the higher the precision, the later the assembly" can ensure that high-precision parts are installed in a more stable assembly environment, thereby minimizing assembly errors. This principle reflects the fine control of assembly order and is the key to ensuring the overall precision and performance of the product.

[0169] Third, the size factor also plays a crucial role in the assembly sequence. According to the logic of "assembling larger parts first", it can ensure that large parts have enough space for positioning and fixing during assembly. This helps to avoid assembly difficulties or quality problems caused by space limitations in the later assembly stage.

[0170] Finally, from the perspective of structural hierarchy, the relative position relationship between parts should follow the principle of "first internal then external". This means that during assembly, internal parts should be installed first, and then external parts should be installed. This principle helps to ensure the correct positioning and fixing of internal parts, providing stable support and foundation for the installation of subsequent external parts.

[0171] Cooperation quantity relationship rule: Through analysis of the three-dimensional model by CAD software, the cooperation relationship of the corresponding parts can be obtained. Parts with more cooperation relationships should be assembled earlier.

[0172] In an exemplary embodiment, to verify the feasibility of the proposed method, an experimental verification and analysis are carried out taking the reducer as an example. The experiment takes Python programming language as the experimental basis, and the configuration of the experimental platform is based on Windows 11 operating system, I9-13900KF CPU (3.00 GHz), GTX4080 GPU, CUDA 11.1, Python 3.8.13, PyTorch1.12.1. During the training process, the experimental environment of all algorithms is the same.

[0173] Target detection and segmentation task configuration and process.

[0174] In the target detection and segmentation task, the definition of positive samples is crucial. The present application adopts the candidate region with an intersection over union (IOU) greater than 0.5 as a positive sample, and only calculates the mask loss on the positive sample.

[0175] In the configuration aspect of the region extraction network, the positive and negative sample ratio is set to 1:3, and 5 different scales (specifically 8, 16, 32, 64, 128 pixels) and 3 different aspect ratios (specifically 1:1, 1:2, 2:1) are used to generate candidate regions. These settings aim to improve the adaptability of RPN to target size and shape.

[0176] In the feature extraction aspect, ResNet50 is adopted as the backbone network, and the feature map generated thereby is used by RPN to generate 300 candidate regions (‌Region of Interest‌, ROIs) for classification and bounding box regression. In order to further enhance the multi-scale nature of feature representation, a feature pyramid network is introduced. FPN provides more rich feature information for RPN by fusing feature maps of different levels. RPN generates candidate regions based on the multi-scale feature maps output by FPN, and then performs a non-maximum suppression (NMS) operation on the candidate regions to reduce the number of overlapping boxes.

[0177] In the detection and segmentation stage, the top 100 candidate regions in the detection score are selected for mask detection. The mask branch is responsible for predicting 7 categories of masks, selecting the mask of the corresponding category according to the classification result, and resetting the size of the mask to the size of the ROI. Finally, the mask is binarized using a threshold of 0.5 to obtain the final mask image. This process effectively realizes the accurate detection and segmentation of the target.

[0178] Performance analysis of the improved mask region convolutional neural network.

[0179] The original image quantity of the data set is 1910, including 344 reducer complete assembly model images, involving 7 parts. After data enhancement, all processed images are divided into training set, validation set and test set according to the ratio of 7:2:1, wherein the training set is the input image of model training, and the validation set is also crucial in the training process, which is used to evaluate and check the performance of the model at intervals. In order to provide accurate supervision information for subsequent assembly model part refinement segmentation, the assembly part images are accurately labeled, such as Figure 9 as shown in Figure 9 The fine labeling diagram of the reducer parts provided by the application. In the experiment, the evaluation indexes of the algorithm model include the average precision mean (Bbox) mAP50, the segmentation mask average precision mean (Seg) mAP50 and the average precision mean (MAR). The evaluation index data of different models is shown in Table 2.

[0180] Table 2

[0181]

[0182] Among them, (Bbox) mAP50 is used to identify and locate the target object in the image, which represents the main index of target position detection. (Seg) mAP50 is used to separate the target object from the background, which represents the segmentation index of the target mask. MAR is used to evaluate the performance of the target detection model, which represents the comprehensive index of detection accuracy in multiple categories. In the training process, the same data set and parameters are used as control variables. When Mask-U3 and Mask-U3-CBAM are used to compare the original Mask R-CNN, the mask branch of the original Mask R-CNN uses FCN Fully (Convolutional Network).

[0183] Compared with the Mask R-CNN algorithm, when the Mask-U3 algorithm replaces the mask branch with UNet3+, its performance shows significant improvement in key indicators. Specifically, (Bbox) mAP50 increases from 82.0% to 83.5%, with an increase of 1.5%; at the same time, (Seg) mAP50 also increases from 81.8% to 83.1%, with an increase of 1.3%. In addition, the average recall rate also shows progress, from 93.7% to 95.4%, with an increase of 1.7%.

[0184] Further, when CBAM attention modules are embedded in the backbone network of the Mask-U3 algorithm to form the Mask-U3-CBAM algorithm, the performance is optimized again. The (Bbox)mAP50 and (Seg)mAP50 are increased to 83.9% and 84.0%, respectively, which are increased by 0.4% and 0.9% compared with the Mask-U3 algorithm. The average recall rate also increases slightly from 95.4% to 95.8%, an increase of 0.4%.

[0185] Finally, compared with the original Mask R-CNN algorithm, the Mask-U3-CBAM algorithm combined with the mask branch replacement and the CBAM attention mechanism module in the backbone network realizes a more significant leap in performance. The (Bbox)mAP50 is greatly increased from 82.0% to 83.9%, an increase of 1.9%; the (Seg)mAP50 is increased from 81.8% to 84.0%, an increase of 2.2%; and the average recall rate is also significantly increased from 93.7% to 95.8%, an increase of 2.1 percentage points.

[0186] The above results not only verify the effectiveness of the mask branch replacement UNet3+, but also further reveal the potential of adding the CBAM attention mechanism module in improving the performance of target detection and instance segmentation.

[0187] The present application designs a set of experiments in the model training stage to explore the mAP50 values of different network structure models involved in the ablation experiment and the convergence of the loss function, and compares the mAP50 curves and loss function curves of different algorithms. The mAP50 values and overall loss function of different network structure models in the training process are shown in Figure 10 Based on the change trend of the mAP50 curve in Figure 10 , the (Seg)mAP50 value curves of the Mask-U3 and Mask-U3-CBAM algorithms have a leading advantage at the initial position compared with the original Mask R-CNN, and gradually stabilize at about the 15th epoch, with the highest values of 0.835 and 0.839, respectively, which are increased by 0.015 (i.e. 1.5%) and 0.019 (i.e. 1.9%) compared with the original Mask R-CNN of 0.820. After observation, it can be seen that the curve with attention mechanism has a faster rising trend and can reach a higher peak value.

[0188] The overall loss function of the Mask-U3 and Mask-U3-CBAM algorithms has a faster convergence speed than the original algorithm, and the numerical value after convergence is lower. The convergence is 0.1232 and 0.1208, respectively, which is decreased by 0.017 (i.e. 1.70%) and 0.0194 (i.e. 1.94%) compared with the original Mask R-CNN loss function of 0.1402.

[0189] According to the observation of the loss function curve before and after embedding the CBAM attention module, after embedding the attention module, the loss function decreases more and fits faster to reach the fitting state.

[0190] The prediction results of the three network models on the reducer model image are shown in Figure 11 Figure 11 Sealing_ring in is a sealing ring, Key is a key, Bearing is a bearing, Axle is a shaft, Gear is a gear, Box is a box, End_plate is an end plate, Figure 11 (1) in the A diagram in is Mask R-CNN, Figure 11 (2) in the A diagram in is Mask-U3, Figure 11 (3) in the A diagram in is Mask-U3-CBAM, which shows the effect of performing instance segmentation on the reducer image based on the Mask R-CNN model. The Mask R-CNN model shows high accuracy in part detection, successfully identifying most of the parts in the image, with only the key in the position set to exist being undetected. In terms of segmentation mask, it is observed that the boundary definition of a small number of parts is not clear enough, especially in the End_plate area, there is a certain degree of incomplete mask coverage phenomenon.

[0191] Subsequently, the present application introduces the Mask-U3 model which uses the UNet3+ network to optimize Mask R-CNN. Compared with the segmentation results of the benchmark model, the accuracy of Mask-U3 in part detection has been significantly improved, and the key in the above-mentioned which was not detected is successfully identified. From the figure, it can be clearly seen that the quality of the segmentation result has been greatly improved, especially in the End_plate area, the problem of incomplete mask coverage has been effectively solved, and the mask can accurately and completely cover the parts to be segmented, which further verifies the effectiveness of the improved model.

[0192] On this basis, the present application further proposes the Mask-U3-CBAM model, which is realized by integrating the CBAM attention mechanism in the backbone network on the basis of Mask-U3. The prediction result figure directly shows the significant improvement of the model in segmentation effect. The Mask-U3-CBAM model not only realizes the accurate segmentation of all parts, but also reaches a very high level in the accurate positioning of the part detection frame and the complete coverage of the segmentation mask.

[0193] As shown in the right part of Figure 11 , the benchmark model and the improved model are used to perform instance segmentation on different parts of the reducer, and the evaluation indicators of each part of the model, i.e. precision, recall and F1 score. ​

[0194] Figure 11 Figure A shows the application of the accuracy index. Accuracy is a key indicator for evaluating the prediction performance of the model. It specifically reflects the proportion of positive samples predicted by the model to be positive samples. In the figure, the bar chart intuitively presents the comparison of the accuracy data of the baseline model Mask R-CNN and the improved models Mask-U3 and Mask-U3-CBAM when segmenting different parts. At the same time, the line chart corresponding to the bar chart clearly reveals the improvement trend of the improved model in terms of accuracy. Through careful observation, the present invention clearly concludes that compared with the baseline model Mask R-CNN, the Mask-U3 and Mask-U3-CBAM models show higher accuracy in instance segmentation tasks, which shows that the proposed improvement strategy is effective.

[0195] Figure 11 (1) in Figure B is Mask R-CNN, Figure 11 (2) in Figure B is Mask-U3, Figure 11 (3) in Figure B is Mask-U3-CBAM, Figure 11 Figure B in the figure focuses on the recall rate metric, which is used to quantify the proportion of all actual positive samples that are correctly predicted as positive samples by the model. In this figure, the bar chart compares the recall rate data of the Mask R-CNN, Mask-U3, and Mask-U3-CBAM models in different part segmentation tasks. The accompanying line chart vividly demonstrates the improvement in recall rate of the improved model. Through in-depth analysis, the present invention determines that the Mask-U3 and Mask-U3-CBAM models have significant advantages over the Mask R-CNN model in instance segmentation recall rate, which further verifies the effectiveness of the improvement strategy.

[0196] Figure 11 (1) in Figure B is Mask R-CNN, Figure 11 (2) in Figure B is Mask-U3, Figure 11 (3) in Figure B is Mask-U3-CBAM, Figure 12The C diagram in the figure shows the evaluation results of the F1 score, which is the harmonic mean of precision and recall, achieving a good balance between the two. In the figure, the column chart intuitively compares the F1 scores of the Mask R-CNN, Mask-U3 and Mask-U3-CBAM models in different part segmentation tasks. At the same time, the line chart corresponding to the column chart clearly reveals the improvement of the improved model in F1 score. After careful analysis, the present application concludes that compared with the benchmark model Mask R-CNN, the Mask-U3 and Mask-U3-CBAM models perform better in instance segmentation F1 score, which again proves the effectiveness and practicality of the improvement strategy.

[0197] Reducer body modeling and reasoning results.

[0198] According to the proposed assembly sequence planning knowledge model construction method, the present application realizes the construction of the ontology model. Figure 12 The example node diagram of the assembly provided by the present application is shown in Figure 13 The assembly hierarchy and assembly structure are used as nodes to form hierarchical relationships, assembly structure classes and assembly part classes. The Mask R-CNN neural network model is used to perform instance segmentation tasks in the assembly model, and the instances contained in the reducer assembly model are parsed. The training set and the validation set composed of the reducer assembly model and related parts are used to obtain the improved Mask R-CNN model in the field of assembly sequence planning, and a primary cylindrical gear reducer in the test set is taken as an example. The assembly model is labeled by instance segmentation, and the obtained instance segmentation is shown in Figure 13 Figure 14 The exploded view of the segmented reducer and parts provided by the present application is shown in

[0199] After obtaining the instance composition of the assembly model, the present application models the attributes of the extracted instances to form binary relationships between instances. The assembly process of the primary cylindrical gear reducer is constructed in a visual form of knowledge model, and the SWRL rule is used for knowledge reasoning to obtain an industrial feasible assembly sequence planning, as shown in Figure 14 Figure 15 The ontology instance and rule reasoning result diagram provided by the present application is shown in

[0200] Then the reasoning machine Pellet (Incremental) is enabled in protege, and the SWRL rule is selected for reasoning. After sorting the reasoning results, the following​​Figure 16 a series of feasible assembly sequences.

[0201] In an exemplary embodiment, Figure 1 As the overall method framework of the application, the overall method framework is guided by instance segmentation of the assembly model, realizes assembly sequence planning of the assembly model from the semantic angle of assembly process knowledge, and mainly includes three steps:

[0202] Step one, image data processing. Collect parts and product images through different assembly product 3D models, perform image enhancement processing on the image edges by using Laplacian sharpening method, and perform noise reduction processing on the edge-enhanced pictures to eliminate the influence of noise.

[0203] Step two, improved Mask R-CNN network model. Based on the Mask R-CNN network structure, the feature extraction network structure is replaced by the UNet3+ network which extracts more detailed features, and the attention mechanism is added to improve the performance of the improved neural network model.

[0204] Step three, ontology module. Perform assembly body information modeling, instantiate the part information obtained through semantic segmentation, and express the empirical knowledge in a regularized manner, and reason to obtain a feasible assembly sequence.

[0205] The application combines deep learning in the field of artificial intelligence, ontology theory and intelligent assembly to propose an intelligent planning method for assembly sequence of mechanical products. First, an improved Mask R-CNN model is used for instance segmentation, combined with Resnet50 for pre-training to improve the generalization ability of the model, which can more accurately extract the part information contained in the assembly body. Second, the ontology semantic information model is constructed by comprehensively considering the assembly prior knowledge, assembly information and part relationship, the hierarchical relationship of product, component and part, and the important regulations such as the weight, precision and size attributes of the part are described, the reasoning rules for assembly sequence planning generation are created by using SWRL language, and the assembly sequence information of the product is reasoned. Finally, the reducer is taken as an example to verify the feasibility and effectiveness of the method.

[0206] ​The application aims to reduce the cost of assembly planning and improve the efficiency of generating assembly sequence by combining semantic segmentation technology and ontology modeling method. Specifically, an improved Mask R-CNN deep learning model is used to accurately segment the assembly body parts in the image, so as to obtain detailed assembly part information. On this basis, a perfect assembly information model is further constructed using ontology, which can comprehensively and accurately describe the parts, assembly relationship and assembly constraints in the assembly body. Through in-depth research and practice, the application successfully combines deep learning technology with the actual needs of the mechanical engineering field, writes semantic rules for assembly sequence reasoning, and finally realizes the efficient generation of product assembly sequence. The experimental results show that the method provided by the application can not only accurately segment each part in the assembly body, but also efficiently process and utilize the assembly information based on the ontology model, so as to generate an efficient and feasible assembly sequence.

[0207] Future work will focus on more refined deep neural network models, and the instance segmentation accuracy and speed of the assembly body model are the focus of the research, especially when dealing with complex assembly bodies and small parts.

[0208] When applying the assembly sequence planning method provided by the application, the execution order of each step shown in Figure 17 may be determined according to the needs, and the application does not limit the execution order of each step.

[0209] The above is an assembly sequence planning method provided by one or more embodiments of the application. Based on the same idea, the application also provides a corresponding assembly sequence planning device, as shown in Figure 17 .

[0210] Figure 1 The assembly sequence planning device provided by the application is a schematic diagram, which comprises:

[0211] The first acquisition module 1701 is used to acquire an initial mask region convolutional neural network; the initial mask region convolutional neural network comprises a backbone network and a mask branch; the backbone network is constructed through a residual network 50 and a pyramid network module.

[0212] The improvement module 1702 is used to improve the full convolutional network in the mask branch by adopting a UNet3+ network structure, and insert a convolution block attention module into each bottleneck residual block layer of the residual network 50, to obtain a mask region convolutional neural network.

[0213] The second acquisition module 1703 is used to acquire an assembly body part image data set, and pre-process the assembly body part image data set to obtain a pre-processed assembly body part image data set.

[0214] The training module 1704 is used to train the mask region convolutional neural network using the preprocessed assembly part image dataset to obtain a target mask region convolutional neural network.

[0215] Extraction module 1705 is used to obtain multiple part images of the target assembly to be assembled, and perform feature extraction on the multiple part images of the target assembly through the target mask region convolutional neural network to obtain a two-dimensional mask image for each part image; each pixel in the two-dimensional mask image represents the part type.

[0216] The third acquisition module 1706 is used to acquire the assembly rules of the target assembly; the assembly rules are determined based on the quality factors, precision factors, size factors, connection relationship factors, position factors and the preset part assembly sequence relationship of each part.

[0217] Construction module 1707 is used to construct the assembly information ontology, and instantiate the assembly information ontology through the two-dimensional mask image of all part images; the instantiated assembly information ontology includes the object attributes of the parts and the data attributes of the parts; the object attributes represent the connection relationship between the parts in the assembly; the data attributes represent the quality, size, accuracy and position relationship of the parts.

[0218] The reasoning module 1708 is used to execute assembly rules according to the instantiated assembly information ontology and perform reasoning to obtain assembly sequence planning information.

[0219] The specific definition of an assembly sequence planning device can be found in the definition of an assembly sequence planning method above and will not be repeated here. The various modules in the above-mentioned assembly sequence planning device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above-mentioned modules.

[0220] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 18 An assembly sequence planning method is provided.

[0221] The present invention also provides Figure 18 The structural diagram of the computer equipment shown in FIG. Figure 1 As shown in the figure, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above ​An assembly sequence planning method is provided.

[0222] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by a computer program instructing relevant hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in each embodiment of the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0223] Each technical feature of the above embodiments can be combined arbitrarily, and in order to make the description simple, not all possible combinations of each technical feature in the above embodiments are described, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.

Claims

1. An assembly sequence planning method characterized by, The method comprises the following steps: An initial mask region convolutional neural network is acquired; the initial mask region convolutional neural network comprises a backbone network and a mask branch; The backbone network is constructed by a residual network 50 and a pyramid network module; The full convolutional network in the mask branch is improved by adopting a UNet3+ network structure, and a convolution block attention module is inserted after each bottleneck residual block layer in the residual network 50 to obtain a mask region convolutional neural network; An assembly part image dataset is acquired, and the assembly part image dataset is preprocessed to obtain a preprocessed assembly part image dataset; The mask region convolutional neural network is trained by using the preprocessed assembly part image dataset to obtain a target mask region convolutional neural network; A plurality of part images of a target assembly to be assembled are acquired, and the plurality of part images of the target assembly are subjected to feature extraction by using the target mask region convolutional neural network to obtain a two-dimensional mask image of each part image; each pixel point in the two-dimensional mask image represents a part type; An assembly rule of the target assembly is acquired; the assembly rule is determined according to a quality factor, an accuracy factor, a size factor, a connection relationship factor, a position factor and a preset part assembly sequence relationship of each part; An assembly information ontology is constructed, and the assembly information ontology is instantiated by using the two-dimensional mask images of all the part images; the instantiated assembly information ontology comprises object attributes of parts and data attributes of the parts; the object attributes represent connection relationships between the parts in the assembly; and the data attributes represent qualities, sizes, accuracies and position relationships of the parts; According to the instantiated assembly information ontology, the assembly rule is executed, and assembly sequence planning information is obtained through reasoning.

2. The method of claim 1, wherein, The mask region convolutional neural network comprises a region proposal network, an anchor box mapping module, an interesting region alignment module and a fixed feature map size module; the feature extraction of the plurality of part images of the target assembly by using the target mask region convolutional neural network to obtain the two-dimensional mask image of each part image specifically comprises the following steps: For any part image of the target assembly, the part image is input into the backbone network, and the residual network 50, the convolution block attention module and the pyramid network module are sequentially connected to process the part image to obtain a processing result of the backbone network; The processing result of the backbone network is input into the region proposal network to obtain an output of the region proposal network; The output of the region proposal network and the processing result of the backbone network are input into the anchor box mapping module to obtain an output of the anchor box mapping module; The output of the anchor box mapping module is input into the interesting region alignment module to obtain a processing result of the interesting region alignment module; The processing result of the interesting region alignment module is sampled to a preset size by using the fixed feature map size module to obtain a sampling result. The sampling result is input into the UNet3+ network structure for up-sampling and down-sampling operations to obtain a mask feature map; the mask feature map has a preset size and a plurality of channels; each channel of the mask feature map represents different feature information. The mask feature map is classified to generate a two-dimensional mask image of the part image.

3. The method of claim 1, wherein, The data attributes include weight level, precision level, size, position relationship, start, end, assembly relationship quantity.

4. The method of claim 1, wherein, The object attributes include part, assembly, assembled, mounted on, connected through, assembled, start installation, end installation and connection.

5. The method of claim 1, wherein, The preprocessing of the assembly part image dataset to obtain a preprocessed assembly part image dataset specifically includes: Each part image in the assembly part image dataset is subjected to diversity processing; the diversity processing includes angle transformation, texture change and size adjustment; The part image subjected to the diversity processing is subjected to sharpening processing by Laplacian sharpening; The part image subjected to the sharpening processing is subjected to median filtering to obtain a preprocessed image; All the preprocessed images are combined into the preprocessed assembly part image dataset.