Intelligent segmentation method for depth map of aero-engine component based on deep convolutional network
By introducing quantitative analysis of perspective difference and boundary deformation rate and non-rigid alignment mechanism, the problem of template registration failure in complex mechanical structures is solved, high-precision automatic segmentation of aviation engine components is achieved, and the cost of manual labeling is reduced.
Patent Information
- Application Number
- CN202510975817.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-09
AI Technical Summary
When existing technologies migrate shared label templates across images, registration fails due to differences in perspective, scale, and occlusion. This is especially true in multi-angle images of complex mechanical structures such as turbine disks, where label misalignment or boundary errors are severe, making it difficult to adapt to scenes with strong structural heterogeneity.
By acquiring depth images of aircraft engine components, the initial template mask is generated using traditional image processing algorithms, and the template registration error is analyzed by combining the view difference ADI and the boundary deformation rate BDR. A non-rigid alignment mechanism is introduced, and a deep convolutional neural network is used for training to generate a model for semantic segmentation of aircraft engine components.
It significantly improves the adaptability and annotation accuracy of template migration, realizes high-precision, automated pixel-level semantic segmentation of aero-engine components, reduces manual annotation costs, and is suitable for industrial vision application scenarios with complex structures.
Smart Images

Figure CN120612487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to an intelligent segmentation method for depth images of aircraft engine components based on a deep convolutional network. Background Art
[0002] Intelligent segmentation of aircraft engine component depth images based on deep convolutional networks (CNNs) utilizes deep convolutional neural networks (CNNs) within deep learning to automatically identify and accurately segment aircraft engine component depth images. This method, through training a model to learn features from a large number of depth images, can intelligently distinguish and extract the boundaries and structural information of different components. This enables efficient and accurate processing of complex mechanical structure images, enhancing the intelligent level of engine component inspection, maintenance, and assembly.
[0003] The existing technology has the following shortcomings:
[0004] When transferring automatic annotations using shared label templates across images, registration often fails due to differences in viewpoint, scale, and occlusion, leading to label misalignment or boundary errors. This is particularly noticeable in multi-angle images of complex mechanical structures such as turbine disks. Because this method relies on fixed geometric relationships, it struggles to adapt to scenes with strong structural heterogeneity. Existing annotation tools lack non-rigid alignment mechanisms based on keypoints, optical flow, or deformation maps, limiting the accuracy and applicability of template transfer. Summary of the Invention
[0005] The purpose of the present invention is to provide an intelligent segmentation method for depth images of aircraft engine components based on deep convolutional networks to address the shortcomings of the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network, comprising:
[0007] Acquiring depth image data of aircraft engine components and preprocessing the depth image;
[0008] Generate component masks using traditional image processing algorithms in the first annotated image, construct an initial template mask, and establish image and label pairs using image annotation tools;
[0009] The initial template mask is transferred to the target image to generate a candidate mask, and the view difference ADI is extracted from the target image to measure the angle difference between the main direction vectors of the corresponding component in the template image and the target image; and the boundary deformation rate BDR is extracted to measure the degree of geometric deformation between the key points of the component contour;
[0010] Determining the template registration error level based on the comprehensive analysis results of ADI and BDR, and executing a non-rigid alignment mechanism for images with a severe registration error level;
[0011] The registered mask and depth map are fed into a deep convolutional neural network for training to generate a model for semantic segmentation of aerospace engine parts.
[0012] The model is applied to unlabeled depth maps for pixel-level semantic segmentation of parts.
[0013] Preferably, generating an initial template mask through an image processing algorithm includes: performing Canny edge detection on the depth map to obtain boundary information; using a region growing algorithm to fill the closed boundary area to form an initial mask; performing morphological processing on the generated mask, including dilation, erosion and hole filling; importing the generated mask into an image annotation tool for manual review and refinement to generate an image label pair.
[0014] Preferably, the method for obtaining the viewing angle difference ADI is as follows: performing principal component analysis on the corresponding component contours in the template image and the target image to obtain the first principal component direction; defining the first principal component direction vector as the component principal direction; assuming the principal direction vector in the template image is v1 and the principal direction vector in the target image is v2; and calculating the angle θ between the two, the expression is: Among them: ∥v1∥, ∥v2∥ respectively represent the modulus of the two vectors; the unit of θ is radians, which is converted to angle expression to calculate the viewing angle difference ADI, and the expression is: ADI is the visual angle difference in degrees, π is the circumference of a circle, and θ is the angle between the main directions of the component in the template and the target image.
[0015] Preferably, the boundary deformation rate BDR is obtained by normalizing the template mask and the target image mask to a uniform scale, extracting N corresponding key points from the contour; the key point in the template image is denoted as P i =(x i ,y i ), the corresponding key point in the target graph is Q i =(x′ i ,y′ i ); for each key point, calculate the Euclidean distance from it to the contour centroid, which are: The deformation rate is expressed as: BDR is the boundary deformation rate, N is the Among them C T 、C G Calculate the number of boundary key points for the centroids of the contours in the template image and the target image respectively; is the distance from the i-th key point to its respective centroid.
[0016] Preferably, the template registration error level is determined based on the comprehensive analysis results of ADI and BDR. For images with a severe registration error level, a non-rigid alignment mechanism is executed, specifically including:
[0017] The perspective difference and boundary deformation rate are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model uses the template registration error score value label predicted by each set of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of all template registration error score value labels as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The template registration error score value is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0018] Preferably, the obtained template registration error score is compared with a preset threshold. If the template registration error score is greater than or equal to the preset threshold, it is considered that the current template registration error is serious and non-rigid alignment needs to be triggered; if the template registration error score is less than the preset threshold, it is considered that the registration error is slight and the affine registration result continues to be used.
[0019] Preferably, the non-rigid alignment mechanism comprises:
[0020] Perform local feature point extraction on the component area in the template image and the target image respectively; suppose the key point set extracted in the template image is In the target graph, Calculate the descriptor similarity between the two and match several candidate point pairs
[0021] A random sampling consensus algorithm is used for robust matching screening, which minimizes the geometric error between matching point pairs and retains the set of interior points that meet most structural constraints.
[0022] The result after eliminating matching point pairs is used to construct the control point set for deformation estimation, denoted as: P = where p i is the control point in the template graph, p′ i is the corresponding control point in the target graph, and b is the total number of points;
[0023] A non-rigid mapping relationship is constructed based on the matching point pairs, and a transformation function with minimum bending energy is constructed to achieve continuous deformation of the template coordinate system to the target image coordinate system. The transformation definition expression is: Where: (x, y) is any pixel in the template; U is the radial basis function; a1, a2, a3 and w i is the transformation parameter, which is solved by least square fitting;
[0024] Apply the desired deformation field to the initial template mask image, and use bilinear interpolation and other methods to map the coordinates of all pixels in the mask to obtain the deformed mask: M ′ (x ′ ,y ′ )=M(f -1 (x′,y′)); where M is the template mask, M′ is the mask after registration, and f -1 is the inverse deformation mapping.
[0025] Preferably, inputting the mask and depth map into a deep convolutional neural network for training includes: determining a semantic segmentation network and constructing a training set, including a preprocessed depth map and a registered mask label map; using a cross entropy or Dice loss function as a supervisory signal; applying an image enhancement strategy, including rotation, scaling, and affine perturbation; and performing iterative training until the loss converges.
[0026] Preferably, applying the trained model to the unlabeled depth map includes: performing normalization and resizing operations on the unlabeled depth map; inputting the image into the trained model to perform forward reasoning and outputting a pixel-level prediction mask; mapping different category labels in the mask to different colors to generate semantic segmentation visualization results.
[0027] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0028] 1. This invention quantitatively analyzes template transfer registration errors by introducing perspective difference and boundary deformation rate, and combines this with a polynomial regression model to intelligently predict and grade registration errors. This effectively addresses the existing issues of template registration failure and label misalignment caused by perspective, scale variations, and complex structural differences. For higher error levels, a non-rigid alignment mechanism based on key point matching and thin plate spline (TPS) transformation is further introduced. This allows the mask to accurately fit the target component boundary even in target images with significant structural deformation, significantly improving the adaptability of template transfer and the accuracy of labeling.
[0029] 2. This paper constructs a closed-loop system from data generation, error determination, non-rigid correction, to model training and deployment, enabling high-precision, automated pixel-level semantic segmentation of depth maps of aircraft engine components. This method not only improves training data quality and model segmentation accuracy, but also significantly reduces manual annotation costs, demonstrating strong engineering feasibility and potential for widespread adoption. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0031] Figure 1 This is a mind map of the method of the present invention. DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0033] For examples, see Figure 1 As shown, the method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network described in this embodiment includes:
[0034] Acquiring depth image data of aircraft engine components and preprocessing the depth image;
[0035] Generate component masks using traditional image processing algorithms in the first annotated image, construct an initial template mask, and establish image and label pairs using image annotation tools;
[0036] The initial template mask is transferred to the target image to generate a candidate mask, and the view difference ADI is extracted from the target image to measure the angle difference between the main direction vectors of the corresponding component in the template image and the target image; and the boundary deformation rate BDR is extracted to measure the degree of geometric deformation between the key points of the component contour;
[0037] Determining the template registration error level based on the comprehensive analysis results of ADI and BDR, and executing a non-rigid alignment mechanism for images with a severe registration error level;
[0038] The registered mask and depth map are fed into a deep convolutional neural network for training to generate a model for semantic segmentation of aerospace engine parts.
[0039] The model is applied to unlabeled depth maps for pixel-level semantic segmentation of parts.
[0040] In an embodiment of the present invention, depth image data of an aircraft engine component is first acquired. This depth image can be obtained using a structured light 3D camera, a time-of-flight camera (ToF camera), a laser scanning system, or a point cloud reconstruction method to ensure that accurate geometric information of the component surface is captured. To enhance the accuracy and stability of subsequent image processing and semantic segmentation models, the acquired depth image must be systematically preprocessed. This preprocessing includes, but is not limited to, the following steps:
[0041] Denoising is performed on the original depth image to reduce random noise caused by lighting interference, material reflections, or sensor jitter. In this embodiment, a median filter is used to remove isolated pixels, and a bilateral filter is used to maintain edge contours while smoothing regional noise. For areas with high structural requirements (such as blade edges or shell seams), a joint filter can be used to enhance boundary preservation by combining RGB image information.
[0042] Normalize the depth image. By unifying the depth range and linearly mapping the depth values to the interval [0, 1] or [0, 255], this improves the consistency of the image representation within the neural network. Furthermore, for scenes with highly concentrated depth values, such as those with highly reflective surfaces or relatively fixed distances, histogram equalization can be used to enhance the image's structural contrast.
[0043] Furthermore, a hole filling operation is performed on the depth map's holes caused by occlusion, reflection, or measurement blind spots. In this embodiment, a neighborhood-based interpolation approach is preferentially used to fill small missing areas to ensure image continuity. For complex or large missing areas, structure-preserving filling is performed by combining the guidance image and edge information. Multi-view reconstruction data can be used to assist in the restoration, if necessary, to improve the integrity and accuracy of key structures (such as blade roots, joints, and bolt holes).
[0044] After completing the above preprocessing, the output depth image will have higher structural clarity and edge fidelity, significantly improving the subsequent component segmentation accuracy based on deep convolutional networks. It is particularly suitable for industrial vision application scenarios in aviation engine components with complex boundaries, highly reflective surfaces, and drastic depth changes.
[0045] In the embodiment of the present invention, a representative depth image of an aircraft engine component is first selected as the source image for initial template construction. The image should have good structural clarity and typical component poses to ensure the universality and transferability of the template.
[0046] On this image, a traditional image processing algorithm is used to generate an initial mask of the component. Specifically, an edge detection operation is first performed on the original depth map, such as using the Canny edge detection operator to extract the main boundary outline of the component; then, the region growing algorithm is combined to fill the area within the closed boundary to form a preliminary target area mask. In addition, the GrabCut image segmentation method can be combined to further refine the boundary based on the foreground / background contrast information to improve the accuracy of the target area. In order to improve the robustness of the segmentation, the above-mentioned multiple processing results can also be fused, and their union or intersection can be taken to enhance the integrity of the mask. Finally, morphological operations (including dilation, erosion, hole filling, etc.) are performed on the mask to eliminate isolated pixels and internal noise, and obtain an initial component mask with a continuous structure and clear boundaries.
[0047] The generated initial mask is then imported into an image annotation tool for manual review and refinement. In this example, commonly used image annotation tools such as CVAT (Computer Vision Annotation Tool) or LabelMe are used to inspect and correct the automatically generated mask region by region, paying particular attention to the accuracy of annotation of component edges, sharp corners, and occluded areas. After annotation is complete, the image and mask are paired to generate a standard "image-label" pair.
[0048] The label map is saved as a single-channel image, where different pixel values correspond to different semantic categories, such as "0" for background and "1" for target components. This category correspondence is recorded in a label mapping file (classes.txt). To facilitate subsequent training and data expansion, images and labels are organized into a unified structure, and the final approved mask is archived as the initial template mask for multi-image sharing and template migration.
[0049] Through this implementation, an accurate initial component mask template can be efficiently constructed, and a structurally complete image and label pair dataset can be established, providing high-quality, standardized basic data for subsequent deep learning-based segmentation model training.
[0050] The initial template mask constructed above is transferred to the target image to achieve rapid annotation and expansion of multiple depth images of aircraft engine components.
[0051] Specifically, the basic feature information of the target image to be transferred is first extracted, including image size, main direction angle, and outline center coordinates, to ensure comparability between the target image and the template image. Then, an affine transformation is used to geometrically adjust the initial template mask, including rotation, scaling, and translation operations, to ensure that the template outline matches the position and posture of the corresponding part in the target image as much as possible.
[0052] First, the structural principal axis angle between the initial template mask and the target image is calculated, and a rotation transformation is performed accordingly; secondly, position alignment is performed by bounding box or centroid alignment; if there is a scale difference, the template is further proportionally adjusted by applying a uniform scaling factor to form a transformed candidate mask.
[0053] By overlaying the candidate mask on the target image, a preliminary migrated segmentation region is obtained. To further improve matching accuracy, edge information can be extracted from the target image and compared with the candidate mask outline. Mask quality can be assessed using edge overlap or Intersection over Union (IoU) metrics. In cases of insufficient edge consistency or significant morphological deviation, a non-rigid registration mechanism can be introduced to perform mask correction.
[0054] Through this migration step, it is possible to quickly perform mask inference on multiple depth maps with similar structures but different poses while maintaining annotation consistency, effectively reducing the cost of manual annotation and providing a foundation for the construction of large-scale training data.
[0055] The angle difference ADI is used to measure the angle difference between the main direction vectors of the same component in the initial template image and the target image, reflecting the degree of deviation in the posture (angle) of the component in the two images.
[0056] Perform principal component analysis (PCA) on the corresponding component contours (i.e., mask areas) in the template image and the target image to obtain the first principal component direction. Define the first principal component direction vector as the component principal direction. Let the principal direction vector in the template image be v1 and the principal direction vector in the target image be v2. Calculate the angle θ between the two, as follows: Among them: ∥v1∥, ∥v2∥ respectively represent the modulus of the two vectors; the unit of θ is radians, which can be converted into angle representation to calculate the viewing angle difference ADI, the expression is: ADI is the visual angle difference, measured in degrees, π is pi, and θ is the angle between the main directions of the component in the template and the target image, reflecting the change in posture. The larger the ADI value, the greater the difference in component posture and the more difficult the template alignment.
[0057] The boundary deformation rate (BDR) is used to measure the deformation degree of the boundary structure of the same component in the template image and the target image, reflecting the geometric deformation between the corresponding points of the contour.
[0058] Normalize the template mask and the target image mask to a unified scale (e.g., within a unit bounding box) to ensure that the deformation rate is not affected by scale.
[0059] Use uniform sampling or key point detection algorithm (such as RDP simplification, centroid angle sampling) to extract N corresponding key points from the contour; the key point in the template image is denoted as P i=(x i ,y i ), the corresponding key point in the target graph is Q i =(x′ i ,y′ i ); for each key point, calculate its Euclidean distance to the contour centroid (center point), which are: Among them C T 、C G are the centroids of the contours in the template image and the target image respectively. The boundary deformation rate is calculated as follows: BDR is the boundary deformation rate, which measures the geometric structure deviation of the target image relative to the template image; N is the number of key points (usually 50 to 100, adjusted according to the complexity of the contour); is the distance from the i-th key point to its centroid; the larger the BDR, the more obvious the deformation and the higher the risk of template misregistration.
[0060] The template registration error level is determined based on the comprehensive analysis results of ADI and BDR. For images with severe registration error levels, a non-rigid alignment mechanism is implemented, specifically including:
[0061] The perspective difference and boundary deformation rate are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model uses the template registration error score value label predicted by each set of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of all template registration error score value labels as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The template registration error score value is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0062] The obtained template registration error score is compared with the preset threshold. If the template registration error score is greater than or equal to the preset threshold, it is considered that the current template registration error is serious and non-rigid alignment needs to be triggered; if the template registration error score is less than the preset threshold, it is considered that the registration error is slight and the affine registration result continues to be used.
[0063] In an embodiment of the present invention, for a target image that is determined to have a serious registration error by the template registration error score, a non-rigid alignment step is further performed to improve the adaptability and registration accuracy of the template in complex deformation scenes.
[0064] First, local feature point extraction is performed on the component area (i.e., the mask area or near the contour boundary) in both the template image and the target image. Preferably, the Scale Invariant Feature Transform (SIFT) algorithm or the Orbital Robust Feature (ORB) algorithm is used to extract a set of key points and their descriptors that are rotationally and scale-invariant.
[0065] Suppose the set of key points extracted from the template graph is In the target graph, respectively
[0066] Calculate the descriptor similarity between the two and match several candidate point pairs
[0067] To eliminate false matching or structurally inconsistent point pairs, the RANSAC (Random Sampling Consensus) algorithm is used for robust matching screening. By minimizing the geometric error between matching point pairs, the set of inliers that meet most structural constraints is retained.
[0068] The result after eliminating matching point pairs is used to construct the control point set for deformation estimation, which is recorded as: where p i is the control point in the template graph, p′ i is the corresponding control point in the target graph, and b is the total number of points;
[0069] A non-rigid mapping relationship is constructed based on the matching point pairs. Preferably, a Thin-Plate Spline (TPS) transformation algorithm is used to achieve continuous deformation from the template coordinate system to the target image coordinate system by constructing a transformation function with minimum bending energy.
[0070] The TPS transformation definition expression is: Where: (x, y) is any pixel in the template; U is the radial basis function; a1, a2, a3 and w i is the transformation parameter, and is solved by least squares fitting. In another optional way, a dense optical flow algorithm (such as or RAFT), estimates the pixel-level motion vector between the template image and the target image, and thus constructs a dense deformation field.
[0071] Apply the above deformation field to the initial template mask image, and use bilinear interpolation and other methods to map the coordinates of all pixels in the mask to obtain the deformed mask: M ′ (x ′ ,y ′ 0=M(f -1 (x′,y′)); where M is the template mask, M′ is the mask after registration, and f -1 is the inverse deformation mapping.
[0072] The resulting non-rigid registration mask M′ will fit the actual contour of the component in the target image more accurately, and is especially suitable for situations where there are structural changes such as nonlinear deformation, local bending, and posture rotation.
[0073] This non-rigid alignment step significantly improves the adaptability of the template in complex images, effectively avoids problems such as boundary mismatch and regional offset caused by rigid affine transformation in multi-view and deformation environments, and improves the matching accuracy of the semantic mask and the consistency of the model training data.
[0074] In an embodiment of the present invention, the semantic mask after registration processing is combined with the corresponding depth image data as an input sample pair and input into a deep convolutional neural network for supervised training, thereby constructing a model for semantic segmentation of depth images of aircraft engine components.
[0075] Specifically, we first organize all the labeled images and construct a training dataset containing several "image-label" pairs. Each pair of data consists of the following two parts:
[0076] Input image: pre-processed depth map of aircraft engine components, in the form of a single-channel grayscale image or a three-channel expanded pseudo-color image;
[0077] Label image: A high-quality mask image obtained after registration, where each pixel corresponds to a semantic category label. It is usually a single-channel, 8-bit image, and the pixel values represent different component categories (such as background, blades, shell, etc.).
[0078] Next, a deep convolutional neural network architecture suitable for semantic segmentation tasks is selected for model construction. In this embodiment, a U-Net or DeepLabV3+ network architecture is used. The network input is a uniformly sized depth map image, and the output is a predicted mask image of the same size as the original image. The encoder part of the network extracts multi-level spatial semantic features, and the decoder part gradually restores the spatial resolution, ultimately outputting the category prediction for each point at the pixel level.
[0079] During the training process, the following technical solutions are adopted:
[0080] Use the cross-entropy loss function (Cross-Entropy Loss) or Dice loss function as a supervisory signal to guide the model to learn the mapping relationship between the input image and the label map;
[0081] If there is a class imbalance problem, class weights or focal loss functions can be introduced to compensate;
[0082] Simultaneously perform data augmentation operations (such as image rotation, scale transformation, affine perturbation, etc.) to improve the model's robustness to component structural deformation and perspective shift;
[0083] Use Adam or SGD optimizer to update the gradient until the loss converges.
[0084] After model training, a new image semantic segmentation model generalizable to unlabeled depth maps is obtained. This model can automatically identify the boundaries and categories of aircraft engine components in depth maps and output a pixel-level predicted mask of the same size as the input image, providing structured visual input for subsequent component recognition, inspection, 3D reconstruction, or intelligent maintenance systems.
[0085] In an embodiment of the present invention, the trained deep convolutional neural network model is applied to unlabeled depth map images of aircraft engine components to achieve pixel-level semantic segmentation of the components in the target image.
[0086] Specifically, an unlabeled depth image of an aircraft engine component is first input. The image undergoes preprocessing operations consistent with the training phase, including normalization, resizing, format conversion and other processing steps to ensure that the input data remains consistent with the input during model training.
[0087] This preprocessed depth map is then fed into a trained deep convolutional neural network. The network performs forward inference on the input image, extracting spatial features layer by layer and performing pixel classification operations, ultimately outputting a predicted mask image that matches the original image size. In this predicted mask image, each pixel position is assigned a specific semantic category label, representing the component category corresponding to that pixel in the image, such as "background," "nozzle," "turbine blade," or "housing."
[0088] To facilitate practical applications, the model output results can be visualized or post-processed, including:
[0089] Map pixel labels to different colors to achieve color visualization of semantic segmentation results;
[0090] Perform edge extraction or morphological processing on the predicted mask to generate structure contours;
[0091] Export segmentation results to standard formats (such as PNG, TIFF, NumPy arrays, or COCO labels) for subsequent industrial tasks such as 3D modeling, CAD registration, defect detection, or robot guidance.
[0092] This model is suitable for depth images of aviation engine components of different models and structural forms. It does not rely on manual labeling or template matching mechanisms, has strong generalization and automatic processing capabilities, and is particularly suitable for large-scale, variable-viewing angle industrial visual recognition scenarios.
[0093] By applying the trained model to unlabeled depth maps, the present invention achieves high-precision, automated semantic segmentation of complex structural components of aircraft engines, effectively reducing the cost of manual participation, improving the efficiency of depth map information structuring, and providing high-quality visual basic data for downstream tasks such as intelligent inspection, assembly, and maintenance.
[0094] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0095] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0096] It should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. In addition, the character " / " herein generally indicates that the objects associated with each other are in an "or" relationship, but it may also indicate an "and / or" relationship, which can be understood by referring to the context. A person of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0097] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An intelligent segmentation method for depth images of aircraft engine components based on a deep convolutional network, characterized by: include: Acquiring depth image data of aircraft engine components and preprocessing the depth image; Generate component masks using traditional image processing algorithms in the first annotated image, construct an initial template mask, and establish image and label pairs using image annotation tools; Migrating the initial template mask to the target image to generate a candidate mask, and extracting the viewing angle difference ADI from the target image to measure the angle difference between the main direction vectors of the corresponding components in the template image and the target image; and the boundary deformation rate (BDR), which is used to measure the degree of geometric deformation between key points of the component contour; Determining the template registration error level based on the comprehensive analysis results of ADI and BDR, and executing a non-rigid alignment mechanism for images with a severe registration error level; The registered mask and depth map are fed into a deep convolutional neural network for training to generate a model for semantic segmentation of aerospace engine parts. The model is applied to unlabeled depth maps for pixel-level semantic segmentation of parts.
2. The method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network according to claim 1 is characterized in that: Generating an initial template mask through an image processing algorithm includes: performing Canny edge detection on the depth map to obtain boundary information; using a region growing algorithm to fill the closed boundary area to form an initial mask; performing morphological processing on the generated mask, including dilation, erosion, and hole filling; and importing the generated mask into an image annotation tool for manual review and refinement to generate image label pairs.
3. The method for intelligent segmentation of depth images of aircraft engine components based on deep convolutional networks according to claim 1 is characterized in that: The method for obtaining the viewing angle difference ADI is as follows: perform principal component analysis on the corresponding component contours in the template image and the target image to obtain the first principal component direction; define the first principal component direction vector as the component principal direction; let the principal direction vector in the template image be v1 and the principal direction vector in the target image be v2; calculate the angle θ between the two, and the expression is: Among them: ∥v1∥, ∥v2∥ respectively represent the modulus of the two vectors; the unit of θ is radians, which is converted to angle expression to calculate the viewing angle difference ADI, and the expression is: ADI is the visual angle difference in degrees, π is the circumference of a circle, and θ is the angle between the main directions of the component in the template and the target image.
4. The method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network according to claim 3 is characterized by: The method for obtaining the boundary deformation rate BDR is as follows: normalize the template mask and the target image mask to a unified scale, extract N corresponding key points from the contour; the key point in the template image is denoted as P i =(x i ,y i ), the corresponding key point in the target graph is Q i =(x′ i , y′ i ); for each key point, calculate the Euclidean distance from it to the contour centroid, which are: Among them C T 、C G The centroids of the contours in the template image and the target image are used to calculate the boundary deformation rate. The expression is: BDR is the boundary deformation rate, N is the number of key points; is the distance from the i-th key point to its respective centroid.
5. The method for intelligent segmentation of depth images of aircraft engine components based on deep convolutional networks according to claim 4 is characterized in that: The template registration error level is determined based on the comprehensive analysis results of ADI and BDR. For images with severe registration error levels, a non-rigid alignment mechanism is implemented, specifically including: The perspective difference and boundary deformation rate are converted into comprehensive feature vectors, which are used as inputs of the machine learning model. The machine learning model uses the template registration error score value label predicted by each set of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of all template registration error score value labels as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The template registration error score value is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
6. The method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network according to claim 5, characterized in that: The obtained template registration error score is compared with the preset threshold. If the template registration error score is greater than or equal to the preset threshold, it is considered that the current template registration error is serious and non-rigid alignment needs to be triggered; if the template registration error score is less than the preset threshold, it is considered that the registration error is slight and the affine registration result continues to be used.
7. The method for intelligent segmentation of depth images of aircraft engine components based on a deep convolutional network according to claim 6, characterized in that: The non-rigid alignment mechanism includes: Perform local feature point extraction on the component area in the template image and the target image respectively; suppose the key point set extracted in the template image is In the target graph, Calculate the descriptor similarity between the two and match several candidate point pairs A random sampling consensus algorithm is used for robust matching screening, which minimizes the geometric error between matching point pairs and retains the set of interior points that meet most structural constraints. The result after eliminating matching point pairs is used to construct the control point set for deformation estimation, which is recorded as: where p i is the control point in the template graph, p′ i is the corresponding control point in the target graph, and b is the total number of points; A non-rigid mapping relationship is constructed based on the matching point pairs, and a transformation function with minimum bending energy is constructed to achieve continuous deformation of the template coordinate system to the target image coordinate system. The transformation definition expression is: Where: (x, y) is any pixel in the template; U is the radial basis function; a1, a2, a3 and w i is the transformation parameter, which is solved by least square fitting; Apply the desired deformation field to the initial template mask image, and use bilinear interpolation and other methods to map the coordinates of all pixels in the mask to obtain the deformed mask: M′(x′, y′)=M(f -1 (x′, y′)); where M is the template mask, M′ is the mask after registration, and f -1 is the inverse deformation mapping.
8. The method for intelligent segmentation of depth images of aircraft engine components based on deep convolutional networks according to claim 7 is characterized in that: Inputting the mask and depth map into the deep convolutional neural network for training includes: determining the semantic segmentation network and constructing a training set, including the preprocessed depth map and the registered mask label map; using the cross entropy or Dice loss function as the supervision signal; applying image enhancement strategies, including rotation, scaling and affine perturbation; and iterative training until the loss converges.
9. The method for intelligent segmentation of depth images of aircraft engine components based on deep convolutional networks according to claim 8, characterized in that: Applying the trained model to unlabeled depth maps involves normalizing and resizing the unlabeled depth map; feeding the image into the trained model for forward reasoning, outputting a pixel-level prediction mask; and mapping different category labels in the mask to different colors to generate semantic segmentation visualizations.
Citation Information
Cited By
Visual segmentation model-based linear edge error measurement method for structural member
CN121837169A