Method and system for extracting and describing position features of multistage parts in two-dimensional image
By constructing a scene graph of multi-level components in a two-dimensional image, using node features and edge features to generate triple representations, combined with coarse detection and fine detection strategies, the problem of insufficient detection accuracy of functional anomaly detection in the prior art is solved, and efficient and robust component position and assembly structure recognition is achieved.
Patent Information
- Application Number
- CN202510326571.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to effectively detect functional abnormalities in industrial products, especially in complex contexts, and it is difficult to establish a multi-target component association relationship model, resulting in insufficient detection accuracy and robustness.
By constructing a scene graph of multi-level components in a two-dimensional image, using node features and edge features to generate triple representations, combining coarse detection, fine detection and hierarchical transmission strategies, a multi-level component spatial position relationship model is built to achieve efficient object detection and abnormal analysis.
It significantly improves the accuracy and robustness of detection, can identify the relative position of parts and the overall assembly structure, reduces false detection and missed detection, enhances the ability to perceive abnormal situations, and is suitable for intelligent detection in complex industrial environments.
Smart Images

Figure CN120235945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and particularly relates to a method and system for extracting and describing multi-level component position features in a two-dimensional image. Background Art
[0002] The intelligent detection of industrial product quality is a hot field in current artificial intelligence research and application. The anomalies detected by visual visible light mainly fall into appearance anomalies and functional anomalies. Appearance anomalies include the analysis of continuous features such as scratches, cracks, depressions, color, and surface roughness, as well as the quantification of features such as dimensional deviation. Functional anomalies include functional failures (the product cannot work properly or perform its designed functions, which may be caused by circuit failures, mechanical component damage, etc.) and performance degradation (the performance of the product does not meet the design standards or expected levels, which may be caused by material aging, wear, etc.). Thus, it can be seen that there are a variety of industrial product anomalies. The existing technologies mainly focus on the research and application of appearance anomaly detection, and most of the relevant data sets are also for appearance anomaly detection, while there is a lack of public data sets and relevant research on functional anomaly detection.
[0003] Early intelligent industrial anomaly detection mainly relied on manual feature extraction and rule-based defect detection methods. Manual image feature extraction processes and analyzes images to extract key information in the images, such as edges, corners, textures, etc., for identifying the shape, size, position, etc. of objects. Common image feature extraction methods include image segmentation, edge detection, geometric features, and data statistical features. After extracting a large number of features, feature selection and optimization are required to improve the performance and generalization ability of the detection model. For example, methods such as statistical control charts and PCA (Principal Component Analysis) can detect possible anomalies through the analysis of data in the production process. However, these methods can often only handle relatively single anomaly patterns and rely on the feature engineering of historical defect data. In the application of machine learning, classical algorithms such as SVM (Support Vector Machine) and KNN (K-Nearest Neighbor Algorithm) are used to identify defects in products. These methods can automate the anomaly detection process to a certain extent, but their performance depends on the quality of feature extraction and data preprocessing, and their ability to process large-scale and diverse data is limited.
[0004] After 2006, with the development of convolutional neural networks, deep learning has gradually been introduced into industrial anomaly detection. Deep learning is a deep neural network structure with multiple convolutional layers. By learning the features of input data, it forms more abstract high-level feature representations from low-level features, expressing data in ways such as vectors and feature maps, thereby improving the effectiveness of deep learning algorithms. Based on the powerful learning ability and feature extraction ability of deep learning in large amounts of data, many researchers have tried to apply deep learning technology to product defect detection to improve product quality. Currently, the main research area focuses on the surface defect detection of individual industrial products, that is, texture defects.
[0005] Functional defects are mainly caused by errors in the overall structure of the product, including deformation, dislocation, defect, and contamination, etc. For example, the bending of a wire, the edge defect of a diode or being in the wrong position, etc. The overall structure of the product is more complex, and there is background interference outside the product. Under different backgrounds, the weakness degrees of different types of defects are different, and even between different instances of the same type of defect, the visibility may vary greatly. Therefore, functional defect detection requires deeper logical reasoning, that is, creating a relational model that associates part features with the structure of components. However, there is little research on existing functional defect detection. How to combine the actual background needs of the industry, establish a multi-object component association relationship model, and implement an intelligent reasoning system for functional anomalies will have good application value.
[0006] Chinese patent document CN110533725A discloses a detection method for large-scale different components of a high-speed rail catenary. Although this method uses a clustering method to reduce the influence of scale differences of different components, its core is still based on the framework of object detection (InceptionV2 + SSD), mainly optimizing the accuracy of object detection, and not constructing the topological or hierarchical structural relationship between components.
[0007] Chinese patent document CN117994237A discloses a detection method for the assembly state of an aeroengine. This method uses a neural network for classification training and uses the correct assembly state diagram as verification data to detect the assembly process. However, this method only relies on image classification and detection technologies and cannot determine whether a certain part should exist and whether the relative position with other components is reasonable. Summary of the Invention
[0008] The purpose of the present invention is to propose a method and system for extracting and describing the position features of multi-level components in a two-dimensional image. By constructing a scene graph of multi-level components in a two-dimensional image and generating a triple representation using node features and edge features, the spatial relationship between components can be accurately extracted and described, realizing efficient object detection and anomaly analysis.
[0009] According to the first aspect of the embodiments of the present disclosure, a method for extracting and describing multi-level component position features in a two-dimensional image is provided, including the following steps:
[0010] Obtain an industrial component two-dimensional image dataset and process it from an abstract level and a specific technical level; the abstract level refers to the standardized definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is the preprocessing of industrial component two-dimensional images;
[0011] For each preprocessed industrial component two-dimensional image, use an object detection algorithm to identify and locate each object in the image, and obtain its position and name information;
[0012] Based on the object detection results, construct a multi-level component spatial position relationship model, and describe the multi-level structure of industrial components by generating node features and edge features.
[0013] According to the second aspect of the embodiments of the present disclosure, a system for extracting and describing multi-level component position features in a two-dimensional image is provided, including:
[0014] An image processing module that obtains an industrial component two-dimensional image dataset and processes it from an abstract level and a specific technical level; the abstract level refers to the standardized definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is the preprocessing of industrial component two-dimensional images;
[0015] An object detection module that, for each preprocessed industrial component two-dimensional image, uses an object detection algorithm to identify and locate each object in the image, and obtain its position and name information;
[0016] A description module that, based on the object detection results, constructs a multi-level component spatial position relationship model, and describes the multi-level structure of industrial components by generating node features and edge features.
[0017] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program running on the memory, and when the processor executes the program, it implements the method for extracting and describing multi-level component position features in a two-dimensional image as described above.
[0018] According to the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method for extracting and describing multi-level component position features in a two-dimensional image as described above.
[0019] The above technical solution adopted by the present invention has the following advantages compared with the prior art: 1) By adopting the strategy of combining rough detection, fine detection and hierarchical transmission, the present invention realizes hierarchical detection of the target, gradually refining the recognition from the whole to the local, effectively reducing the interference of complex backgrounds and improving the stability of detection. In the fine detection stage, the detection range of the third-level components is restricted by using hierarchical dependence relationships, reducing the occurrence of false detections and missed detections. At the same time, combining the prior structural information of the components ensures that each second-level component only detects the third-level components that it may contain, making the detection results more in line with the actual situation of industrial assembly, thus significantly improving the accuracy and reliability of recognition.
[0020] 2) After detection, the present invention constructs a multi-level scene graph, which can not only identify the positions of individual components, but also understand their relative relationships in the overall assembly structure. By extracting spatial features such as relative positions, angles, and distances, it has a stronger ability to recognize the assembly structure. It not only improves the understanding of the spatial positions of components, but also can be used for assembly consistency inspection, defect detection, and anomaly recognition, significantly enhancing the intelligent level and the ability to perceive abnormal situations.
[0021] 3) The hierarchical consistency correction mechanism proposed by the present invention can automatically adjust when the detection results do not conform to the hierarchical relationship. For example, if a second-level component lacks the third-level components that it should contain, a re-inspection can be triggered, thereby reducing omissions in assembly detection. This optimization strategy makes the present invention have stronger robustness and adaptability in complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application.
[0023] Figure 1 is a flowchart of the method for extracting and describing the position features of multi-level components in a two-dimensional image;
[0024] Figure 2 is a schematic diagram of the processing of the two-dimensional image dataset of industrial components;
[0025] Figure 3 is a schematic diagram for defining the first-level component;
[0026] Figure 4 is a schematic diagram for defining the first-level component;
[0027] Figure 5 is a flowchart of the target detection;
[0028] Figure 6 is a schematic diagram for generating the spatial position relationship. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0030] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.
[0031] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0032] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Similarly, it should be noted that each block in the flowchart and / or block diagram, as well as the combinations of blocks in the flowchart and / or block diagram, may be implemented using a dedicated hardware-based system for performing the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions.
[0033] Embodiment 1:
[0034] As Figure 1 shown, this embodiment provides a method for extracting and describing multi-level component position features in a two-dimensional image, including the following steps:
[0035] S1. Obtain an industrial component two-dimensional image dataset, as Figure 2 shown, and process it from an abstract level and a specific technical level; the abstract level refers to standardizing the definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is to preprocess the industrial component two-dimensional image to ensure that subsequent object detection and relationship modeling can be carried out under a unified data format and quality;
[0036] Specifically, the present invention divides each target to be detected into three constituent levels: the first-level component (top layer) corresponds to the overall target to be detected, the second-level component corresponds to the local component obtained after the first decomposition of the target to be detected; the third-level component corresponds to the basic component obtained after the second decomposition of the second-level component (local component) and cannot be further divided. More levels can be divided according to the assembly situation of the specific target.
[0037] S1.1 Based on the collected two-dimensional image data of industrial components, process it at the abstract level, that is, define the hierarchical structure;
[0038] S1.1.1 First-level component (the whole image scene):
[0039] The first-level component corresponds to the whole image of the product to be detected collected by each camera. Therefore, the difference between first-level components depends on the hardware deployment. This is because during the quality inspection process, multiple cameras are often used to collect images of the same product from different angles to ensure a comprehensive inspection of the product. However, the target components collected at each angle may have different composition forms due to the positional relationship. As Figure 3 shown.
[0040] The whole image collected by cameras at different spatial positions is used as the first-level component, denoted as where N1 is the total number of cameras (division of the target standard area), r n corresponds to the image collected by the nth camera. The first-level component is identified by the camera number n, n ∈ N1.
[0041] S1.1.2 Second-level component (main component):
[0042] After decomposing each original image of the first-level component once according to the physical structure, the second-level component is obtained. This is often the core area division of the entire target to be detected. Assuming that there are a total of N2 types of second-level components, a certain first-level component r is composed of specific m m second-level components, denoted as: k
[0043]
[0044] where represents the i-th second-level component of the m-th first-level component, m k ∈ {1, 2,..., N2}, which also means As Figure 4 shown.
[0045] S1.1.3 Third-level component (non-separable component):
[0046] The tertiary components are further disassembled from each secondary component into components that cannot be further divided. Suppose there are N3 types of tertiary components in total. The relationship between the secondary components and the tertiary components is described as follows:
[0047]
[0048] Among them, represents the j-th component detected in the n-th secondary region of the image captured by the m-th camera, where j = 1, 2, …, N3; it also indicates that that is, the tertiary components are nested within the secondary components, and the secondary components are nested within the primary components.
[0049] S1.2 For the convenience of subsequent object detection, specific technical-level preprocessing will also be performed on the two-dimensional image data of industrial components.
[0050] S1.2.1 Image size normalization
[0051] Since the images in the industrial scenario may come from different acquisition devices and have different resolutions and scales, directly inputting them into the model will lead to inconsistent feature distributions. Therefore, all images are uniformly adjusted to the same size (H0, W0):
[0052] I←resize(I, H0, W0)
[0053] where I is the original image, and the unified size is adjusted to avoid the decrease in detection accuracy caused by different resolutions.
[0054] S1.2.2 Brightness and illumination equalization
[0055] In the industrial environment, the illumination conditions vary greatly. For example, the illumination conditions are different at different shooting times, and shadow occlusion affects the detection accuracy. Therefore, the present invention uses histogram equalization (HE) to adjust the image brightness distribution:
[0056] I←HE(I)
[0057] where HE(·) is the histogram equalization operation, which makes the brightness distribution of the image more balanced and reduces the detection deviation caused by illumination changes.
[0058] S1.2.3 Noise removal
[0059] In the actual scenario, the image may include noise (such as dust, light spots, etc.), which affects the accuracy of the object detection model. Therefore, the present invention uses Gaussian filtering (Gaussian Blur) for denoising:
[0060] I←G σ *I
[0061] Among them, G σ is the Gaussian filter kernel, and σ controls the degree of smoothing.
[0062] S2. For each preprocessed two-dimensional industrial component image, use the object detection algorithm to identify and locate each object in the image, and obtain its position and name information. During the detection process, in order to avoid repeated detection of the same object, non-maximum suppression (NMS) is used to screen candidate boxes, and the choice strategy of candidate boxes is further optimized by combining DIoU (Distance-IoU);
[0063] S2.1 Detect components at different levels, and the hierarchical object detection algorithm is as Figure 5 shown.
[0064] S2.1.1 Coarse detection: According to the relationship between the first-level components and the second-level regions, use the object detection model to scan the image as a whole, and quickly identify the possible second-level component regions.
[0065] Let the detection model be f θ , its input is the preprocessed image I, and the output is the detection box B and its category c:
[0066] (B, c) = f θ (I)
[0067] Among them, B represents the detection box, including the center coordinates width and height c represents the corresponding category label. For each second-level component, it belongs to its first-level component region, denoted as
[0068] S2.1.2 Fine detection: On the basis of the coarse detection, further analyze and process the identified regions to accurately identify the category and position of the components. Crop the component images identified in the coarse detection, and also use the detection model f θ to detect the cropped image, and extract the information of the third-level components:
[0069] (B, c) = f θ (I)
[0070] Among them, B represents the detection box, including the center coordinates width and height c represents the corresponding category label. For each third-level component, it belongs to its second-level component region, denoted as
[0071] In the fine detection stage, the number of detection boxes increases, and the same target may be covered by multiple candidate boxes. To ensure the uniqueness and accuracy of the detection results, non-maximum suppression (NMS) is used to remove overlapping candidate boxes, and DIoU is combined to further optimize target screening.
[0072] S2.1.3 Hierarchical transmission: During the fine detection process, the information of the identified secondary components can be used to assist in identifying the tertiary components they contain, so as to improve the detection accuracy and hierarchical consistency.
[0073] Since each type of secondary component usually contains a fixed type and quantity of tertiary components, when the detection results do not conform to the hierarchical relationship, a recheck can be triggered through the consistency check mechanism:
[0074]
[0075] Among them, represents the set of categories of tertiary components that are inherently related to the secondary component . If a secondary component lacks the expected tertiary components, the area will be re-detected and the detection results will be adjusted to reduce detection omissions.
[0076] For example, the secondary component bracket fixation includes the tertiary components bolt and nut, and the secondary component wire clamp fixation includes the tertiary component buckle. By constructing a structural dependency relationship, the interference of irrelevant categories can be effectively reduced and the detection efficiency can be improved.
[0077] S2.1.4 Post-processing and optimization:
[0078] During the target detection process, the same component may be detected by multiple overlapping candidate boxes, resulting in repeated predictions. To screen out the optimal target detection box, the present invention uses non-maximum suppression (NMS, Non-Maximum Suppression) to screen candidate boxes according to the IoU (Intersection over Union) and remove low-quality detection results.
[0079] In addition, to further optimize the NMS strategy, the present invention introduces the DIoU (Distance-IoU) metric. DIoU-NMS can more accurately distinguish adjacent targets, reduce false eliminations, and thus improve the detection integrity of small components. In addition, the optimized NMS strategy can ensure the reasonable selection of detection boxes, make the detection results more consistent, and improve the hierarchical relationship matching degree of components.
[0080] S3. Based on the target detection results, a multi-level component spatial position relationship model is constructed. By generating node features and edge features, the multi-level structure of industrial components is described, as Figure 6 shown.
[0081] S3.1 Multi-level Node Feature Construction
[0082] For the detected components, they are classified according to their hierarchical levels into:
[0083] ● Primary Component Node (Overall scene),
[0084] ● Secondary Component Node (Main component),
[0085] ● Tertiary Component Node (Detail component).
[0086] Each node is obtained by extracting high-dimensional visual features from its detection region in the image through pre-trained CNN, denoted as:
[0087]
[0088] where l ∈ {1, 2, 3} represents different hierarchical levels, and d v denotes the node feature dimension using symbols.
[0089] Each node v i can be represented as a matrix containing feature values, denoted as:
[0090] v i = [x i , y i , class i , confidence i
[0091] S3.2 Edge Feature Construction and Multi-level Relationship Modeling
[0092] To fully capture the spatial relationships and hierarchical dependencies between components, the present invention constructs two types of edge features:
[0093] (1) Intra-level Edge Feature
[0094] For two component nodes within the same hierarchical level and define their central coordinates as (x i , y i ) and (x j , y j ). The intra-level edge feature is defined as:
[0095] Δx ij = x j - x i , Δy ij = y j - y i
[0096]
[0097] Thus, we have:
[0098]
[0099] These edge features describe the relative position changes between components at the same level, such as the spatial sorting relationship between components at the same level, for example, the relative distribution of multiple nuts and multiple support parts.
[0100] (2) Cross-level edge features
[0101] A first-level component (such as an overall bracket) usually contains multiple second-level components (such as insulators, suspension devices, etc.). Let a certain first-level node and the second-level nodes it contains respectively have central coordinates and At the same time, let the width and height of the detection box of the first-level node be Define the relative position normalized from the first level to the second level:
[0102]
[0103] And calculate the normalized distance:
[0104]
[0105] Finally, the cross-level edge feature from the first level to the second level is defined as:
[0106]
[0107] This definition can capture the relative position of the second-level components in the entire scene (first-level components) and eliminate the influence of scale differences, making the model more stable. For example, the cross-level relationship between the bracket (first-level component) and the insulator (second-level component) can help detect whether the suspension device is correctly connected to the bracket.
[0108] Similarly, a second-level component (such as an insulator) contains multiple third-level components (such as bolts, nuts), and the relational formula is as follows:
[0109]
[0110] Finally, the cross-level edge feature from the second level to the third level is defined as:
[0111]
[0112] This definition enables the edge features to not only reflect the absolute distance but also consider the influence of the secondary component dimensions, depicting the relative position of the tertiary components within the parent nodes. For example, in the combination of a nut and a support rod, this edge feature can describe whether the nut is correctly positioned on the support rod.
[0113] S3.3 Triplet Representation and Scene Graph Generation
[0114] Based on the construction of the above node and edge features, the present invention expresses the spatial relationships in the graph in the form of triplets, that is, each edge corresponds to a triplet. For components at the same level:
[0115]
[0116] For components across levels:
[0117]
[0118] Finally, the entire component structure can be represented as a set of scene graphs containing all triplets:
[0119]
[0120] Each row (v i , e ij , v j ) represents the relationship between components i and j, that is, this row describes the spatial relationship between the two components and their respective features.
[0121] This modeling method ensures that the entire industrial scene is encoded into a structured scene graph, in which the spatial relationships of all components are quantified and stored, providing a basis for subsequent reasoning and detection. Through this method, the present invention can not only express the hierarchical information of components but also utilize technologies such as graph neural networks for efficient spatial relationship reasoning and anomaly detection.
[0122] Existing methods usually have difficulty effectively modeling the spatial dependence relationships between components at different levels, resulting in a decline in detection accuracy and limited anomaly recognition ability. The present invention constructs a multi-level scene graph, combines node features and edge features to generate triplet representations, accurately depicts the hierarchical relationships and spatial layouts between components, thereby improving detection accuracy, enhancing the model generalization ability, and supporting anomaly detection and assembly quality assessment, and is applicable to intelligent detection tasks in complex industrial environments.
[0123] Embodiment 2:
[0124] This embodiment provides a system for extracting and describing the position features of multi-level components in a two-dimensional image, including:
[0125] An image processing module that obtains a two-dimensional image dataset of industrial components and processes it at an abstract level and a specific technical level. The abstract level refers to the standardized definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is the preprocessing of two-dimensional images of industrial components.
[0126] A target detection module that, for each preprocessed two-dimensional image of an industrial component, uses a target detection algorithm to identify and locate each target in the image and obtain its position and name information.
[0127] A description module that, based on the target detection results, constructs a multi-level component spatial position relationship model and describes the multi-level structure of industrial components by generating node features and edge features.
[0128] Example 3:
[0129] An electronic device includes a memory, a processor, and a computer program running on the memory. When the processor executes the program, it implements the above method for extracting and describing the position features of multi-level components in a two-dimensional image, including:
[0130] Obtain a two-dimensional image dataset of industrial components and process it at an abstract level and a specific technical level. The abstract level refers to the standardized definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is the preprocessing of two-dimensional images of industrial components.
[0131] For each preprocessed two-dimensional image of an industrial component, use a target detection algorithm to identify and locate each target in the image and obtain its position and name information.
[0132] Based on the target detection results, construct a multi-level component spatial position relationship model and describe the multi-level structure of industrial components by generating node features and edge features.
[0133] Example 4:
[0134] A computer-readable storage medium stores a computer program that, when executed by a processor, implements the above method for extracting and describing the position features of multi-level components in a two-dimensional image, including:
[0135] Obtain a two-dimensional image dataset of industrial components and process it at an abstract level and a specific technical level. The abstract level refers to the standardized definition according to the physical structure of industrial components at the first, second, and third levels, and the specific technical level is the preprocessing of two-dimensional images of industrial components.
[0136] For each preprocessed two-dimensional image of industrial parts, use the object detection algorithm to identify and locate each object in the image, and obtain its position and name information;
[0137] Based on the object detection results, construct a multi-level spatial position relationship model of parts. By generating node features and edge features, describe the multi-level structure of industrial parts.
[0138] Those skilled in the art should understand that the above-mentioned modules or steps of the present disclosure can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present disclosure is not limited to any specific combination of hardware and software.
[0139] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0140] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A method for extracting and describing multi-level component position features in a two-dimensional image, characterized in that: The following steps are involved: Obtain a two-dimensional image data set of industrial parts and process it from an abstract level and a specific technical level; the abstract level refers to the standardized definition of the first, second and third levels according to the physical structure of the industrial parts, and the specific technical level is to pre-process the two-dimensional images of the industrial parts; For each preprocessed two-dimensional image of industrial parts, use the target detection algorithm to identify and locate each target in the image and obtain its location and name information; Based on the target detection results, a multi-level component spatial position relationship model is constructed, and the multi-level structure of industrial components is described by generating node features and edge features.
2. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 1, characterized in that: According to the physical structure of industrial parts, standardization is carried out at the first, second and third levels. The specific implementation methods are as follows: The entire two-dimensional image of the industrial component obtained by cameras at different spatial positions is taken as the first-level component and represented as Where N1 is the total number of cameras, r n Corresponding to the image captured by the nth camera, the first-level component is identified by the camera number n, n∈N1; Each first-level component original image After decomposing the physical structure once, we get the second-level components. Assuming that the second-level components include N2 types in total, a first-level component r m By specific m k It is composed of two secondary components, expressed as: in represents the mth first-level component and the i-th second-level component, m k ∈{1,2,…,N2}; The third-level parts are further decomposed into indivisible parts by splitting each second-level component. Assuming that there are N3 third-level parts in total, the relationship between the second-level components and the third-level parts can be described as follows: in It represents the j-th component detected in the n-th secondary component area of the image captured by the m-th camera, where j=1, 2, ..., N3.
3. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 1, characterized in that: The target detection algorithm is used to identify and locate each target in the image: First, perform a rough detection, specifically: Assume the detection model is f θ , whose input is the preprocessed two-dimensional image I of industrial parts, and whose output is the detection box B and its category c: (B,c)=f θ (I) in, B represents the detection box, including the center coordinates Width and Height c represents the corresponding category label; for each secondary component, it belongs to its primary component area, denoted as Then, fine inspection is performed, specifically: the two-dimensional image of the industrial parts identified in the coarse inspection is cropped, and the inspection model f is also used θ Detect the cropped image and extract the information of the three-level components: (B,c)=f θ (I) in, B represents the detection box, including the center coordinates Width and Height c represents the corresponding category label; each third-level component belongs to the second-level component area, denoted as In the fine detection stage, the number of detection frames increases, and the same target will be covered by multiple candidate frames. The non-maximum suppression strategy NMS is used to remove overlapping candidate frames. The DIoU indicator is introduced into the non-maximum suppression strategy NMS to improve the detection integrity of fine industrial parts.
4. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 3, characterized in that: Each secondary component Contains three-level components of fixed types and quantities, so when the inspection results do not conform to the hierarchical relationship, a re-inspection is triggered through the consistency check mechanism: in, Represents secondary component An inherently associative collection of three-level component categories.
5. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 1, characterized in that: Edge features include edge features at the same level, which are obtained as follows: For two industrial component nodes in the same level and Define the center coordinates (x i ,y i ) and (x j ,y j ), then the edge features at the same level Defined as: Δx ij =x j -x i ,Δy ij =y j -y i Thus:
6. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 1, characterized in that: Edge features include cross-level edge features, which are obtained as follows: The first-level component contains multiple second-level components; if a certain level node The secondary nodes it contains With center coordinates and At the same time, the detection box width and height of the first-level node are set to Define the relative position of primary to secondary normalization: And calculate the normalized distance: The cross-level edge feature from level 1 to level 2 is defined as: Similarly, a secondary component contains multiple third-level components, and the relationship formula is as follows: The cross-level edge feature from level 2 to level 3 is defined as:
7. The method for extracting and describing multi-level component position features in a two-dimensional image according to claim 1, characterized in that: The multi-level structure of industrial parts is described as follows: The spatial relationship in the two-dimensional image of industrial parts is expressed as a triplet, that is, each edge corresponds to a triplet. For industrial parts at the same level: For cross-level industrial components: The entire industrial parts structure is represented as a scene graph set containing all triples: Each row (v i ,e ij ,v j ) represents the relationship between parts i and j, that is, this row describes the spatial relationship between the two parts and their respective characteristics.
8. A system for extracting and describing multi-level component position features in a two-dimensional image, characterized in that: include: The image processing module obtains a two-dimensional image data set of industrial parts and processes it from an abstract level and a specific technical level; the abstract level refers to the standardized definition of the first, second and third levels according to the physical structure of the industrial parts, and the specific technical level is to pre-process the two-dimensional images of the industrial parts; The target detection module uses the target detection algorithm to identify and locate each target in each preprocessed two-dimensional image of industrial parts and obtains its location and name information; The description module builds a multi-level component spatial position relationship model based on the target detection results, and describes the multi-level structure of industrial components by generating node features and edge features.
9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the method for extracting and describing multi-level component position features in a two-dimensional image is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a method for extracting and describing multi-level component position features in a two-dimensional image is implemented.
Citation Information
Patent Citations
High-speed railway overhead line system multi-part positioning method based on structural reasoning network
CN110533725A
Method for detecting assembly state of aero-engine
CN117994237A