AR-assistance-based frame support installation condition detection method, computer system and computer readable storage medium
By introducing a frame bracket detection model with a bilateral attention guidance module in the frame bracket installation detection, the problem of difficulty in taking into account global information and detailed information in the existing technology is solved, and accurate detection and visual guidance on the installation of the frame bracket are achieved.
Patent Information
- Application Number
- CN202510011316.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-27
AI Technical Summary
The existing identification model is difficult to take into account the global information and detailed information required for the classification and positioning of the frame bracket, which makes it difficult to meet the accurate detection requirements for the installation of the frame bracket.
The AR-assisted frame bracket installation detection method is used to perform image positioning and classification detection through the frame bracket detection model, and the bilateral attention guidance module is used to guide the details and semantic branches to each other, generating richer feature maps for accurate detection.
The detection results are displayed by AR equipment, and the visual guidance effect of the installation results of the frame bracket is improved, and the global information and detailed information required for classification and positioning can be taken into account, meeting the accurate detection requirements of the installed frame bracket.
Smart Images

Figure CN120047863A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent detection of frame installation, and particularly relates to a method for detecting the installation condition of a frame bracket based on AR assistance, a computer system, and a computer-readable storage medium. Background Art
[0002] With the rapid development of the global manufacturing industry, especially in the field of automobile manufacturing, the installation process of frame brackets has become increasingly complex. The traditional installation method relies on manual inspection or simple tools, which has obvious deficiencies: on the one hand, workers need to frequently view drawings or installation instructions, which is time-consuming and laborious, and it is easy to make mistakes during the interpretation process, resulting in improper installation or incorrect installation of the brackets; on the other hand, the diversity of frame types and bracket models, as well as the complexity of the workshop environment (such as insufficient light, narrow space, or noise interference), further increases the difficulty of installation for workers and the challenge of abnormal detection. Especially in the case where abnormal installations (such as missing installation, incorrect installation, loosening) are not detected in time, it may seriously affect subsequent assembly and product quality.
[0003] In the prior art, for the installation structure, machine vision methods such as the YOLO model can be used to identify the installation condition. For example, the Chinese invention patent with the application publication number CN114782778A inputs a state picture into an identification model to judge the state of the aviation fan rotor in the state picture, determine which assembly process the current aviation fan rotor is in, and compare the current assembly process with the judged assembly process to judge whether there is an assembly error; however, existing identification models often have difficulty taking into account the global information and detailed information required for classification and positioning respectively, resulting in difficulty in meeting the accuracy requirements for classifying and positioning the currently installed frame brackets. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting the installation condition of a frame bracket based on AR assistance, a computer system, and a computer-readable storage medium, which is used to solve the problem that existing identification models often have difficulty taking into account the global information and detailed information required for classification and positioning respectively, resulting in difficulty in meeting the accuracy requirements for classifying and positioning the currently installed frame brackets.
[0005] To achieve the above purpose, the present invention provides a method for detecting the installation condition of a frame bracket based on AR assistance, obtaining a detection result for positioning and classifying the frame bracket in the image through a frame bracket detection model; judging the installation condition of the frame bracket according to the detection result, and displaying the detection result and the judgment result in an AR device; the construction method of the frame bracket detection model includes:
[0006] 1) Input the frame support image samples in the training set into the model. Select an initial feature map obtained by feature extraction through the backbone network in the model, and input it into the detail branch and the semantic branch respectively;
[0007] 2) The feature maps in the detail branch and the feature maps in the semantic branch are input into the bilateral attention guidance module to obtain the outputs for the detail branch and the semantic branch respectively, so as to generate new feature maps in the detail branch and the semantic branch respectively; Repeat 2) until the number of times of inputting the corresponding feature maps into the bilateral attention guidance module reaches the set number of times;
[0008] 3) Fuse the finally obtained feature maps in the detail branch and the semantic branch. The fusion result passes through the detection head to obtain the detection result; Update the parameters of the model through the corresponding loss function, and repeat 1)-3) until stopping to obtain the constructed model;
[0009] The processing method of the bilateral attention guidance module for the input feature maps includes: After processing the feature maps in the input detail branch to the same scale as the feature maps in the input semantic branch, through spatial attention processing, and then fusing with the feature maps in the input semantic branch to obtain the output of the bilateral attention guidance module to the semantic branch; The feature maps in the input semantic branch are processed through channel attention, and then fused with the feature maps in the input detail branch to obtain the output of the bilateral attention guidance module to the detail branch.
[0010] Further, the method of processing through spatial attention and then fusing with the feature maps in the input semantic branch to obtain the output of the bilateral attention guidance module to the semantic branch includes:
[0011] Divide the channels of the processed feature maps into a set number of groups, and each group generates a spatial attention map through spatial attention; Divide the channels of the feature maps in the input semantic branch into the same set number of groups, and each group is fused with the spatial attention map generated by the corresponding group of channels of the processed feature maps to obtain the output of the bilateral attention guidance module to the semantic branch.
[0012] Further, the method of fusing the finally obtained feature maps in the detail branch and the semantic branch includes:
[0013] Select an initial feature map obtained by feature extraction through the backbone network in the model or a feature map in the detail branch, and segment it through the boundary segmentation head to obtain a rough boundary map; Fuse the rough boundary map with the finally obtained feature maps in the detail branch to obtain a boundary feature map; After processing the boundary feature map through the self-attention mechanism, then fuse it with the result obtained by processing the finally obtained feature maps in the semantic branch, or fuse it with the finally obtained feature maps in the semantic branch to obtain the fusion result.
[0014] Further, the method for processing the feature map in the finally obtained semantic branch includes: after fusing the feature map in the finally obtained semantic branch with the feature map in the finally obtained detail branch, sequentially processing it through the channel attention and self-attention mechanisms.
[0015] Further, the method for obtaining the feature map in the finally obtained semantic branch includes:
[0016] The newly generated feature map in the semantic branch obtained according to the output of the bilateral attention guidance module to the semantic branch for the last time in 2) is processed through each average pooling layer and the global average pooling layer with different pooling kernel scales respectively, and then the processing results of the global average pooling layer and each average pooling layer are fused with the feature map to obtain the feature map in the finally obtained semantic branch.
[0017] Further, the number of the average pooling layers is not less than 4, and the method for fusing the processing results of the global average pooling layer and each average pooling layer with the feature map includes:
[0018] The processing result of the average pooling layer with the smallest pooling kernel scale is respectively fused with the processing results of the two average pooling layers with the second smallest and the third smallest pooling kernel scales to obtain the second fusion feature and the third fusion feature; the result obtained by respectively fusing the third fusion feature with the processing results of other average pooling layers with larger pooling kernel scales and the global average pooling layer is further fused with the processing result of the average pooling layer with the smallest pooling kernel scale, the second fusion feature and the third fusion feature, and the result of this joint fusion is fused with the feature map.
[0019] Further, the method for judging the installation situation of the frame bracket according to the detection result includes:
[0020] Comparing the detection result of positioning the frame bracket image with a preset installation position template to judge whether the installation position of the frame bracket is correct;
[0021] Comparing the detection result of classifying the frame bracket image with a preset installation category template to judge whether the installation category of the frame bracket is correct.
[0022] Further, the method for displaying the detection result and the judgment result in the AR device includes:
[0023] Mapping the detection frame corresponding to the detection result of positioning the frame bracket in the image to the field of view of the AR device, and superimposing and displaying the category name corresponding to the detection result of classifying the frame bracket in the image with the detection frame in the AR device; using the color corresponding to the judgment result as the color for displaying the detection frame in the field of view of the AR device.
[0024] Furthermore, the method of mapping the detection box into the field of view of the AR device includes:
[0025] Collect the three-dimensional point cloud data of the current scene; use the SLAM algorithm to estimate the pose matrix of the AR device in the world coordinate system of the current scene in real time through the continuously collected camera image data and the three-dimensional point cloud data of the current scene; convert the coordinates of the detection box to the world coordinate system of the current scene through the pose matrix; and then convert the coordinates of the detection box in the world coordinate system of the current scene to the rendering coordinate system of the AR device through the conversion matrix from the AR device rendering coordinate system to the world coordinate system, so as to map the detection box into the field of view of the AR device.
[0026] The above technical solution of the present invention provides a brand-new AR-assisted detection method for the installation situation of the vehicle frame bracket, and its beneficial effects include:
[0027] The detection results of positioning and classifying the vehicle frame bracket by the AR device and the judged installation situation of the vehicle frame bracket are displayed, which can effectively improve the visualization effect of guiding the installation result of the vehicle frame bracket. At the same time, not only a detail branch with a smaller downsampling rate is adopted to fully capture and preserve detail features, and a semantic branch with a larger downsampling rate is adopted to fully capture the overall context information in the image, but also a bilateral attention guidance module is adopted to achieve the effect of mutual guidance between the features in the detail branch and the features in the semantic branch; specifically, the bilateral attention guidance module uses the feature map in the detail branch after enhancing the detail features through spatial attention to fuse with the input feature map in the semantic branch as the output to the semantic branch, achieving the effect that the detail branch guides the semantic branch to learn the detail features lacking in the semantic branch; similarly, the bilateral attention guidance module also uses the feature map in the semantic branch after enhancing the global features through channel attention to fuse with the input feature map in the detail branch as the output to the detail branch, achieving the effect that the semantic branch guides the detail branch to learn the global features lacking in the detail branch. After being processed by the bilateral attention guidance module, the new feature map in the detail branch can contain the global information learned from the semantic branch, and the new feature map in the semantic branch can contain the detail information learned from the detail branch, effectively utilizing the complementary information of the two branches, making the detail information and global information in the finally obtained feature map for obtaining the detection result more abundant, and the fusion effect of the two types of information is also better, being able to take into account the global information and detail information required for classification and positioning respectively, so as to meet the accuracy requirements for the classification and positioning of the installed vehicle frame bracket.
[0028] The present invention also provides a computer system, including a processor configured to execute executable program instructions for implementing the above-described method for detecting the installation condition of a vehicle frame bracket based on AR assistance.
[0029] The computer system of the present invention can achieve the same beneficial effects as the above-described method for detecting the installation condition of a vehicle frame bracket based on AR assistance.
[0030] The present invention also provides a computer-readable storage medium storing computer program instructions for implementing the above-described method for detecting the installation condition of a vehicle frame bracket based on AR assistance when the computer program instructions are executed.
[0031] The computer-readable storage medium of the present invention can achieve the same beneficial effects as the above-described method for detecting the installation condition of a vehicle frame bracket based on AR assistance. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention;
[0033] Figure 2 is a schematic structural diagram of a vehicle frame bracket detection model constructed in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention;
[0034] Figure 3 is a schematic structural diagram of a bilateral attention guidance module in the vehicle frame bracket detection model constructed in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention;
[0035] Figure 4 is a schematic structural diagram of a boundary perception module in the vehicle frame bracket detection model constructed in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention;
[0036] Figure 5 is a schematic structural diagram of a semantic perception module in the vehicle frame bracket detection model constructed in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention;
[0037] Figure 6 is a schematic structural diagram of a multi-scale pyramid pooling module in the vehicle frame bracket detection model constructed in an embodiment of the method for detecting the installation condition of a vehicle frame bracket based on AR assistance according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0039] Embodiment of AR-assisted Detection Method for Installation Condition of Frame Bracket
[0040] This embodiment provides a technical solution for an AR-assisted detection method for the installation condition of a frame bracket. This method uses AR technology to display the detection results and installation conditions for positioning and classifying the frame bracket, optimizes the visualization guidance effect for the installation of the frame bracket, and for the detection model for the installation condition of the frame bracket used to achieve positioning and classification, on the basis of setting a detail branch for saving local detail information of weak defects and tiny defects and a semantic branch focusing on extracting global context information to capture high-level abstract features in the image, also guides the two branches to learn each other's features to make up for the missing information of themselves, so as to take into account the global features and detail features required for classification and positioning respectively, thereby meeting the accuracy requirements for the classification and positioning of the installed frame bracket.
[0041] Refer to Figure 1 , the method specifically includes: obtaining the detection results for positioning and classifying the frame bracket in the image through the frame bracket detection model; judging the installation condition of the frame bracket according to the detection results, and displaying the detection results and the judgment results in the AR device; wherein, refer to Figure 2 , the construction method of the frame bracket detection model (i.e., the target detection module in the intelligent edge controller in Figure 1 ) includes:
[0042] 1) Input the frame bracket image samples in the training set into the model, select an initial feature map obtained by feature extraction through the backbone network in the model, and input it into the detail branch and the semantic branch respectively;
[0043] 2) The feature maps in the detail branch and the feature maps in the semantic branch are input into the bilateral attention guidance module (i.e., the BAGM module in Figure 2 ), and the outputs to the detail branch and the semantic branch are obtained respectively, so as to generate new feature maps in the detail branch and the semantic branch respectively; repeat 2) until the number of times of inputting the corresponding feature maps into the bilateral attention guidance module reaches the set number of times; it should be noted that the downsampling rate of the feature maps in the detail branch is less than that of the feature maps in the semantic branch;
[0044] 3) Fuse the finally obtained feature maps in the detail branch and the semantic branch, and the fusion result passes through the detection head (i.e., the Seg Head in Figure 2 ) to obtain the detection results; update the parameters of the model through the corresponding loss function, and repeat 1)-3) until stopped to obtain the constructed model;
[0045] Refer to Figure 3, the processing method of the bilateral attention guidance module for the input feature map (i.e., the feature map input to the bilateral attention guidance module) includes: after processing the feature map in the detail branch of the input bilateral attention guidance module (referred to as the feature map in the input detail branch) to the same scale as the feature map in the semantic branch of the input bilateral attention guidance module (referred to as the feature map in the input semantic branch), through spatial attention processing, and then fusing with the feature map in the input semantic branch to obtain the output of the bilateral attention guidance module to the semantic branch; passing the feature map in the input semantic branch through channel attention processing, and then fusing with the feature map in the input detail branch to obtain the output of the bilateral attention guidance module to the detail branch.
[0046] It can be seen that this method displays the detection results of positioning and classifying the frame bracket through the AR device and the judged installation situation of the frame bracket, effectively improving the visualization effect of guiding the installation result of the frame bracket. At the same time, on the basis of the detail branch with a smaller downsampling rate to fully capture and preserve detail features and the semantic branch with a larger downsampling rate to fully capture the overall context information in the image, a bilateral attention guidance module is also adopted to achieve the effect of mutual guidance between the features in the detail branch and the features in the semantic branch; the bilateral attention guidance module uses the spatial attention generated by the detail branch (i.e., the feature map in the detail branch after enhancing the detail features through spatial attention), and fuses with the feature map in the input semantic branch as the output to the semantic branch, realizing the effect of the detail branch guiding the semantic branch to learn the detail features lacking in the semantic branch; the bilateral attention guidance module also uses the channel attention generated by the semantic branch (i.e., the feature map in the semantic branch after enhancing the global features through channel attention), and fuses with the feature map in the input detail branch as the output to the detail branch, realizing the effect of the semantic branch guiding the detail branch to learn the global features lacking in the detail branch. After the processing of the bilateral attention guidance module, both the detail branch and the semantic branch can effectively utilize the complementary information of each other, making the detail information and global information in the feature map used to obtain the detection result more abundant, so as to take into account the global information and detail information required for classification and positioning respectively, and meet the accuracy requirements for the classification and positioning of the installed frame bracket.
[0047] In this embodiment, as Figure 2 shown, a total of 2 bilateral attention guidance modules are adopted (i.e., the set number in "the number of times of inputting the corresponding feature map into the bilateral attention guidance module reaches the set number" is 2). Therefore, in addition to the feature map f obtained by inputting the initial feature map into the detail branch 4 in the detail branch, there are also 2 sets of new feature maps, namely f 5 , f 6; Similarly, in addition to the feature map g obtained by inputting the initial feature map into the semantic branch, the feature maps in the semantic branch also include two new groups of feature maps, namely g 4 and g 5 ; Then f 6 and g 4 are jointly input into the first bilateral attention guidance module for mutual guidance. According to the outputs of the first bilateral attention guidance module to the detail branch and the semantic branch respectively, new feature maps f 4 in the detail branch and new feature maps g 5 in the semantic branch are generated; After that, f 5 and g 5 are jointly input into the second bilateral attention guidance module for mutual guidance. According to the outputs of the second bilateral attention guidance module to the detail branch and the semantic branch respectively, new feature maps f 5 in the detail branch and new feature maps g 6 in the semantic branch are generated; In other embodiments, only one bilateral attention guidance module may be used to obtain only one group of new feature maps; or N (N is greater than or equal to 3) bilateral attention guidance modules may be used to obtain N groups of new feature maps; It should be noted that in this embodiment, in the process of generating the new feature map g 6 in the semantic branch according to the output of the first bilateral attention guidance module to the semantic branch, at least one operation of increasing the downsampling rate (such as one downsampling operation) is required to convert the feature map with a 16-fold downsampling rate of the output of the first bilateral attention guidance module to the semantic branch into the new feature map g 5 with a 32-fold downsampling rate; Similarly, the feature map g 5 with a 16-fold downsampling rate obtained by inputting the initial feature map into the semantic branch is obtained by inputting the initial feature map with an 8-fold downsampling rate obtained by feature extraction through the backbone network in the model into the semantic branch and performing at least one operation of increasing the downsampling rate. In other embodiments, the initial feature map selected through feature extraction by the backbone network in the model is not necessarily 8-fold downsampling rate, and the downsampling rates of f 4 , f 4 , f 5 , f 6 and g 4 , g 5 , g 6 can also be selected according to actual needs, as long as it is ensured that the downsampling rate of the feature maps in the detail branch is less than that of the feature maps in the semantic branch.
[0048] Specifically, referring to Figure 3 , the above method of obtaining the output of the bilateral attention guidance module to the semantic branch by spatial attention processing and then fusing it with the feature maps in the input semantic branch includes:
[0049] The channels of the processed feature map are divided into set groups (the number of the set groups in this embodiment is 16, and the number of groups can be set according to actual needs in other embodiments), and each group generates a spatial attention map through spatial attention; the channels of the feature map in the input semantic branch are also divided into set groups, and each group is fused with the spatial attention map generated by the grouping of the channels of the corresponding processed feature map (the fusion method in this embodiment is element-by-element multiplication, and other fusion methods can also be used in other embodiments) to obtain the output of the bilateral attention guidance module to the semantic branch.
[0050] Specifically, in this embodiment, given a feature map in the detail branch of the input bilateral attention guidance module Where H and W represent the height and width of the image respectively; in the bilateral attention guidance module, two convolutions are first used (3×3 convolutions are used in this embodiment, and the specific convolution scale does not affect the overall logical architecture, ( Figure 3 The feature map in the detail branch of the bilateral attention guidance module is downsampled to the feature map g in the semantic branch of the bilateral attention guidance module. i At the same scale, when i=4, the downsampling step size of the second convolution is n=1; when i=5, n=2, Figure 3 Then, the processed feature map (i.e., the feature map g in the semantic branch of the input) is downsampled to i The channels of the feature map in the detail branch of the input of the same scale are evenly divided into 16 groups, and each group generates a spatial attention map; the feature map g in the semantic branch of the input bilateral attention guidance module i The channels of are also divided into 16 groups, and the feature maps after grouping are fused with each group of the feature maps in the semantic branch of the bilateral attention guidance module after grouping, so as to guide the semantic branch to learn local detail information; for the convenience of expression, the feature maps in the semantic branch of the bilateral attention guidance module can be referred to as "feature maps in the input semantic branch", and the feature maps in the detail branch of the bilateral attention guidance module can be referred to as "feature maps in the input detail branch". The overall process can be expressed as:
[0051]
[0052] Among them, g" i+1 represents the output of the bilateral attention guidance module to the semantic branch, which is used to generate a new feature map g in the semantic branch. i+1 ; and denote the spatial attention function, convolution downsampling function, and element-wise multiplication, respectively. Denote the feature map in the detail branch of the input downsampled to the same scale as the feature map g in the semantic branch of the input, i.e.: i Same as the scale of the input in the detail branch of the input, i.e.:
[0053]
[0054] Where, C2d 3×3 Denotes a 3×3 two-dimensional convolution. Then group f i ' to get Where j represents the j-th group (1 ≤ j ≤ 16). The detail branch generates spatial attention through grouping to adjust the importance of sub-features in the semantic branch. The semantic features of different groups can learn and suppress noise in a targeted manner and autonomously enhance the spatial detail feature representation. Then use Add each group of features element-wise and obtain the spatial attention through the Sigmoid function. The calculation process of each group can be expressed as:
[0055]
[0056] Where, g i,j Denotes the j-th group in g i , Sum(f i ' ,j , C) denotes the element-wise addition of f i,j in the channel dimension, σ denotes the Sigmoid function, and ⊙ denotes element-wise multiplication. The in the formula of the above calculation process of each group actually belongs to the existing process of forming spatial attention, guiding the semantic branch to focus on the key details in the image while extracting context information, and the element-wise multiplication belongs to the existing feature fusion method, so it will not be elaborated here.
[0057] For the feature map of the semantic branch input to the bilateral attention guidance module First, use a 1×1 convolution ( Figure 3 denoted as Conv1×1 in Figure 3 ) and global average pooling (
[0058]
[0059] f" i+1 Denotes the output of the bilateral attention guidance module to the detail branch, used to generate the new feature map f i+1 ; the in the above formula denotes gi The process of channel attention processing belongs to the prior art, and element-wise multiplication belongs to the existing feature fusion method, which will not be elaborated here; among them, and represent the cross-channel interaction function and the global average pooling function respectively. Let represent the pooled features, that is:
[0060]
[0061] where Res, GAP, and C2d 1×1 represent feature reshaping, global average pooling, and 1×1 two-dimensional convolution respectively. Then use to perform cross-channel interaction to form channel attention, which is specifically expressed as:
[0062]
[0063] where C1d 1×9 represents 1×9 one-dimensional convolution. This process uses the semantic branch to generate channel attention to guide the detailed information to better learn the context information.
[0064] In addition, in this embodiment, the method of fusing the feature maps in the finally obtained detailed branch and the semantic branch refers to Figure 4 , including:
[0065] Select an initial feature map obtained by feature extraction through the backbone network in the model or a feature map in the detailed branch, and segment it through a boundary segmentation head (i.e., Figure 2 Edge Head in) to obtain a rough boundary map; fuse the rough boundary map with the feature map in the finally obtained detailed branch to obtain a boundary feature map; process the boundary feature map through the self-attention mechanism, which is equivalent to completing Figure 2 a series of processes of the EAM module (also called the boundary awareness module) in to obtain the output result. The boundary feature map is then fused with the result obtained by processing the feature map in the finally obtained semantic branch, or the boundary feature map is fused with the feature map in the finally obtained semantic branch to obtain the fusion result.
[0066] Specifically, referring to Figure 4 , in this embodiment, select the feature map f 4 in the detailed branch, and segment it through the boundary segmentation head to obtain a rough boundary map Then fuse the boundary map with the feature map in the finally obtained detailed branch to enhance the boundary information and obtain a boundary feature map After that, the boundary features and detail features are fully fused through the self-attention mechanism; the overall process of the EAM module can be expressed as:
[0067]
[0068] Among them and represent the function corresponding to the self-attention mechanism (for fully fusing boundary features and detail features) and the boundary feature enhancement function based on feature fusion, both of which adopt the self-attention mechanism and feature fusion method in the prior art. First, use to obtain the boundary feature map e', and its specific process can be expressed as:
[0069]
[0070] Subsequently, f 6 and e' pass through a 1×1 convolution ( Figure 4 denoted as Conv1×1 in to obtain Finally, the self-attention mechanism is used to calculate the similarity between the boundary features and the detail features, and the detail features closely related to the boundary information are given higher weights, so as to improve the sensitivity and response ability of the model to the boundary features. The specific process of the function
[0071]
[0072] Among them represents matrix multiplication (element-wise multiplication), represents element-wise addition; is the feature map after fusing the detail information and the boundary, that is, the feature map obtained by fusing the feature maps in the finally obtained detail branch and the semantic branch; the boundary perception module can make the detail branch more focused on the boundary features of the image, thereby improving the segmentation performance of weak boundaries.
[0073] Moreover, in this embodiment, the way to process the feature map in the finally obtained semantic branch includes: after fusing the feature map in the finally obtained semantic branch with the feature map in the finally obtained detail branch, and then processing it through the channel attention and the self-attention mechanism in turn.
[0074] On this basis, in this embodiment, after processing the boundary feature map through the self-attention mechanism, it is fused with the result 6 obtained by processing the feature map g in the finally obtained semantic branch. In this case, the way to process the feature map in the finally obtained semantic branch is through Figure 2The SAM module (also known as the semantic awareness module) in Figure 5 includes: After fusing the feature maps in the finally obtained semantic branch with the feature maps in the finally obtained detail branch, the result is processed successively through the channel attention and self-attention mechanisms. In other embodiments, after processing the boundary feature maps through the self-attention mechanism, they can also be directly fused with the feature map g in the finally obtained semantic branch 6 to obtain a fusion result, that is, without processing the feature map g in the finally obtained semantic branch 6 .
[0075] Specifically, the SAM module, that is, the semantic awareness module, can adaptively learn the key information between cross-scale feature maps through the global channel attention mechanism and the self-attention mechanism, thereby reducing the semantic differences existing between different branches. The structure of the SAM module refers to Figure 5 , and the whole process can be expressed as:
[0076]
[0077] Among them, and respectively represent the feature fusion function based on the self-attention mechanism and the semantic feature extraction function based on the channel attention mechanism. can be expressed as:
[0078]
[0079] Among them, L represents the fully connected operation. Subsequently, a 1×1 convolution ( Figure 5 represented as Conv1×1 in and are respectively mapped to and Finally, the self-attention mechanism is adopted to calculate the similarity between the detail features and the semantic features, and higher weights are assigned to the associated features in the two branches. The specific process of the function can be expressed as:
[0080]
[0081] Among them, is the feature after fusing semantic information and detail information, that is, the result obtained by processing the feature map g in the finally obtained semantic branch 6 . In fact, both the channel attention mechanism and the self-attention mechanism adopted in the above SAM module belong to the prior art and will not be elaborated here.
[0082] Moreover, in this embodiment, the obtaining method of the feature map in the finally obtained semantic branch includes:
[0083] The new feature map in the semantic branch generated from the output of the bilateral attention guidance module to the semantic branch for the last time in 2) is processed by the MPPM module (also known as the multi-scale pyramid pooling module) to obtain the feature map in the finally obtained semantic branch; in this embodiment, the MPPM module processes the new feature map in the semantic branch generated from the output of the bilateral attention guidance module to the semantic branch for the last time in 2) (taking g in this embodiment 6 ) The processing method is referred to Figure 6 , specifically: it is processed by each average pooling layer and the global average pooling layer with different pooling kernel scales respectively, and then the processing results of the global average pooling layer and each average pooling layer are fused with this feature map to obtain the feature map in the finally obtained semantic branch. In the MPPM module, pooling kernels of different scales can output features with different receptive fields, improving the richness of features.
[0084] In this embodiment, the number of average pooling layers is not less than 4. Then, in order to avoid the damage caused by the fusion between features with too large differences, the method of fusing the processing results of the global average pooling layer and each average pooling layer with this feature map includes:
[0085] The processing result of the average pooling layer with the smallest pooling kernel scale is fused with the processing results of the two average pooling layers with the second smallest and the third smallest pooling kernel scales respectively to obtain the second fusion feature and the third fusion feature; the result obtained by fusing the third fusion feature with the processing results of other average pooling layers with larger pooling kernel scales and the global average pooling layer respectively is then fused with the processing result of the average pooling layer with the smallest pooling kernel scale, the second fusion feature and the third fusion feature together, and the result of this joint fusion is fused with this feature map. That is, features that are adjacent are preferably selected for fusion to avoid the damage caused by the fusion between features with too large differences.
[0086] In this embodiment, there are 4 average pooling layers with different pooling kernel scales in the MPPM module and 1 global average pooling layer; referring to Figure 6 , specifically, for the feature map input to the MPPM module The output of each average pooling layer and the output of the global average pooling layer at each scale (i = {1, 2, 3, 4, 5}) can be expressed as:
[0087]
[0088] Among them, DW 3×3 is a 3×3 depthwise separable convolution, Figure 6is denoted as DWConv 3×3 (depthwise separable convolution, like the element-wise addition operation, is part of feature fusion, used to improve the fusion effect, more efficient than ordinary convolution, and it is also possible to use ordinary convolution to form the fusion module), Up represents the upsampling operation, GAP is global pooling, and P k,s denotes that the pooling kernel size is k ∈ {3, 7, 11}( Figure 6 is denoted as Kernel = 3, Kernel = 7, Kernel = 11 in Figure 6 and the stride is s ∈ {2, 3, 5}( is denoted as Stride = 2, Stride = 3, Stride = 5 in Figure 6 which is the average pooling operation, that is, the separate processing of each average pooling layer with different pooling kernel scales. Output features of different scales are obtained
[0089] Based on the above construction method of the frame bracket detection model, the method for judging the installation situation of the frame bracket according to the detection result (that is Figure 1 the function of the template matching module in
[0090] includes: comparing the detection result of localizing the frame bracket image with a preset installation position template to determine whether the installation position of the frame bracket is correct;
[0091] comparing the detection result of classifying the frame bracket image with a preset installation category template to determine whether the installation category of the frame bracket is correct.
[0092] Specifically, comparing the detection result of localizing the frame bracket image with a preset installation position template to determine whether the installation position of the frame bracket is correct is achieved by comparing the IoU (Intersection over Union) of the detection bounding box and the template bounding box. Both the detection bounding box and the template bounding box are represented in the form of a rectangle, and the format is:
[0093] B = (X min , Y min , X max , Y max )
[0094] where, (X min , Y min(X max , Y max ) is the coordinate of the upper left corner of the rectangular frame, and (X
[0095] First, perform confidence screening on the detection results of localizing the frame bracket image, and eliminate the detection results with a confidence level lower than T conf . The screening formula is:
[0096] D = (d ∈ Results | C(d) > T conf )
[0097] where D represents the screening result of performing confidence screening on the detection results of localizing the frame bracket image; C(d) represents the confidence level of the detection result, and T conf is the set confidence threshold.
[0098] Judge whether the position is correct by calculating the IoU between the detection bounding box and the template bounding box:
[0099]
[0100] where B is the detection bounding box, and B T is the template bounding box.
[0101] The judgment result of whether the installation position of the frame bracket is correct is specifically: if IoU > T iou (T iou is the predefined IoU threshold), it is judged that the installation position of the frame bracket is correct, otherwise it is judged that the installation position of the frame bracket is abnormal.
[0102] Similarly, comparing the detection results of classifying the frame bracket image with the pre-set installation category template to judge whether the installation category of the frame bracket is correct specifically includes comparing the category T detected of the detection result with the template category T template to judge whether they are consistent:
[0103]
[0104] In the formula, Matchclass represents the judgment result of whether the installation category of the frame bracket is correct, True represents that the installation category of the frame bracket is judged to be correct, and False represents that the installation category of the frame bracket is judged to be incorrect; T detected = T template means that the category T detected of the detection result is compared with the template category T template and they are consistent.
[0105] In addition, in this embodiment, the way of displaying the detection results and judgment results in the AR device (that is Figure 1The functions of the data transmission module and augmented reality technology include:
[0106] Map the detection box corresponding to the detection result of locating the frame bracket in the image to the field of view of the AR device, and superimpose and display the class name corresponding to the detection result of classifying the frame bracket in the image on the above detection box in the AR device; use the color corresponding to the judgment result as the color of the above detection box displayed in the field of view of the AR device. In this embodiment, the AR device used is an AR glasses. In other embodiments, other AR devices can also be used.
[0107] Specifically, in this embodiment, the data of the detection results of locating and classifying the frame bracket image are received through subscribing to a specified topic via the MQTT protocol; the received data mainly includes the camera number, the bracket category T detected (the data of the detection result of classifying the frame bracket image), the bracket position B detected (the data of the detection result of locating the frame bracket image), and the bracket status S (including the judgment result of whether the installation position and installation category of the frame bracket are correct), and parse and extract the key fields from the received JSON data.
[0108] According to the bracket status value S, map the bracket status to different color tags for display in the AR glasses. For example:
[0109]
[0110] Map the bracket status to different colors of the detection box according to the Label respectively, green for correct, red for wrong, and blue for missing.
[0111] Moreover, in this embodiment, the method of mapping the detection box to the field of view of the AR device includes:
[0112] Collect the three-dimensional point cloud data of the current scene; use the SLAM algorithm to estimate the pose matrix of the AR device in the world coordinate system of the current scene in real time through the continuously collected camera image data of the current scene and the three-dimensional point cloud data of the current scene; for example, first collect and pre-store the three-dimensional point cloud data of the current scene, then use the SLAM algorithm to collect the image data of the current scene through the camera of the AR device to identify the environmental features, and match the collected image data of the current scene with the pre-stored three-dimensional point cloud data of the current scene, calculate the current pose of the AR device relative to the pre-stored point cloud data, and then the pose matrix of the AR device in the world coordinate system of the current scene can be determined;
[0113] Afterwards, the coordinates of the detection box are transformed into the world coordinate system of the current scene through the pose matrix; then, through the transformation matrix from the rendering coordinate system of the AR device to the world coordinate system, the coordinates of the detection box in the world coordinate system of the current scene are transformed into the rendering coordinate system of the AR device, so as to map the detection box into the field of view of the AR device. Thus, it is possible to superimpose and display the detection box and the bracket name of the bracket in the field of view of the AR device, and the status information of the bracket is represented by frames of different colors. The color and label name can change dynamically according to the bracket status S, and as the data is received and updated, the AR device automatically refreshes the view.
[0114] Embodiment of computer system
[0115] This embodiment provides a technical solution for a computer system. The computer system includes a processor, and executable program instructions are stored in the processor. The executable program instructions are used to be executed to implement the method for detecting the installation condition of the vehicle frame bracket based on AR assistance in the above-mentioned embodiment of the method for detecting the installation condition of the vehicle frame bracket based on AR assistance.
[0116] Since the specific working principle and effect of the computer system in this embodiment have been described in detail in the embodiment of the method for detecting the installation condition of the vehicle frame bracket based on AR assistance, they will not be elaborated here.
[0117] Embodiment of computer-readable storage medium
[0118] This embodiment provides a technical solution for a computer-readable storage medium. Computer program instructions are stored in the storage medium. The computer program instructions are used to be executed to implement the method for detecting the installation condition of the vehicle frame bracket based on AR assistance in the above-mentioned embodiment of the method for detecting the installation condition of the vehicle frame bracket based on AR assistance.
[0119] Since the specific working principle and effect of the computer-readable storage medium in this embodiment have been described in detail in the above-mentioned embodiment of the method for detecting the installation condition of the vehicle frame bracket based on AR assistance, they will not be elaborated here.
[0120] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention.
Claims
1. A method for detecting the installation status of a frame bracket based on AR, characterized in that: The vehicle frame bracket detection model is used to obtain detection results for positioning and classifying the vehicle frame bracket in the image; the installation status of the vehicle frame bracket is judged according to the detection results, and the detection results and judgment results are displayed in the AR device; The frame bracket detection model is constructed in the following ways: 1) Inputting the frame bracket image samples in the training set into the model, selecting an initial feature map obtained by feature extraction of the backbone network in the model, and inputting them into the detail branch and the semantic branch respectively; 2) The feature map in the detail branch and the feature map in the semantic branch are input into the bilateral attention guidance module to obtain outputs to the detail branch and the semantic branch respectively, so as to generate new feature maps in the detail branch and the semantic branch respectively; 2) is repeated until the number of times the corresponding feature map is input into the bilateral attention guidance module reaches the set number; 3) Fusing the feature maps in the detail branch and the semantic branch finally obtained, and obtaining the detection result through the detection head; updating the parameters of the model through the corresponding loss function, and repeating 1)-3) until stopping to obtain the constructed model; The bilateral attention guidance module processes the input feature map in the following way: processing the feature map in the input detail branch to the same scale as the feature map in the input semantic branch, processing it through spatial attention, and then fusing it with the feature map in the input semantic branch to obtain the output of the bilateral attention guidance module to the semantic branch; processing the feature map in the input semantic branch through channel attention, and then fusing it with the feature map in the input detail branch to obtain the output of the bilateral attention guidance module to the detail branch.
2. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 1, characterized in that: The method of performing spatial attention processing and fusing the feature map in the input semantic branch to obtain the output of the bilateral attention guidance module to the semantic branch includes: The channels of the processed feature map are divided into set groups, and each group generates a spatial attention map through spatial attention; the channels of the feature map in the input semantic branch are also divided into set groups, and each group is fused with the spatial attention map generated by the grouping of the channels of the corresponding processed feature map to obtain the output of the bilateral attention guidance module to the semantic branch.
3. The AR-assisted frame bracket installation condition detection method according to claim 1 or 2, characterized in that: The methods for fusing the feature maps in the final detail branch and the semantic branch include: An initial feature map obtained by feature extraction of the backbone network in the model or a feature map in the detail branch is selected, and a rough boundary map is obtained by segmentation through a boundary segmentation head; the rough boundary map and the feature map in the detail branch obtained finally are fused to obtain a boundary feature map; after processing the boundary feature map through a self-attention mechanism, it is fused with the result obtained by processing the feature map in the semantic branch obtained finally, or it is fused with the feature map in the semantic branch obtained finally to obtain a fusion result.
4. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 3 is characterized in that: The method of processing the feature map in the final semantic branch includes: fusing the feature map in the final semantic branch with the feature map in the final detail branch, and then processing them through channel attention and self-attention mechanisms in sequence.
5. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 1 or 2, characterized in that: The method of obtaining the feature map in the final semantic branch includes: The new feature map in the semantic branch generated by the output of the bilateral attention guidance module to the semantic branch for the last time in 2) is processed by each average pooling layer and the global average pooling layer with different pooling kernel scales, and then the processing results of the global average pooling layer and each average pooling layer are fused with the feature map to obtain the final feature map in the semantic branch.
6. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 5, characterized in that: The number of the average pooling layers is not less than 4, and the method of fusing the processing results of the global average pooling layer and each average pooling layer with the feature map includes: The processing results of the average pooling layer with the smallest pooling kernel scale are fused with the processing results of the two average pooling layers with the second and third smallest pooling kernel scales to obtain the second fused feature and the third fused feature; the third fused feature is fused with the processing results of other average pooling layers with larger pooling kernel scales and the global average pooling layer, and then the results are fused together with the processing result of the average pooling layer with the smallest pooling kernel scale, the second fused feature and the third fused feature, and the fused result is fused with the feature map.
7. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 1 or 2, characterized in that: Methods for judging the installation condition of the frame bracket based on the test results include: Comparing the detection result of positioning the frame bracket image with the pre-set installation position template to determine whether the frame bracket installation position is correct; The detection result of classifying the frame bracket image is compared with a pre-set installation category template to determine whether the frame bracket installation category is correct.
8. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 7, characterized in that: The method of displaying the detection result and the judgment result in the AR device includes: The detection frame corresponding to the detection result of locating the frame bracket in the image is mapped to the field of view of the AR device, and the category name corresponding to the detection result of classifying the frame bracket in the image is superimposed and displayed with the detection frame in the AR device; the color corresponding to the judgment result is used as the color of the detection frame displayed in the field of view of the AR device.
9. The method for detecting the installation condition of a frame bracket based on AR assistance according to claim 8, characterized in that: Ways to map the detection box to the field of view of the AR device include: Collect three-dimensional point cloud data of the current scene; use the SLAM algorithm to estimate the pose matrix of the AR device in the world coordinate system of the current scene in real time through the camera image data and the three-dimensional point cloud data of the current scene continuously collected; convert the coordinates of the detection frame into the world coordinate system of the current scene through the pose matrix; and then convert the coordinates of the detection frame in the world coordinate system of the current scene into the rendering coordinate system of the AR device through the conversion matrix from the rendering coordinate system of the AR device to the world coordinate system, so as to map the detection frame to the field of view of the AR device.
10. A computer system comprising a processor, wherein the processor is configured to execute executable program instructions, wherein: The executable program instructions are used to be executed to implement the AR-assisted frame bracket installation condition detection method according to any one of claims 1-9.
11. A computer-readable storage medium, wherein computer program instructions are stored in the storage medium, characterized in that: The computer program instructions are used to implement the AR-assisted frame bracket installation condition detection method as described in any one of claims 1 to 9 when executed.
Citation Information
Patent Citations
Assembly state monitoring method and system based on machine vision technology
CN114782778A