A method for object positioning and attitude recognition based on 3D vision

Through the object positioning and attitude recognition method based on 3D vision, 3D point cloud data and image data are used, combined with the conditional diffusion model guided by hierarchical feature and differential transformer network, the accuracy of object positioning and attitude recognition in complex environments is solved, and the recognition effect is achieved with high accuracy and robustness.

CN119625068BActive Publication Date: 2025-06-13SHENZJEM SOFTWELL TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510147458.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

In the case of complex lighting and occlusion, the accuracy of object positioning and posture recognition is difficult to ensure, and the geometric features of 3D data are not fully utilized.

Method used

3D vision-based object positioning and attitude recognition methods are adopted, and 3D point cloud data and image data are collected, image preprocessing and feature extraction are performed, and object positioning and attitude recognition are used to guide hierarchical feature and differential transformer network.

Benefits of technology

It improves the accuracy of object positioning and posture recognition, enhances the robustness and performance of the system, and can effectively process 3D point cloud data in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625068B_ABST
    Figure CN119625068B_ABST
Patent Text Reader

Abstract

The present invention provides a method for object positioning and attitude recognition based on 3D vision. The method includes: collecting 3D point cloud data and image data corresponding to the 3D point cloud data; preprocessing the image data to obtain a quality assessment result, and based on the quality assessment result, performing image restoration processing to generate optimized image data; inputting the optimized image data into a pre-constructed conditional diffusion model guided by hierarchical features to generate a feature representation of the optimized image data; and inputting the feature representation of the optimized image data into a differential transformer network to obtain a position and attitude prediction result of the object. Through the innovative design of image quality assessment driven by dual-representation interaction, a conditional diffusion model guided by hierarchical features, and a differential transformer network, the present invention improves the accuracy and robustness of object positioning and attitude recognition in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and computer vision, and particularly relates to a method for object positioning and pose recognition based on 3D vision technology. Background Art

[0002] With the rapid development of industrial automation and intelligent manufacturing, the precise positioning and pose recognition of objects play an increasingly important role in industrial production. At present, the object positioning and pose recognition in industrial environments mainly adopt traditional computer vision methods based on 2D images and point cloud processing methods based on a single 3D sensor.

[0003] Traditional 2D image vision methods achieve object recognition and positioning through image feature matching, but their performance drops significantly under complex lighting and occlusion conditions. Although the point cloud processing method of a single 3D sensor can obtain the depth information of an object, it is difficult to ensure the positioning accuracy due to data noise and occlusion problems.

[0004] In recent years, although some researchers have tried to introduce deep learning technology to solve these problems, the existing simple deep learning models have not fully utilized the geometric features of 3D data, resulting in unstable recognition results. In addition, factors such as complex lighting and occlusion in industrial environments seriously affect the image quality, and there is also a lack of effective processing and feature extraction methods for 3D point cloud data. These problems all affect the accuracy of positioning and recognition. Summary of the Invention

[0005] In view of this, this application provides a method for object positioning and pose recognition based on 3D vision, which improves the accuracy of object positioning and pose recognition.

[0006] An embodiment of this application provides a method for object positioning and pose recognition based on 3D vision, including:

[0007] Collect 3D point cloud data and image data corresponding to the 3D point cloud data;

[0008] Preprocess the image data to obtain a quality assessment result, and based on the quality assessment result, perform image restoration processing to generate optimized image data;

[0009] Input the optimized image data into a pre-constructed conditional diffusion model guided by hierarchical features to generate a feature representation of the optimized image data;

[0010] Input the feature representation of the optimized image data into a differential transformer network to obtain the position and pose prediction results of the object;

[0011] Output the precise spatial position and pose information of the object according to the position and pose prediction results of the object.

[0012] The embodiment of the present application further provides an object positioning and pose recognition device based on 3D vision, and the device includes:

[0013] A processor; and,

[0014] A memory communicatively connected to the processor; wherein,

[0015] The memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is enabled to execute the above-mentioned object positioning and pose recognition method based on 3D vision.

[0016] The embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the above-mentioned object positioning and pose recognition method based on 3D vision.

[0017] The embodiment of the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the above-mentioned object positioning and pose recognition method based on 3D vision are implemented.

[0018] The present application has the following technical effects:

[0019] Through the innovative dual-representation interaction-driven image quality assessment method, the image processing ability in complex environments is improved, providing high-quality input data for subsequent feature extraction.

[0020] Based on the design of a hierarchical feature-guided conditional diffusion model, the generation and optimization of high-quality features are realized, effectively integrating the geometric features of 3D point cloud data and the visual features of image data.

[0021] The differential transformer network is used for object positioning and pose recognition, and through multi-scale feature fusion and an adaptive weight mechanism, the overall performance and robustness of the system are significantly improved. Description of the Drawings

[0022] In order to more clearly illustrate the disclosed embodiments in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0023] Figure 1 It is a schematic flowchart of the object positioning and pose recognition method based on 3D vision provided by the embodiment of the present application;

[0024] Figure 2It is a schematic flowchart of the image data preprocessing method provided by an embodiment of the present application;

[0025] Figure 3 It is a schematic structural diagram of the two-stream neural network provided by an embodiment of the present application;

[0026] Figure 4 It is a schematic structural diagram of the conditional diffusion model provided by an embodiment of the present application;

[0027] Figure 5 It is a schematic structural diagram of the differential transformer network provided by an embodiment of the present application;

[0028] Figure 6 It is a schematic structural diagram of the 3D vision-based object positioning and pose recognition device provided by an embodiment of the present application. Detailed implementation manners

[0029] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] It should be clear that the following uses specific specific examples to illustrate the implementation manners of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0031] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0032] It should also be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The figures only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout pattern may also be more complex.

[0033] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0034] The object positioning and pose recognition method based on 3D vision provided by the embodiments of the present application has an overall process as Figure 1 shown, including the following steps:

[0035] Obviously, the above embodiments are only examples given for clear illustration, rather than limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

[0036] S1: Collect 3D point cloud data and the image data corresponding to the 3D point cloud data.

[0037] In this solution, a laser scanner (such as LiDAR) can be used to efficiently obtain high-precision 3D point cloud data. These devices calculate the distance of an object by emitting laser beams and measuring the return time, thereby generating a 3D point cloud.

[0038] S2: Preprocess the image data to obtain a quality assessment result, and based on the quality assessment result, perform image restoration processing to generate optimized image data.

[0039] Among them, the image preprocessing process in step S2 is as Figure 2 shown, specifically including two stages: quality assessment and image restoration.

[0040] In the quality assessment stage, first construct a two-stream neural network structure, as Figure 3 shown. This network extracts low-level features and high-level features of the image data respectively. For the low-level features, a multi-scale convolutional layer is used to extract basic features including edges and textures, where the multi-scale convolutional layer includes convolutional kernels of 3×3, 5×5, and 7×7. For example, for a vehicle body image, the low-level features include basic features such as the texture and edge contour of the vehicle body surface. For the high-level features, a deep convolutional network with an attention mechanism is used to extract semantic-level features, such as the semantic information of components such as doors and windows.

[0041] To improve the effect of feature extraction, this application adopts a contrastive learning strategy to optimize the feature extraction parameters of the two-stream neural network. Specifically, images from different perspectives of the same region are constructed as positive sample pairs, and images from different regions are used as negative sample pairs. The network parameters are optimized by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. For example, for a certain local area of the vehicle body, images collected from different angles should have similar feature representations.

[0042] Then, based on a pre-designed interaction module, feature fusion is performed on the dual-representation features. The cross-attention mechanism is used to calculate the correlation matrix between the low-level feature map and the high-level feature map, and dynamically adjust the weights of different features. For example, when a scratch is detected on the vehicle body surface, the low-level features can capture the texture features of the scratch, while the high-level features can identify that this is a defect in the door area. The interaction between the two types of features can more accurately evaluate the image quality.

[0043] In the image restoration stage, first, adaptive denoising is performed on image regions with a quality score lower than a preset threshold (such as 0.5). The algorithm dynamically adjusts the denoising parameters according to the type and degree of the noise. For example, non-local means denoising is used for noise caused by uneven illumination, and wavelet threshold denoising is used for sensor noise.

[0044] Next, local adaptive histogram equalization technology is used to improve the image details. The specific method is to divide the image into multiple small blocks (such as 16×16 pixels), perform histogram equalization on each block, and use bilinear interpolation to ensure smooth transitions between blocks. For example, for a local area of the vehicle body with insufficient light, the contrast of this area is enhanced to make the details clearer.

[0045] Finally, the overall image quality is optimized through a multi-scale fusion strategy. The image is decomposed into a pyramid structure of different scales, enhanced processing is performed separately at each scale, and then the final optimized image data is obtained through weighted fusion. Among them, the weight coefficients are adaptively determined by the quality score, and higher enhancement weights are assigned to regions with lower quality scores.

[0046] In S2, the image data is preprocessed to obtain a quality assessment result, including:

[0047] S2.1: Based on the constructed two-stream neural network, the low-level features and high-level features of the image data are respectively extracted to obtain dual-representation features including the low-level features and the high-level features, where: the low-level features include the edge and texture information of the image data, and the high-level features include the semantic information of the image data.

[0048] More specifically, S2.1.1: For the low-level features, a multi-scale convolutional layer is used to extract basic features including edges and textures, where the multi-scale convolutional layer includes at least one of convolutional kernels of 3×3, 5×5, and 7×7.

[0049] For the extraction of low-level features in the two-stream neural network, a multi-scale convolutional layer is first used for processing. Specifically, a multi-scale convolutional module containing convolutional kernels of three sizes, 3×3, 5×5, and 7×7, is designed. Among them, the 3×3 convolutional kernel is mainly used to capture fine texture information in the image, such as fine scratches on the vehicle body surface and subtle changes in the paint surface; the 5×5 convolutional kernel focuses on extracting medium-scale edge features, such as the contour lines of vehicle body parts and window frames; the 7×7 convolutional kernel pays attention to larger-scale regional features, such as the boundaries of large components like doors and hoods. Through this design of multi-scale convolution, basic features of different scales in the image can be comprehensively captured. In addition, a batch normalization layer and a ReLU activation function are introduced after each convolutional layer to enhance the non-linear expression ability of feature extraction.

[0050] Therefore, using at least one of the convolutional kernels of 3×3, 5×5, and 7×7 can be applied to a variety of different scenarios and improve the prediction accuracy for different regional features.

[0051] S2.1.2: For the high-level features, a deep convolutional network with an attention mechanism is used to extract the semantic-level features;

[0052] For the extraction of high-level features, a deep convolutional network with an attention mechanism is used. Among them, the backbone of the deep convolutional network adopts the ResNet structure, and the gradient vanishing problem of the deep network is effectively alleviated through residual connections.

[0053] In addition, a self-attention module is added between convolutional layers to enable the network to adaptively focus on key regions in the image.

[0054] For example, when processing a vehicle body image, the attention mechanism will give priority to focusing on regions with significant semantic information such as door handles, vehicle emblems, and characteristic components.

[0055] In other words, such high-level features not only contain the semantic information of the object but also retain the spatial relationship between components, providing an important basis for subsequent localization and pose recognition.

[0056] S2.1.3: Construct images of the same region from different perspectives as positive sample pairs and images of different regions as negative sample pairs, and optimize the feature extraction parameters of the two-stream neural network by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs.

[0057] To further improve the effect of feature extraction, this application innovatively introduces a contrastive learning strategy to optimize the feature extraction parameters of the two-stream neural network. First, images are collected from different perspectives of the same vehicle body area to construct positive sample pairs.

[0058] For example, two images taken of the front door area at 45 degrees and 60 degrees respectively form a pair of positive samples.

[0059] Meanwhile, images of different areas are selected as negative samples, such as images of the front door and the trunk.

[0060] During the training process, a contrastive loss function is used to guide the network to learn feature representations. By maximizing the feature similarity between positive sample pairs and minimizing the feature similarity between negative sample pairs, the network can learn more discriminative feature representations. In addition, to enhance the robustness of the model, factors such as different lighting conditions and different shooting distances are also considered when constructing sample pairs. Through this strategy, the network can extract more robust and distinguishable features, effectively improving the accuracy of subsequent processing.

[0061] It should be noted that to meet the requirements of different scenarios, this application provides a flexible adjustment space for the specific parameter settings of the multi-scale convolutional layer. For example, when processing high-resolution images, larger-sized convolutional kernels (such as 9×9) can be added; in scenarios with limited computing resources, only 3×3 and 5×5 convolutional kernels can be used. Similarly, different schemes can be selected for the specific implementation of the attention mechanism according to actual needs, such as channel attention, spatial attention, or a combination of both. The sample construction strategy in contrastive learning can also be adjusted according to specific application scenarios, such as increasing or decreasing the range of perspective differences, adjusting the ratio of positive and negative samples, etc. This flexible design makes this solution have good adaptability and scalability.

[0062] S2.2: Based on a pre-designed interaction module, perform feature fusion on the dual-representation features, and dynamically adjust the relative weights of the low-level features and the high-level features through an attention mechanism to obtain fused features.

[0063] After obtaining the dual-representation features, it is necessary to effectively integrate these features through a feature fusion module.

[0064] First, design an interaction module to achieve the fusion of dual representations, and use a cross-attention mechanism to calculate the correlation matrix between the low-level feature map and the high-level feature map. For example, when a scratch is detected on the vehicle body surface, the low-level features can capture the texture features of the scratch, while the high-level features can identify that this is a defect in the door area. Through the dynamic weight adjustment of the attention mechanism, the system can adaptively balance the importance of low-level and high-level features according to the characteristics of different regions, so as to obtain a more comprehensive and accurate feature representation.

[0065] S2.3: Based on the fused features, determine the image quality score as the quality assessment result.

[0066] Specifically, the implementation of the interaction module uses a multi-layer perceptron structure.

[0067] First, map the low-level and high-level features to the same feature space, and then calculate the attention weights between them. This process can be expressed as: first calculate the query matrix Q, the key matrix K, and the value matrix V, then obtain the attention weights through the softmax function, and finally perform weighted fusion to obtain the fused features. This mechanism ensures that the advantages of different-level features can be fully utilized during the feature fusion process, improving the quality assessment result of the feature representation.

[0068] And based on the quality assessment result, perform image restoration processing to generate optimized image data, including:

[0069] S2.4: For image regions with a quality score lower than the preset threshold, adaptively adjust the denoising parameters according to the noise type and degree to obtain a denoised image.

[0070] Based on the fused features, use a fully connected layer network to calculate the image quality score. This fully connected layer network maps the fused features to a normalized quality score between 0 and 1, where 1 represents the highest quality and 0 represents the lowest quality. For example, a quality score of 0.9 indicates that the image quality is good, and 0.3 indicates that the image quality is poor and needs to be restored. This scoring process not only considers the basic quality attributes of the image but also includes the quality requirements related to specific tasks, providing an important basis for subsequent image restoration processing.

[0071] For image regions with a quality score lower than the preset threshold (such as 0.5), the system starts an adaptive denoising processing flow. The denoising algorithm will dynamically adjust the parameters according to the type and degree of the noise. For example, for noise caused by uneven illumination, the non-local means denoising algorithm is used; for sensor noise, the wavelet threshold denoising method is used. The algorithm first estimates the statistical characteristics of the noise and then adaptively sets the denoising parameters according to these characteristics. This adaptive mechanism ensures that the denoising process can effectively remove noise interference while retaining useful information.

[0072] S2.5: Divide the denoised image into blocks, and use local adaptive histogram equalization technology to process the images of each block to obtain an enhanced image.

[0073] To further improve the image quality, the system uses local adaptive histogram equalization technology to enhance the image. The specific approach is to divide the image into multiple small blocks, and a typical block size is 16×16 pixels. Histogram equalization is performed on each block separately, and bilinear interpolation is used to ensure smooth transitions between blocks and avoid blocky artifacts. This local processing method can better adapt to the lighting conditions in different regions of the image and improve the visibility of details. For example, for a locally dark area of the vehicle body, the details can be made more clearly visible by enhancing the contrast in that area.

[0074] S2.6: Through a multi-scale fusion strategy, decompose the enhanced image into a pyramid structure of different scales, perform enhancement processing on the images at each scale respectively, and fuse the images at each scale by weighting to obtain the optimized image data, where the weight values of the images at each scale are determined based on quality scores.

[0075] Finally, the image is globally optimized through a multi-scale fusion strategy. This process first decomposes the image into a pyramid structure of different scales, and enhancement processing is performed on each scale separately. During the processing at different scales, the system will select appropriate enhancement parameters according to the characteristics of that scale. For example, at a higher scale, the focus is mainly on adjusting the overall brightness and contrast, while at a lower scale, more attention is paid to enhancing details. After the processing is completed, the results at each scale are synthesized into the final optimized image through weighted fusion. The weight coefficients are adaptively determined based on the previous quality scores, and regions with lower quality scores will be assigned higher enhancement weights, which ensures that processing resources can be more invested in areas that need improvement.

[0076] It should be noted that during the entire image preprocessing process, the system continuously monitors the processing effect. After each processing step is completed, a quality assessment is performed on the intermediate result. If it is found that the processing effect of a certain step is not ideal, the system will automatically adjust the parameters for optimization. In addition, considering the different requirements of different application scenarios, parameter adjustment interfaces are reserved for each processing module in this solution, and operators can flexibly adjust processing parameters according to specific needs, such as denoising intensity, contrast enhancement degree, etc. This design greatly improves the practicality and adaptability of the solution.

[0077] Step S3: Input the optimized image data into a pre-constructed conditional diffusion model based on hierarchical feature guidance to generate a feature representation of the optimized image data.

[0078] In step S3, as Figure 4 shown, first, a feature extraction network based on the PointNet++ architecture is used to extract the geometric features of the 3D point cloud data.

[0079] For the body point cloud, the network captures local and global structure information through hierarchical sampling and grouping operations. For example, the first layer extracts local surface features (such as flatness, curvature), the second layer extracts component-level features (such as the geometry of the door, hood), and the third layer extracts the overall vehicle contour features.

[0080] Meanwhile, ResNet50 is used as the backbone network to extract the visual features of the image data, and the transfer of detailed information is maintained through residual connections. During the conditional diffusion process, it includes two stages: forward diffusion and reverse denoising.

[0081] In the forward diffusion stage, Gaussian noise is gradually injected into the features to generate a sequence of noise features at T time steps. In the reverse denoising stage, a denoising network with a U-Net structure is designed to gradually restore the contaminated features.

[0082] To ensure that the generated features are consistent with the physical properties of the input data, this application designs a conditional embedding network to encode the physical constraints of the object (such as body size, symmetry) as conditional vectors. During the diffusion process, the conditional vectors are injected into the denoising network at each time step to guide the feature generation process to meet the physical constraints.

[0083] In addition, this application also constructs a feature pyramid structure to organize and guide the feature generation process. As Figure 4 shown, the pyramid contains three levels: the bottom level maintains the original resolution for preserving the basic features; the middle level uses 2x downsampling to preserve the component-level features; the top level uses 4x downsampling to preserve the overall-level features. At different stages of the diffusion process, the feature guidance information of the corresponding levels is introduced respectively: the top-level features are introduced in the early stage of diffusion to provide global guidance, the middle-level and top-level features are introduced simultaneously in the middle stage of diffusion to provide detail constraints, and the bottom-level features are mainly introduced in the late stage of diffusion to restore the fine structure.

[0084] S3 specifically includes:

[0085] S3.1: Use a feature extraction network based on the PointNet++ architecture to extract the geometric features of the 3D point cloud data.

[0086] When processing the 3D point cloud data, this application uses a feature extraction network based on the PointNet++ architecture. This feature extraction network captures local and global structure information through hierarchical sampling and grouping operations.

[0087] First, extract local surface features in the first layer, including geometric attributes such as flatness and curvature of the surface. For example, for vehicle body point cloud data, features such as surface flatness and bending degree can be identified.

[0088] Then, extract component-level features in the second layer, such as the geometric shapes of components like doors and hoods. Finally, extract overall contour features in the third layer to obtain the overall shape information of the vehicle body. Through this hierarchical feature extraction method, the system can comprehensively obtain the geometric feature information of the object.

[0089] To improve the effect of feature extraction, the network adopts a hierarchical sampling and grouping strategy. In the sampling stage, the farthest point sampling algorithm is used to select key points to ensure that the sampling points are evenly distributed on the object surface. In the grouping stage, the ball query method is used to group the neighborhoods of each key point to construct local feature regions. By setting different ball radii, local features of different scales can be obtained. In addition, the network also introduces a dense connection mechanism to fuse features at different levels and enhance the feature expression ability.

[0090] S3.2: Use a backbone network based on ResNet50 to extract the visual features of the image data, and obtain multi-level visual features through residual connections.

[0091] For the processing of image data, the system uses ResNet50 as the backbone network for visual feature extraction.

[0092] The deep structure based on ResNet50 can effectively learn the hierarchical representation of images, and its unique residual connection mechanism can effectively alleviate the gradient vanishing problem in the training of deep networks.

[0093] Specifically, the deep structure network based on ResNet50 contains multiple residual blocks. Each residual block consists of multiple convolutional layers and realizes residual learning through shortcut connections. This design not only improves the training efficiency of the network but also maintains the effective transmission of detailed information. For example, when processing vehicle body images, low-level features retain detailed texture information, and high-level features contain semantic-level understanding.

[0094] S3.3: Based on a conditional diffusion process containing T time steps, gradually remove the noise components in the features through a denoising network with a U-Net structure at each time step to obtain a preliminary feature representation;

[0095] In the conditional diffusion process, the system designs a processing flow including two stages: forward diffusion and reverse denoising. In the forward diffusion stage, a sequence of noise features at T time steps is generated by gradually injecting Gaussian noise.

[0096] For example, from t = 1 to t = T, the geometric and visual information in the features is gradually contaminated by noise. The noise intensity at each time step is predefined and usually follows a pattern of increasing from small to large. This progressive noise injection process helps the model learn the multi-scale representation of features.

[0097] In the reverse denoising stage, the system uses a denoising network with a U-Net structure to gradually recover the contaminated features.

[0098] The U-Net network consists of an encoder and a decoder, and realizes the fusion of features at different levels through skip connections. At each time step, the denoising network receives the current noisy features and time encoding as inputs, predicts and removes the noise components. This step-by-step denoising method can better preserve the structural information of the features. To improve the denoising effect, the network also introduces a self-attention mechanism, enabling it to focus on the key regions in the features.

[0099] S3.4: Encode the physical constraint conditions of the object into a conditional vector, inject the conditional vector into the denoising network at each time step, ensure that the generated features are consistent with the physical properties of the input data, and obtain the feature representation of the optimized image data.

[0100] To ensure that the generated features are consistent with the physical properties of the input data, this application innovatively introduces a conditional control mechanism.

[0101] First, encode the physical constraints of the object into a conditional vector. These constraints include physical properties such as vehicle body size and symmetry. The conditional vector is generated using an encoder-decoder structure, which compresses the high-dimensional physical constraint information into a low-dimensional conditional space. Then, during the denoising process at each time step, inject the conditional vector into the denoising network to guide the feature generation process to meet the physical constraints. For example, when generating vehicle body features, the conditional vector will ensure that the generated result maintains left-right symmetry and the relative position relationship between components meets the actual assembly requirements.

[0102] It should be noted that to adapt to different application scenarios, each network module in this solution provides flexible configuration options.

[0103] For example, parameters such as the number of sampling points and the ball query radius of the PointNet++ network can be adjusted according to the point cloud density; the number of layers of ResNet50 can be trimmed according to the computing resources; the number of time steps T in the conditional diffusion process can be set according to the accuracy requirements. This flexible design enables the system to achieve a good balance between accuracy and efficiency. At the same time, considering various situations that may be encountered in practical applications, the system also designs an exception handling mechanism, which can effectively handle abnormal situations such as noisy point clouds and blurred images.

[0104] And optionally, S3.5: Before feature extraction, construct a feature pyramid structure including a bottom layer, a middle layer, and a top layer, where:

[0105] Before feature guidance, the present application first constructs a feature pyramid structure including a bottom layer, a middle layer, and a top layer. This pyramid structure uses a step-by-step downsampling method to capture multi-scale information of the image at different resolutions. This hierarchical design enables the system to simultaneously focus on local details and global structures, thus providing a more comprehensive feature representation. It should be noted that at each layer of the pyramid, channel alignment is performed through adaptive pooling and 1×1 convolution to ensure the consistency of feature dimensions.

[0106] S3.5.1: The bottom layer maintains the original resolution and is used to preserve the basic features of the image data.

[0107] At the bottom layer of the pyramid, the features maintain the original resolution and are used to preserve the basic features of the image data.

[0108] This layer mainly focuses on the detailed information of the object surface, such as local features like the texture and edges of the vehicle body surface. For example, for a vehicle body image, the bottom layer features can accurately capture fine features such as scratches, bumps, and seams on the surface. To ensure the integrity of the features, a dense feature map is used at the bottom layer, and complete feature information is retained at each pixel position. In addition, an attention mechanism is introduced at this layer to enhance the feature extraction ability for key local regions.

[0109] S3.5.2: The middle layer uses a 2-fold downsampling of the original resolution and is used to preserve the component-level features of the image data.

[0110] The middle layer features use a 2-fold downsampling of the original resolution and are mainly used to preserve the component-level features.

[0111] The features at this level can capture the medium-scale structure of the object, such as the shape and layout information of vehicle body components like doors, hoods, and headlights. The downsampling process uses a max-pooling operation with a stride of 2, which not only reduces the spatial dimension of the feature map but also retains important structural information. To enhance the expressive ability of the features, a multi-scale convolution module is introduced at this layer to capture the multi-scale features of the components through receptive fields of different sizes. At the same time, residual connections are used to ensure the effective transmission of information and avoid feature degradation.

[0112] S3.5.3: The top layer uses a 4-fold downsampling of the original resolution and is used to preserve the overall-level features of the image data.

[0113] At the top layer of the pyramid, a 4-fold downsampling of the original resolution is used to preserve the overall-level features.

[0114] The features of this layer mainly focus on the overall contour of the object and the large-scale structural relationships, such as the overall shape of the vehicle body, the relative positional relationships between various components, etc. Through further downsampling, the system can obtain a larger receptive field, thereby better understanding the global structure of the object. In this layer, in addition to the conventional convolution operations, a global context module is added to enhance the global representation ability of the features by modeling long-range dependencies.

[0115] To make the feature pyramid structure more efficient and accurate, the present application designs a feature interaction mechanism between layers. First, a bidirectional connection is established between adjacent layers, allowing the up and down flow of feature information. For example, the detailed features of the bottom layer can be passed to the middle layer through the upsampling path to help better locate the component boundaries; at the same time, the semantic information of the middle layer can also guide the extraction of the bottom layer features through the downsampling path. In addition, during the feature extraction process of each layer, an adaptive feature recalibration module is added to dynamically adjust the weights of the features according to their importance.

[0116] In practical applications, the specific parameters of the pyramid structure can be flexibly adjusted according to requirements. For example, for high-resolution images, the number of pyramid layers can be increased; for scenarios with limited computing resources, the number of feature channels can be appropriately reduced. At the same time, the downsampling ratio between different layers can also be adjusted according to the specific task to achieve the best balance between accuracy and efficiency. In addition, considering the real-time requirements, the present solution also optimizes the feature extraction process, such as using separable convolutions to reduce the computational amount and adopting a feature cropping strategy to remove redundant information.

[0117] S3.6: At different stages of the diffusion process, introduce the corresponding hierarchical feature guidance information respectively, specifically including:

[0118] S3.6.1: Introduce top-level features in the early stage of diffusion.

[0119] In different stages of the conditional diffusion process, the present application designs a hierarchical feature guidance mechanism to guide feature generation by dynamically introducing the corresponding hierarchical feature information at different time steps.

[0120] First, in the early stage of diffusion, that is, when t is close to T, mainly introduce the top-level features of the pyramid to provide global guidance. At this stage, the system focuses on the overall structure and large-scale features of the object, such as the basic contour of the vehicle body, the rough layout of the main components, etc. Since there is more noise in the features at this time, using high-level semantic features can provide more reliable structural constraints to avoid the generation process deviating from the correct direction.

[0121] S3.6.2: Introduce middle-level and top-level features simultaneously in the middle stage of diffusion.

[0122] In the middle stage of diffusion, the system simultaneously introduces middle-level and top-level features for guidance.

[0123] The features at this stage already have a certain structure, but still need to be refined. The middle-level features mainly focus on component-level details, such as the shape of the car door, the position of the headlights, etc., while the top-level features continue to provide global constraints. The combined use of the two-level features can gradually restore the component-level detailed features while maintaining the accuracy of the overall structure. To achieve effective feature fusion, the system uses an attention mechanism to calculate the correlation between the current noisy features and the guiding features of each layer, and controls the intensity of feature injection through a gating unit.

[0124] S3.6.3: In the late stage of diffusion, focus on introducing low-level features.

[0125] In the late stage of diffusion, that is, when t is close to 0, the system focuses on introducing low-level features to restore the fine structure. At this time, the noise in the features is already small, which is more suitable for restoring local detailed information. The low-level features contain rich local texture and edge information, which are crucial for restoring the subtle features of the object surface. For example, in the vehicle body positioning scenario, this stage can accurately restore the texture details, edge contours and other features of the vehicle body surface. To ensure the accuracy of detail restoration, the system uses a smaller denoising step size at this stage and increases the injection weight of the low-level features.

[0126] S3.7: Maintain the detail information of the low-level features through residual connections, and output a feature representation that fuses multi-level information, specifically including:

[0127] S3.7.1: Add lateral connections to each layer of the feature pyramid.

[0128] Throughout the process of feature guidance, this application maintains the detail information of the low-level features through a residual connection mechanism.

[0129] Specifically, lateral connections are added to each layer of the feature pyramid to fuse the original features of the corresponding resolution with the currently generated features. These lateral connections form a fast channel for information, enabling the detailed features at the bottom layer to be directly transmitted to the generation process, avoiding information loss during multi-layer transmission. For example, when processing vehicle body images, the clarity of subtle features such as vehicle body surface texture and welding lines can be maintained through this connection.

[0130] S3.7.2: Based on the fusion method of adaptive weighted summation, fuse the original features of the corresponding resolution with the currently generated features;

[0131] Feature fusion adopts the method of adaptive weighted summation, and the weight coefficients are dynamically adjusted through learnable parameters. The system designs a feature correlation evaluation module to determine the fusion weights according to the content similarity between the original features and the generated features. When the original features contain important structural information, the system will automatically increase their weights; when the quality of the original features is poor, more reliance is placed on the generated features. This adaptive fusion mechanism ensures that the final features can take into account both the fidelity of the original information and the quality of the generated features.

[0132] To further improve the quality of feature generation, this application also introduces a feature enhancement module in the residual connection. This module contains multiple residual units, and each unit is equipped with a channel attention and a spatial attention mechanism. Channel attention is used to enhance important feature channels, and spatial attention is used to highlight key spatial positions. Through this fine-grained feature enhancement, the system can better maintain and restore the feature information of key regions. For example, in the vehicle body pose recognition task, the system will automatically enhance the feature expression of the regions around key feature points (such as vehicle corner points, feature lines, etc.).

[0133] It should be noted that the entire feature guidance and fusion process is end-to-end trainable. In the training stage, the system simultaneously optimizes the parameters of the feature extraction, feature guidance, and feature fusion modules through backpropagation. To improve the training efficiency, a progressive training strategy is adopted: first, the model is pre-trained with a smaller number of time steps, and then the number of time steps is gradually increased for fine-tuning. In addition, the system also designs a multi-scale discriminator to evaluate the quality of the generated features, and further improves the authenticity and detail performance of the features through adversarial training.

[0134] Step S4: Input the feature representation of the optimized image data into the differential transformer network to obtain the position and pose prediction results of the object.

[0135] In step S4, as Figure 5 shown, the differential transformer network adopts a multi-head self-attention structure with three attention heads, which respectively focus on the overall contour, local details, and spatial position relationships of the object. Feature transformation is performed through a two-layer fully connected feed-forward network. The first layer uses the ReLU activation function to expand the feature dimension, and the second layer maps the features back to the original dimension. At the same time, learnable position encodings are attached to each feature vector to inject spatial position information into the feature representation.

[0136] In the feature processing process, based on the constructed feature pyramid structure, multi-scale processing is performed on the output features of the differential transformer. Multiple transformer modules are run in parallel at different scales. The larger-scale transformer focuses on the global structure to determine the approximate position of the entire vehicle, and the smaller-scale transformer focuses on local details for precise key point localization. Then, an adaptive weight fusion mechanism is adopted to dynamically adjust the fusion weights according to the confidence levels predicted at each scale.

[0137] During the training phase, a multi-task loss function based on position localization and pose recognition is constructed. Smooth L1 loss is used for position localization to reduce the impact of outliers, and a weighted combination of cross-entropy loss and regression loss is adopted for pose recognition. A phased training strategy is adopted. In the first phase, a relatively large learning rate (such as 0.001) is used to train only the position localization branch for quick convergence. In the second phase, a smaller learning rate (such as 0.0001) is used to train both the position localization and pose recognition branches for fine-tuning. During the training process, gradient clipping is used to prevent gradient explosion.

[0138] S4 specifically includes:

[0139] S4.1: Based on the constructed feature pyramid structure, multi-scale processing is performed on the output features of the differential transformers in the differential transformer network. Multiple transformer modules are run in parallel at different scales to capture spatial information of the same object at different granularities, where: large-scale features are used to determine the position of the object; small-scale features are used to locate the key points of the object.

[0140] The differential transformer network adopts a multi-scale feature processing architecture. Based on the constructed feature pyramid structure, hierarchical processing is performed on the output features. Multiple transformer modules are run in parallel at different scales to capture spatial information of the same object at different granularities. Larger-scale transformer modules mainly focus on the overall structure and global features of the object, while smaller-scale transformer modules focus on the precise characterization of local details. For example, in the vehicle body pose recognition task, large-scale features are used to determine the approximate position and main pose angles of the whole vehicle, while small-scale features are used to accurately locate key feature points on the vehicle body, such as wheel centers, door handles, etc.

[0141] To achieve efficient multi-scale feature processing, the system designs a hierarchical transformer structure. At the largest scale, downsampling with a stride of 4 is used to obtain a low-resolution feature map, and the transformer module mainly processes the overall contour information of the vehicle body at this scale. At the medium scale, downsampling with a stride of 2 is adopted, and the corresponding transformer module is responsible for processing the spatial relationship of the main components of the vehicle body. At the smallest scale, the original resolution is maintained, and the transformer module focuses on processing local fine features. This multi-scale parallel processing method not only improves the computational efficiency but also makes full use of complementary information at different scales.

[0142] S4.2: Based on the adaptive weight fusion mechanism, the fusion weights are dynamically adjusted according to the confidence levels predicted at each scale, and based on the spatial attention module, feature representations of important regions are generated, and the original feature information is retained through residual connections.

[0143] In each scale of the transformer module, a specific attention mechanism is designed to enhance feature representation. For large-scale features, a global self-attention mechanism is adopted, enabling each position to obtain global context information. For medium-scale features, local window attention is used to establish feature correlations within a fixed-size window. For small-scale features, a deformable attention mechanism is employed, which can adaptively adjust the attention range according to the feature content. This hierarchical attention design ensures efficient capture of corresponding feature information at each scale.

[0144] In the feature fusion stage, this application designs an adaptive weight fusion mechanism. This mechanism dynamically adjusts the fusion weights of features at different scales based on the confidence levels of the prediction results at each scale.

[0145] Specifically, the system first calculates the confidence scores for the prediction results at each scale, which reflect the reliability of the predictions at that scale. Then, these confidence scores are converted into fusion weights through the softmax function, ensuring that the sum of the weights is 1. The larger the weight value, the greater the contribution of the prediction result at that scale to the final fusion. This adaptive weighting method can effectively balance the importance of prediction results at different scales.

[0146] To further improve the effect of feature fusion, the system introduces a spatial attention module. This module can adaptively identify and highlight the feature representations of important regions. Specifically, first, the spatial attention map of the feature map is calculated to represent the importance of each spatial position. Then, the attention map is multiplied by the original features with weights to enhance the feature intensity of the important regions. This mechanism ensures that the system can pay more attention to regions crucial for localization and pose recognition, such as the feature edges of the vehicle body and key connection points.

[0147] In addition, this application retains the original feature information through residual connections. On the feature processing path at each scale, skip connections are added to directly transfer the input features to the output end. This design not only helps with the backpropagation of gradients but also maintains the integrity of the features. At the same time, to handle the size differences between features at different scales, the system adopts learnable upsampling and downsampling modules to ensure size alignment during feature fusion. Through experimental verification, this multi-scale feature extraction and fusion scheme can significantly improve the accuracy of position localization and pose recognition.

[0148] It should be noted that in practical applications, the system will dynamically adjust the parameters of each module according to the requirements of specific tasks. For example, in scenarios with high precision requirements, the number of small-scale transformers can be increased; in scenarios with high real-time requirements, the number of processing scales can be appropriately reduced. In addition, the system also provides a feature clipping function, which can selectively retain the most important feature channels according to the limitations of computing resources, improving the processing efficiency while ensuring performance.

[0149] S4.3: Use the position prediction regression branch with a hybrid structure based on the classification branch and the regression branch to predict the coordinate values of the object in three-dimensional space. Among them, the classification branch is used to predict the range of the object's pose, and the regression branch is used to fine-tune the pose angle of the object.

[0150] In the position and pose prediction stage, the present application designs a prediction network with a hybrid structure based on the classification branch and the regression branch. This network adopts an architecture of a shared encoder and a dual-path decoder. Among them, the classification branch is responsible for predicting the rough range of the object's pose, and the regression branch is responsible for accurately predicting the three-dimensional space coordinate values and the pose angle. The design of this hybrid structure makes full use of the complementarity of the classification and regression tasks, and can provide accurate positioning results while ensuring the prediction robustness. For example, in the vehicle body pose recognition task, the classification branch first divides the pose into several intervals (such as one interval every 30 degrees) to predict the general orientation of the vehicle body, and then the regression branch fine-tunes on this basis to obtain the accurate pose angle.

[0151] For position prediction, the regression branch adopts a multi-layer perceptron structure and directly outputs the coordinate values (x, y, z) of the object in three-dimensional space. To improve the prediction accuracy, the system introduces an adaptive weight mechanism in the regression branch to dynamically adjust the loss weights according to the uncertainty of different coordinate components. For example, when the prediction uncertainty in the depth direction (z-axis) is relatively large, the system will automatically reduce the loss weight in this direction to avoid the negative impact of unreliable depth information on the overall prediction. In addition, the regression branch also outputs the predicted uncertainty estimate, which is of great significance for subsequent filtering and tracking.

[0152] In terms of pose prediction, the classification branch first divides the Euler angle space into discrete intervals. The value range of each direction (pitch angle, roll angle, yaw angle) is evenly divided into multiple categories, and the system predicts the probability that the object's pose falls into each interval through a softmax classifier. This discretization process can effectively handle the periodicity and ambiguity problems in pose representation. On this basis, the regression branch further outputs the angle offset relative to the center of the predicted interval to obtain the accurate pose angle. This strategy of combining rough classification and fine regression significantly improves the accuracy and stability of pose prediction.

[0153] S4.4: The steps to generate the differential transformer network include:

[0154] Construct a multi-head self-attention module with three attention heads, where: the first attention head focuses on the overall contour of the object, the second attention head focuses on the local details of the object, and the third attention head focuses on the spatial position relationship of the object.

[0155] To improve the robustness of the prediction, this application constructs a multi-head self-attention module with three attention heads in the differential transformer network. The first attention head mainly focuses on the overall contour of the object and captures the main shape features of the object through the global receptive field. The second attention head focuses on the local details of the object, such as key regions like feature points and edges. The third attention head is responsible for modeling the spatial position relationship between different parts of the object. The outputs of these three attention heads are fused through weighted summation to obtain the final feature representation, where the fusion weights are adaptively determined by learnable parameters.

[0156] In the specific implementation, each attention head adopts the scaled dot-product attention mechanism. First, query vectors (Q), key vectors (K), and value vectors (V) are generated through three groups of linear projections. Then, the dot product between Q and K is calculated to obtain the attention scores, and the normalized attention weights are obtained through a scaling factor and the softmax operation. Finally, the attention weights are multiplied by the value vectors to obtain the weighted features. To enhance the expressive power of the attention mechanism, the system also introduces relative position encoding so that the attention calculation can consider the spatial relationship between features.

[0157] Construct a feed-forward network with two fully connected layers. The first layer uses the ReLU activation function to expand the feature dimension, and the second layer maps the features back to the original dimension and ensures the stability of the feature distribution through layer normalization.

[0158] In terms of feature transformation, this application designs a feed-forward network with two fully connected layers. The first layer uses the ReLU activation function to map the features to a higher-dimensional space, enhancing the non-linear expressive power of the network. The second layer maps the features back to the original dimension and ensures the stability of the feature distribution through layer normalization. This design not only helps with the non-linear transformation of features but also effectively prevents the problems of gradient vanishing and gradient explosion. In addition, the system adds a dropout layer between the two fully connected layers to enhance the generalization ability of the model by randomly deactivating some neurons.

[0159] Append learnable position encoding to each feature vector to inject spatial position information into the feature representation and obtain the feature representation with position information.

[0160] To make full use of spatial location information, the system attaches learnable position encodings to each feature vector. These position encodings combine absolute position encoding and relative position encoding, enabling both the expression of the absolute spatial position of features and the description of the relative position relationships between features. During the training process, the parameters of the position encodings are updated simultaneously with other parts of the network, allowing the system to adaptively learn the optimal position representation. This position-aware feature representation is crucial for accurately understanding the spatial structure and pose information of objects.

[0161] It should be noted that the specific parameters of the above-mentioned various modules can be flexibly adjusted according to application requirements. For example, the number of attention heads can be increased or decreased according to the task complexity, the dimension of the fully connected layer can be scaled according to computing resources, and different implementation schemes of position encoding can also be selected according to specific scenarios. At the same time, the system provides an interface for model compression, supporting the optimization of the model's computational efficiency through techniques such as knowledge distillation and channel pruning, enabling it to better adapt to different application scenarios.

[0162] Next, a multi-task loss function based on position localization and pose recognition is constructed. Among them, position localization is based on smooth L1 loss, and pose recognition adopts a weighted combination of cross-entropy loss and regression loss. The cross-entropy loss is used for classifying poses, and the regression loss is used for predicting angles.

[0163] For the two core tasks of position localization and pose recognition, this application designs a loss function based on a multi-task learning framework. This loss function comprehensively considers the accuracy of position localization and the accuracy of pose recognition, and achieves balanced optimization of the two tasks through reasonable weight allocation. Among them, smooth L1 loss is used for position localization to reduce the influence of outliers, and pose recognition combines cross-entropy loss and regression loss to achieve progressive optimization from rough classification to accurate regression. For example, in the vehicle body localization task, the system first learns the approximate pose range through cross-entropy loss, and then accurately predicts the angle value through regression loss.

[0164] The smooth L1 loss function of the position localization branch is an improvement on the standard L1 loss. When the prediction error is less than the preset threshold β (for example, β = 1.0), a quadratic function form is used to calculate the loss, which can provide a smoother gradient near the target; when the error is greater than the threshold, the traditional L1 norm is used to reduce the influence of outliers. This design not only ensures the sensitivity to small errors but also improves the robustness of the model to abnormal samples. The mathematical expression of the loss function is: when |x| < β, loss = 0.5x² / β; when |x| ≥ β, loss = |x| - 0.5β, where x represents the difference between the predicted value and the true value.

[0165] For the pose recognition task, the cross-entropy loss is mainly used to optimize the classification branch. The system first divides the pose space into multiple discrete intervals. For example, the 360-degree space is evenly divided into 12 intervals, each interval being 30 degrees. Then, the softmax function is used to calculate the probability that the predicted value falls into each interval, and the cross-entropy loss is used to measure the difference between the predicted probability and the true label. This classification-based processing method can effectively handle the periodicity problem in pose representation and avoid the training instability that may be caused by direct regression.

[0166] In the regression branch, the system uses a weighted smooth L1 loss to optimize the angle prediction. Considering that the importance of different pose parameters (pitch angle, roll angle, yaw angle) may be different, the system assigns different weights to each parameter. These weights can be adjusted according to the specific application scenario. For example, in some tasks, more attention may be paid to the accuracy of the yaw angle, and its loss weight can be appropriately increased. In addition, the system also considers the periodic characteristics of the pose and selects the minimum angle difference when calculating the loss, that is, when the angle difference exceeds 180 degrees, 360 degrees minus this difference is used.

[0167] Finally, a staged training method is adopted to train the preset learning model. In the first stage, based on the first learning rate, the position localization branch of the learning model is trained. In the second stage, based on the second learning rate, both the position localization and pose recognition branches are trained. During the training process, a gradient clipping method is introduced for the gradient, where the first learning rate is greater than the second learning rate.

[0168] In terms of the training strategy, this application adopts a staged training method. In the first stage, the position localization branch is mainly trained, and a relatively large learning rate (such as 0.001) is used to quickly optimize the position prediction ability. The focus of the training in this stage is to enable the model to accurately locate the spatial position of the target object. During the training process, the position error on the validation set is dynamically monitored. When the error decline tends to level off, the second stage of training is entered.

[0169] In the second stage, both the position localization and pose recognition branches are trained, and a relatively small learning rate (such as 0.0001) is used for fine optimization. In this stage, the system balances the training difficulty of the two tasks by adjusting the weight coefficients in the multi-task loss. Specifically, when it is found that the training loss of a certain task is significantly higher than that of other tasks, the system will automatically increase the loss weight of this task to ensure that each task can be fully optimized.

[0170] To improve the stability of training, a gradient clipping mechanism is introduced during the training process. When the calculated gradient exceeds the preset threshold (for example, setting the threshold to 10.0), the system will scale down the gradient value proportionally to avoid training instability caused by overly large gradients. In addition, a learning rate decay strategy is adopted to gradually reduce the learning rate in the later stage of training, helping the model converge to a better local minimum.

[0171] It should be noted that batch normalization technology is adopted throughout the training process to accelerate training convergence and improve the model's generalization ability. In each mini-batch, the system calculates the mean and variance of the features and performs normalization processing. At the same time, to prevent overfitting, regularization techniques such as weight decay and early stopping are introduced. The combined use of these training techniques ensures that the model can have good generalization performance while ensuring accuracy.

[0172] Step S5: According to the predicted results of the object's position and pose, output the precise spatial position and pose information of the object.

[0173] In step S5, first, the sliding window averaging method is used to perform temporal smoothing on the prediction results of consecutive multiple frames.

[0174] For example, for the prediction results of consecutive 5 frames, weighted averaging is used to obtain a more stable position estimate, where the weight of the most recent frame is the largest and decreases frame by frame. Then, through a pre-calibrated coordinate system transformation matrix, the prediction results in the camera coordinate system are transformed into the world coordinate system of the industrial robot.

[0175] Finally, the processed position and pose information are encapsulated into a standard data structure, including: the three-dimensional coordinates of the center point of the vehicle body (accurate to the millimeter level), the pose angle (accurate to 0.1 degree), the spatial coordinates of key feature points, and the confidence score of the prediction result. At the same time, a visualization result is generated for display on the human-machine interaction interface, with detection boxes, pose arrows, and key point markers superimposed on the original image, and different colors are used to represent the confidence level.

[0176] In a specific application scenario of this application, this method can be used for body positioning and pose recognition on an automotive manufacturing production line. The system collects point cloud data and image data of the vehicle body through a 3D camera. After the above processing process, the spatial position and pose information of the vehicle body can be accurately obtained, providing an accurate positioning basis for subsequent automated assembly and quality inspection. In actual tests, this method can still maintain a high positioning accuracy in a complex industrial environment, with an average position error of less than 1 mm and a pose angle error of less than 0.5 degrees.

[0177] Among them, in S5, outputting the precise spatial position and pose information of the object includes:

[0178] S5.1: Apply the sliding window averaging method to perform temporal smoothing on the prediction results of consecutive multiple frames, and obtain the smoothed prediction results through weighted averaging, where the weights decrease frame by frame according to the input time sequence of the frames.

[0179] In the final result output stage, the present application first uses the sliding window averaging method to perform temporal smoothing on the prediction results of consecutive multiple frames. This method uses an exponentially decaying weighted average strategy to assign different weights to the prediction results at different times. Specifically, for a sliding window containing N frames (e.g., N = 5), the weight of the latest frame is the largest, and the weight gradually decreases as time goes back. For example, if the prediction results of 5 consecutive frames show that the coordinates of the center point of the vehicle body are ( , , ), ( , , )...( , , ), then the final smoothed result can be expressed as: p = + +... + ,

[0180] where > >... > and Σ = 1.

[0181] This smoothing strategy can effectively suppress the noise and jitter of single-frame prediction and provide more stable output results.

[0182] To better handle mutation situations in temporal smoothing, the system also designs an anomaly detection mechanism. When the deviation of the prediction result of a certain frame from the historical prediction exceeds a preset threshold, the system will automatically adjust the weight of this frame. If large deviations occur in consecutive multiple frames, it is considered that the target has indeed undergone rapid movement. At this time, the weights of historical frames will be reduced so that the output can respond more quickly to the movement changes of the target. In addition, the system will also record the confidence score of each frame prediction and use the confidence as a weight adjustment factor during weighted averaging to increase the contribution of reliable prediction results.

[0183] S5.2: Through a pre-calibrated coordinate system transformation matrix, transform the smoothed prediction results from the camera coordinate system to the target coordinate system to obtain the position and attitude angle of the object in the target coordinate system.

[0184] After obtaining the smoothed prediction results, they need to be transformed from the camera coordinate system to the target coordinate system (such as the robot world coordinate system). This step is achieved through a pre-calibrated coordinate system transformation matrix, including the rotation matrix R and the translation vector t. Specifically, if the position of the target in the camera coordinate system is p_cam = (x, y, z) and the attitude angles are (α, β, γ), then through rigid body transformation, the position in the world coordinate system can be obtained as p_world = Rp_cam + t. At the same time, the attitude angles also need to be correspondingly transformed according to the relationship between the two coordinate systems to obtain the Euler angles (θ, , ψ).

[0185] Special attention is paid to the accuracy issue during the coordinate system transformation process. First, the system uses a high-precision camera calibration algorithm to obtain the internal and external parameters, and reduces the error by taking the average value of multiple calibrations. Second, in actual applications, online calibration is regularly performed to update the transformation matrix through a calibration object with a known position to compensate for the calibration error caused by environmental changes. In addition, the system also implements an accurate transformation based on hand-eye calibration to ensure that the transformed coordinates can meet the accuracy requirements of the industrial level.

[0186] The last step is to encapsulate the processing results into a standard data structure and generate a visualization result. The data structure contains the following key information: the three-dimensional coordinates of the target center point (accurate to the millimeter level), the attitude angles (accurate to 0.1 degrees), the spatial coordinates of the key feature points (such as the four corners of the vehicle body, the endpoints of the feature lines, etc.), and the confidence scores of each prediction result. Each data field is attached with timestamp information, which is convenient for subsequent time series analysis and trajectory tracking. At the same time, the system will mark the outliers. For example, when the confidence of a certain prediction result is lower than the threshold, a corresponding warning mark will be added to the data.

[0187] S5.3: Encapsulate the transformed position and attitude angles into a standard data structure containing three-dimensional coordinates, attitude angles, key feature point coordinates, and confidence scores, and generate a visualization result.

[0188] In terms of visual display, the system superimposes multiple layers of information on the original image. First, the detection boxes are drawn, and different colors are used to represent the confidence levels: green represents high confidence (>0.9), yellow represents medium confidence (0.7 - 0.9), and red represents low confidence (<0.7). Then, the attitude arrows are displayed, and the attitude angles of the object are represented by three orthogonal arrows respectively. The key points are marked with different-shaped markers, and the relevant feature points are connected by line segments to intuitively display the spatial structure of the object.

[0189] In addition, the system is also designed with complete data recording and analysis functions. Each positioning result will be stored in the database, including information such as the original image, intermediate results of the processing process, and final output. These data can be used for subsequent offline analysis, such as accuracy evaluation, performance optimization, etc. The system also provides a data export interface, supporting the export of results into common data formats (such as JSON, CSV, etc.) for easy data exchange and integration with other systems.

[0190] It should be noted that the design of the entire output module fully considers the actual application requirements. The system supports multiple data interface protocols (such as TCP / IP, industrial fieldbus, etc.) and can push the processing results to the downstream control system in real time. At the same time, considering the special requirements of the industrial field, the system also provides a fault diagnosis and exception handling mechanism, which can detect and report the abnormal conditions of the positioning system in a timely manner to ensure the reliable operation of the entire system.

[0191] The embodiment of the present invention also provides an object positioning and attitude recognition device based on 3D vision, as Figure 6 shown, including a processor and a memory. The memory is used to store computer programs. When the processor executes the computer programs, the object positioning and attitude recognition method as described above is implemented.

[0192] The system of the embodiment of the present disclosure can execute the method provided by the embodiment of the present disclosure, and its implementation principle is similar. The actions performed by each module in the system of each embodiment of the present disclosure correspond to the steps in the method of each embodiment of the present disclosure. For the detailed function description of each module of the system, reference can be specifically made to the description in the corresponding method shown above, and details will not be repeated here.

[0193] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the technical concept of the solution of the present disclosure, using other similar implementation means based on the technical idea of the present disclosure also belongs to the protection scope of the embodiments of the present disclosure.

Claims

1. A method for object positioning and posture recognition based on 3D vision, characterized in that: include: Collect 3D point cloud data and image data corresponding to the 3D point cloud data; Preprocessing the image data to obtain a quality assessment result, and performing image restoration processing based on the quality assessment result to generate optimized image data; Inputting the optimized image data and the 3D point cloud data into a pre-built conditional diffusion model guided by hierarchical features to generate a feature representation of the optimized image data, including: A feature extraction network based on the PointNet++ architecture is used to extract the geometric features of the 3D point cloud data; a backbone network extraction network based on ResNet50 is used to extract the visual features of the optimized image data, and multi-level visual features are obtained through residual connections; based on a conditional diffusion process including T time steps, a denoising network with a U-Net structure is used at each time step to gradually remove the noise components in the features to obtain a preliminary feature representation; wherein, in the forward diffusion process, the geometric and visual information in the features are gradually contaminated by noise; in the reverse denoising process, the contaminated features are gradually restored; the physical constraints of the object are encoded as conditional vectors, and the conditional vectors are injected into the denoising network at each time step to ensure that the generated features are consistent with the physical properties of the input data, so as to obtain the feature representation of the optimized image data; Before the feature guidance is performed, the method further includes: constructing a feature pyramid structure including a bottom layer, a middle layer and a top layer, wherein the bottom layer maintains the original resolution for storing the basic features of the optimized image data, the middle layer adopts a 2-fold downsampling of the original resolution for storing the component-level features of the optimized image data, and the top layer adopts a 4-fold downsampling of the original resolution for storing the overall-level features of the optimized image data; at different stages of the diffusion process, feature guidance information of corresponding levels is introduced respectively, wherein the top layer features are introduced in the early stage of diffusion, the middle layer and the top layer features are introduced at the same time in the middle stage of diffusion, and the bottom layer features are introduced in the late stage of diffusion; the detailed information of the bottom layer features is maintained through residual connections, and a feature representation that integrates multi-level information is output, including: adding lateral connections at each layer of the feature pyramid; integrating the original features of the corresponding resolution with the currently generated features based on an adaptive weighted summation fusion method; Inputting the feature representation of the optimized image data into a differential transformer network to obtain the position and posture prediction results of the object; According to the position and posture prediction results of the object, the precise spatial position and posture information of the object is output.

2. The object positioning and posture recognition method according to claim 1, characterized in that: The preprocessing of the image data to obtain a quality assessment result includes: Based on the constructed two-stream neural network, low-level features and high-level features of the image data are extracted respectively to obtain dual representation features including the low-level features and the high-level features, wherein the low-level features include edge and texture information of the image data, and the high-level features include semantic information of the image data; Based on a pre-designed interaction module, the dual representation features are fused, and the relative weights of the low-level features and the high-level features are dynamically adjusted through an attention mechanism to obtain fused features; Based on the fused features, an image quality score is determined as the quality assessment result.

3. The object positioning and posture recognition method according to claim 2, characterized in that: The constructed two-stream neural network extracts low-level features and high-level features of the image data respectively, including: For the low-level features, a multi-scale convolution layer is used to extract basic features including edges and textures, wherein the multi-scale convolution layer includes at least one of 3×3, 5×5 and 7×7 convolution kernels; For the high-level features, a deep convolutional network with an attention mechanism is used to extract semantic level features; Images of the same area from different perspectives are constructed as positive sample pairs, and images of different areas are constructed as negative sample pairs. The feature extraction parameters of the two-stream neural network are optimized by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs.

4. The object positioning and posture recognition method according to claim 1, characterized in that: The performing of image restoration processing based on the quality assessment result to generate optimized image data comprises: For image areas with quality scores lower than a preset threshold, the denoising parameters are adaptively adjusted according to the type and degree of noise to obtain a denoised image; Dividing the denoised image into blocks, and using a local adaptive histogram equalization technique to process the image of each block to obtain an enhanced image; Through a multi-scale fusion strategy, the enhanced image is decomposed into pyramid structures of different scales, and the images at each scale are enhanced respectively, and the images at each scale are fused in a weighted manner to obtain the optimized image data, wherein the weight value of the image at each scale is determined based on the quality score.

5. The object positioning and posture recognition method according to claim 1, characterized in that: The step of inputting the feature representation of the optimized image data into a differential transformer network to obtain the position and posture prediction results of the object includes: Based on the constructed feature pyramid structure, multi-scale processing is performed on the output features of the differential transformer in the differential transformer network, wherein multiple transformer modules are run in parallel at different scales to capture spatial information of different granularities of the same object, wherein large-scale features are used to determine the position of the object, and small-scale features are used to locate key points of the object; Based on the adaptive weight fusion mechanism, the fusion weight is dynamically adjusted according to the confidence of the prediction at each scale, and based on the spatial attention module, the feature representation of the important area is generated, and the original feature information is retained through the residual connection; A position prediction regression branch based on a hybrid structure of a classification branch and a regression branch is used to predict the coordinate value of the object in three-dimensional space, wherein the classification branch is used to predict the range of the posture of the object, and the regression branch is used to fine-tune the posture angle of the object.

6. The object positioning and posture recognition method according to claim 1 or 5, characterized in that: Generate the differential transformer network as follows: Construct a multi-head self-attention module with three attention heads, where the first attention head focuses on the overall outline of the object, the second attention head focuses on the local details of the object, and the third attention head focuses on the spatial position relationship of the object; Construct a two-layer fully connected feedforward network. The first layer uses the ReLU activation function to expand the feature dimension. The second layer maps the features back to the original dimension and ensures the stability of feature distribution through layer normalization. Attach a learnable position code to each feature vector, inject the spatial position information into the feature representation, and obtain a feature representation with position information; Construct a multi-task loss function based on position positioning and posture recognition, where position positioning is based on smoothing loss, and posture recognition adopts a weighted combination of cross entropy loss and regression loss, where cross entropy loss is used to classify posture and regression loss is used to predict angle; A phased training method is adopted to train the preset learning model, wherein the first phase trains the position positioning branch of the learning model based on a first learning rate, and the second phase trains the position positioning and posture recognition branches simultaneously based on a second learning rate, and during the training process, the gradient introduces a gradient clipping method, wherein the first learning rate is greater than the second learning rate.

7. The object positioning and posture recognition method according to claim 1, characterized in that: Outputting the accurate spatial position and posture information of the object according to the position and posture prediction results of the object includes: The sliding window averaging method is used to smooth the prediction results of multiple consecutive frames in time sequence, and the smoothed prediction results are obtained by weighted averaging, where the weights are reduced frame by frame according to the input time sequence of the frames; The smoothed prediction result is converted from the camera coordinate system to the target coordinate system through a pre-calibrated coordinate system conversion matrix to obtain the position and attitude angle of the object in the target coordinate system; The converted position and attitude angle are encapsulated into a standard data structure containing three-dimensional coordinates, attitude angles, key feature point coordinates and confidence scores to generate visualization results.

Citation Information

Patent Citations

  • Bidirectional fusion 6D object pose estimation method

    CN118799393A

  • Shadow removing method and system based on diffusion, segmentation and super-resolution model

    CN119048357A