Target visual attitude estimation method based on convolutional neural network

By introducing attention mechanism and multi-feature fusion module into the MobileNet model and establishing a virtual image synthesis sample library, the accuracy and efficiency problems of target object pose information measurement are solved, and high-precision and fast pose estimation are achieved.

CN120125657APending Publication Date: 2025-06-10CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510180816.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has challenges in the accurate annotation and measurement of the attitude information of target objects, especially in irregular motion and harsh lighting conditions, the effectiveness of visual measurement methods is significantly reduced.

Method used

A real-time pose visual estimation algorithm based on deep learning is proposed. By combining attention mechanism and multi-feature fusion module into the lightweight MobileNet model, the information extraction capability is improved and a virtual image synthesis sample library for pose changes is established.

Benefits of technology

High accuracy (average error of 0.424°) and fast measurements (20 milliseconds per image) are achieved, providing an effective solution for real-time visual measurement of target object pose.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125657A_ABST
    Figure CN120125657A_ABST
Patent Text Reader

Abstract

The invention discloses a target visual attitude estimation method based on a convolutional neural network. The method comprises the following steps: 1) establishing a data set; step 2) data preprocessing; 3) designing a neural network model; step 4) model training; 5) retraining the model; step 6) color quantification; 7) estimating the attitude; the construction of a data set is simplified, the efficiency is improved, and a high-quality data basis is provided for subsequent network training; compared with a traditional image matching algorithm, the algorithm is remarkably improved in the aspects of speed and precision, and the innovative algorithm provides a new insight and method for the research target of spatial attitude measurement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of attitude estimation, and particularly relates to a method for visual attitude estimation of an object based on a convolutional neural network. Background Art

[0002] The information generated by a target object during spatial movement includes position information and attitude information. Attitude information is crucial for obtaining insights related to the trajectory of the movement process, enabling more accurate judgment and prediction of the trajectory of the target object. This information further promotes the progress in the fields of kinematics and target dynamic tracking technology.

[0003] The attitude positioning of a target object has always been the focus of academic research and is widely applied in multiple fields, including the biomedical industry, environmental monitoring, manufacturing, and quality control processes. Early methods for measuring the spatial attitude angles of a target object relied on laser detection. A ground-based laser was used to determine the three-axis attitude of a spacecraft, and the axis angle was determined by the rotation of laser polarization. Although this method has high resolution and measurement accuracy, it requires specialized laser emission scanning equipment, resulting in high equipment requirements and complex operations. It is applicable to identifying targets with relatively minimal relative movement within medium to long distances. In 2001, Parker adopted a physical modeling method based on the laws of target motion. They determined the parameters describing the system dynamics (inertia matrix) in an inertial fixed coordinate system to determine the components of the angular momentum vector. Combining the inertial reference vector, they estimated the direction of the angular momentum vector in the fixed coordinate system. Valpiani proposed a new efficient and accurate satellite attitude estimation algorithm. Based on Bayesian nonlinear estimation of the phase space geometry of a Hamiltonian system, a new iterative filter for solving target attitude estimation was derived. However, these methods are limited to the three-axis solutions of non-rotating bodies and the single-axis solutions of simple rotating rigid bodies - targets with predictable motion trajectories in a specific environment. In contrast, target objects often exhibit irregular movements in space, making it difficult to identify their inertial parameters and challenging to conduct force analysis.

[0004] Advances in camera technology have led to the emergence and development of measurement methods using imaging devices, including monocular and binocular stereovision systems. These methods have gradually matured. Wang et al. proposed a monocular vision measurement system based on geometric features, which uses monocular vision to determine the attitude of a rotating vehicle when the geometric features of the target object are known. Martinelli employed a sensor assembly including a monocular camera, three orthogonal accelerometers, and three orthogonal gyroscopes. The information from the sensor data obtained within a short time interval can determine the speed and attitude (roll and pitch angles) of the target. Based on the vision method, template matching techniques for measuring spatial attitude have been developed. Spanish researcher de Saxe used a monocular camera, a template matching image processing algorithm, and an unscented Kalman filter to measure the hinge angle of an articulated heavy vehicle. The experimental results show that the root mean square error of joint angle measurement is 0.64° - 0.79°, and the maximum error is 1.91° - 2.76°. However, the calculation is very time-consuming, and the frame rate reaches about 1 frame per second. Song used the CAD file of the target part to create a 3D model. Offline, they generated a target model template library with multiple attitude information in different observation directions. By extending the model matching algorithm to the six-degree-of-freedom attitude detection of complex structural components, a position measurement error within 2 mm and an attitude measurement error of about 3° were achieved, and the measurement time was shortened to within 500 ms per single measurement. Although vision measurement methods provide the highest spatial resolution, they require close observation, and their effectiveness is significantly reduced under irregular target motion or poor lighting conditions.

[0005] However, there are huge challenges in accurately annotating the attitude information of the target object. Although different attitudes correspond to unique spatial information, capturing them through a camera often results in similar image appearances, especially when its shape is relatively regular and simple. This ambiguity makes manual annotation of thousands of samples not only time-consuming but also error-prone. In addition, establishing a comprehensive sample library requires extremely high image quality and accurate attitude information, making data acquisition extremely difficult. Moreover, achieving an effective and accurate measurement of the attitude of the target object remains a key challenge. Although precise methods often require a large amount of time investment, fast measurement techniques often affect accuracy. Therefore, it is crucial to develop an efficient and accurate method for measuring the attitude of the target object. Summary of the Invention

[0006] Aiming at the limitations of existing methods, the present invention proposes a real-time attitude vision estimation algorithm based on deep learning. By incorporating an attention mechanism and a multi-feature fusion module into a lightweight MobileNet model, the information extraction ability is significantly improved. A virtual image synthesis sample library for attitude changes is established.

[0007] The technical solution of the present invention is as follows:

[0008] A target visual pose estimation method based on a convolutional neural network, comprising the following steps:

[0009] Step 1) Dataset establishment: Using spherical virtual particles as the target object, generating a virtual sample library, designing textures on the surface of the virtual particles, and simulating a three-dimensional textured sphere; Capturing images of the virtual particles at different poses by keeping the focal length and sampling position fixed; And annotating the image information;

[0010] Step 2) Data preprocessing: Preprocessing the images, and using a contour detection algorithm to extract the target area, and adjusting the target area to 224×224 pixels to meet the input requirements of the neural network model;

[0011] Step 3) Neural network model design:

[0012] Select the neural network model MobileNet, remove the top network of the network model, introduce the SE module and the pyramid pooling module PPM to enhance the feature extraction ability of the network model; Use global average pooling GAP to integrate the feature vectors, and replace the Softmax layer model with the composite multi-layer perceptron MLP model to construct the MobileNetPlus network model;

[0013] Step 4) Model training:

[0014] Dividing the dataset proportionally includes: training set, test set, setting model training parameters including: epochs, batch size, initial learning rate, minimum learning rate threshold; Using a composite loss function, using the mean squared error MSE as the loss function when the difference between the predicted value and the true value is less than the preset threshold; When the difference between the predicted value and the true value is equal to or greater than the preset threshold, then use the Huber loss function, and set the weight decay coefficient, and select the mean absolute error MAE as the evaluation index for model training; Select the AdamW optimizer and the cosine annealing scheduler to prevent the model from falling into local optimum;

[0015] Step 5) Model retraining:

[0016] Using image data augmentation to retrain the neural network model for the training set images, and the image data augmentation includes grayscale conversion, edge cropping, random adjustment of saturation and brightness, image data normalization, and addition of random noise;

[0017] Step 6) Color quantization:

[0018] To reduce color deviation or noise in the image caused by lighting and shadow factors, resulting in the color distribution deviating from the training image set, color gamut transformation and K-Means clustering algorithm are used to reprocess the test image using image color quantization;

[0019] Step 7) Pose estimation:

[0020] Input the processed test set into the retrained neural network model to estimate the spatial pose of the target particle.

[0021] Furthermore, the specific process of step 1) is as follows:

[0022] 1.1) Texture design:

[0023] Based on texture features, use the VTK simulation library in Python to design a particle surface texture pattern, and the design follows the following rules:

[0024] a. Each view corresponds to a direction with a unique spatial coordinate;

[0025] b. In each view, the ratio of black and white lines is 1:0.9 - 1.1;

[0026] c. The annular ratio of the stripe spacing is greater than 5%;

[0027] d. The stripe directions in adjacent regions should be kept consistent;

[0028] 1.2) Generation and annotation of the dataset:

[0029] After the texture design is completed, use the VTK simulation library to create a dataset; the pose range of the target particle is the global spatial pose angle of 360°, and quantitative sampling is performed on each region using images with different poses, with a sampling interval of 2° - 3°, generating a series of particle images with different poses, and creating a synthetic sample library;

[0030] During the simulation image generation process, establish a file naming rule to encode the three-axis pose angle information of the image into the file name; extract the corresponding pose parameters by parsing the file name and convert them into a numerical vector as the annotation data for each image.

[0031] Furthermore, the specific process of step 2) is as follows:

[0032] 3.1) Model selection:

[0033] The MobileNet model uses the depthwise separable convolution method for parameter optimization, separating the steps of feature extraction and feature fusion:

[0034] First, each channel of the input feature map is separated, and a single convolutional kernel is used to perform a separate convolution operation on each channel. The depth of each convolutional kernel is fixed at 1, that is, each convolutional kernel only operates on a single channel of the input feature map. After performing depthwise convolution, pointwise convolution is carried out to fuse the feature map. The operation of pointwise convolution is to use a 1x1 convolutional kernel, and the size of its convolutional kernel is 1×1×M, where M is the number of channels of the output information of the previous layer. Each pointwise convolutional kernel performs a weighted combination of the feature map of the previous step along the channel direction to generate a new feature map.

[0035] 3.2) Model Optimization:

[0036] Remove some of the top network layers of the original MobileNet model, and retain the depthwise separable convolutional layer, BN layer, and ReLU layer as the feature extraction part of the improved model.

[0037] After the last depthwise separable convolutional stack structure of the MobileNet model, an SE module is embedded. The SE module consists of three steps: global compression, channel excitation, and scale transformation. First, the global compression operation compresses the input features from the H×W×C dimension into a 1×1×C channel descriptor. The channel excitation step uses two fully connected layers to generate channel attention weights. The first fully connected layer uses a scaling parameter ratio r to reduce the features, reducing the number of channels from C to C / r, thereby reducing the number of channels and computational costs. The second fully connected layer restores the channel dimension to the original input channel number C and generates normalized channel attention weights through the Sigmoid function. The scale transformation operation multiplies the weight values with the two-dimensional matrix corresponding to the channels of the original feature map to obtain the final output features after weight adjustment.

[0038] After the SE module, a pyramid pooling module PPM is introduced. The PPM includes multi-scale partitioning: the input feature map is evenly divided into sub-regions of different scales; dimensionality reduction processing: each sub-region is downsampled using a 1x1 convolution to reduce the number of channels, reducing the computational complexity and the number of parameters; upsampling: the downsampled sub-region feature maps are upsampled to the same size as the original input feature map; feature fusion: the original input feature map is concatenated with the upsampled sub-region feature maps to form the final feature representation.

[0039] The Softmax output layer is replaced by a multi-layer perceptron MLP model for pose estimation as the prediction output of the network. The MLP contains one or more hidden layers and an output layer. All layers are fully connected layers, and the fully connected layers all use GeLU as the activation function.

[0040] Furthermore, the specific process of step 6) is as follows:

[0041] 6.1) For a given data set, an initial number of clusters k and the centroids of the initial clusters are predefined; the Euclidean distance is used to measure the similarity between data objects and the cluster centers; the Euclidean distance between a data object X and a cluster centroid Ci is calculated as follows:

[0042]

[0043] where X is the data object, Ci is the centroid of the i-th cluster, and m is the dimension of the data object; Xj and Cij are the j-th attribute values of X and Ci respectively; since the similarity is inversely proportional to the distance between data objects, a higher similarity corresponds to a smaller distance; based on the distances between samples, the data set is divided into multiple clusters, and the positions of the cluster centroids are continuously updated according to the similarity between data objects and the cluster centroids. Through iteration, the sum of squared errors within the clusters SSE gradually decreases. When SSE no longer changes or the objective function converges, the clustering terminates, and the final result is obtained; the SSE of the entire data set is calculated as follows:

[0044]

[0045] where k is the number of clusters, and the smaller the SSE, the better the clustering performance;

[0046] 6.2) According to the K-means algorithm, first select the predefined k color clusters; for each pixel in the image, calculate the Euclidean distance between the pixel's color and the centers of the k clusters it belongs to, and assign the pixel to the nearest cluster; calculate the average value of the three color channel values of the pixels in each cluster as the new cluster center;

[0047] 6.3) Repeat steps 6.1)-6.2) until the cluster centers no longer change or the maximum number of iterations is reached;

[0048] 6.4) Convert the image from RGB to the Lab color space, then perform color quantization in the LAB color space, and finally convert it back to RGB.

[0049] Further, the specific steps of step 2) are as follows:

[0050] 2.1) Image preprocessing: The image is enhanced using the multi-scale retinex with color restoration (MSRCR) algorithm;

[0051] 2.2) Target area localization: The position of the particles in the image is located using the Hough circle transform algorithm;

[0052] 2.3) Profile extraction and target area extraction: Subsequently, the outermost contour of the particle is determined using the contour extraction algorithm, and then the target area is extracted; the image is extracted through fitting and masking algorithms; the ellipse is fitted to the extracted contour using the fitting ellipse function, and a mask is generated to isolate the target area;

[0053] 2.4) Zoom in and out: The extracted target area is enlarged to 224×224 pixels.

[0054] Design concept of the present invention:

[0055] Dataset construction: Using spherical particles as the target object, a virtual sample library is generated, a texture pattern is applied to the surface of the virtual particles, and a virtual three-dimensional textured sphere is simulated. By keeping the fixed focal length and sampling position, images of the virtual particles are captured at different poses. Then, the spatial pose information is marked in the image file name.

[0056] Pose estimation algorithm: A pre-trained lightweight mobile network is used as the basis of the pose estimation network. The top network of the original model is removed, and a squeeze-and-excitation block and a pyramid pooling module are introduced to enhance the feature extraction ability of the network. Finally, global average pooling (GAP) is used to integrate the feature vectors, and a composite multi-layer perceptron (MLP) model is used to replace the Softmax layer model.

[0057] Implementation: In practical applications, a monocular camera captures images of target particles at different poses. These images are processed, and the trained network predicts the three-dimensional spatial pose angles of each processed image. The output represents the estimated pose of the network for the corresponding image, achieving the desired pose measurement result.

[0058] Advantages of the present invention are as follows: The algorithm has high accuracy (average error of 0.424°) and fast measurement time (20 milliseconds per image). This innovative method provides a promising solution for real-time visual measurement of the pose of target objects, opening up new possibilities for its application in various fields. Description of the drawings

[0059] Figure 1 It is the method flow structure diagram of the present invention;

[0060] Figure 2 It is the flow chart of the dataset production of the present invention;

[0061] Figure 3 It is the texture image and simulation diagram designed by the present invention;

[0062] Figure 4 It is the schematic diagram of naming the image set of the present invention after being processed with pose angles and extraction labels;

[0063] Figure 5 Schematic diagram of the standard convolution and depthwise separable convolution processes of the present invention;

[0064] Figure 6 Schematic diagram of the standard convolution block structure and depthwise separable convolution block structure of the present invention;

[0065] Figure 7 Flowchart of the SE module of the present invention;

[0066] Figure 8 Network structure diagram of the present invention;

[0067] Figure 9 Image in the sample library of the present invention and its image after color quantization; Figure (a) is the original image, and Figure (b) shows the contrast (4, 16, 32, and 64 colors) after applying color quantization processing in the LAB and RGB color gamuts;

[0068] Figure 10 Spatial attitude angle of the present invention;

[0069] Figure 11 Spatial attitude angle prediction error and average error of different models of the present invention;

[0070] Figure 12 Physical construction diagram of the present invention;

[0071] Figure 13 Comparison result before and after image preprocessing of the present invention, where Figure (a) is the original image and Figure (b) is the image after image preprocessing;

[0072] Figure 14 Performance diagrams of Mobilenet and MobilenetPlus of the present invention;

[0073] Figure 15 Error curve distribution of the validation sets of Mobilenet and MobilenetPlus of the present invention, where the abscissa represents the number of samples and the ordinate represents the spatial attitude error;

[0074] Figure 16 Two preprocessed test images in the database of the present invention and their corresponding images; Figures (a) and (b) are the actual images to be measured, and Figures (c) and (d) are the corresponding images in the sample library;

[0075] Figure 17 Data augmentation effect diagram of the present invention; Figure (a) is the original image, Figure (b) is the grayscale image, Figure (c) is the image after cropping the main part, Figure (d) is the image after adjusting brightness and contrast, Figure (e) is the normalized image, and Figure (f) is the image after adding random noise;

[0076] Figure 18 The average error in three directions before and after data augmentation of the present invention;

[0077] Figure 19 The color quantization effect diagram of the present invention; Figure (a) is the original image, and Figures (b) to (e) are the images after color quantization of 4, 16, 32, and 64 respectively; Figures (f) and (g) are the color distributions of the original image and the processed image (16 colors) respectively;

[0078] Figure 20 The error comparison after different color quantization processes of the present invention, and the unit of each grid line is 0.05 degrees.

[0079] Figure 21 The comparison diagram of the error and time consumption of the proposed algorithm of the present invention and the traditional matching algorithm on each axis. Detailed implementation manners

[0080] The present invention will be further described below in conjunction with the accompanying drawings of the specification and embodiments.

[0081] A method for visual pose estimation of target objects based on deep learning

[0082] Step 1: Dataset establishment

[0083] As Figure 2 shown, a virtual pose dataset based on image synthesis is constructed; in order to achieve the pose estimation of spherical particles, color textures are designed on the particle surface in the simulation library, and computer-generated simulation images under different spatial poses are used; at the same time, an actual particle with a diameter of 4 cm and the same texture is designed and produced. First, a texture image with obvious feature differences and reduced complexity is designed to ensure differentiation from different angles; second, an image generation script is written; third, the image information is labeled.

[0084] 1.1 Texture design

[0085] Texture patterns are used to effectively distinguish the poses of target particles, which is crucial for understanding how the visual system perceives textures. Identify the basic perceptual dimensions of textures (also known as texture features), and 6 basic texture features include: roughness, contrast, directionality, line similarity, regularity, and coarseness; based on the theoretical framework regarding how texture distribution affects visual perception rules, use the VTK simulation library of Python to design a particle surface texture pattern, and the design follows the following rules:

[0086] a. Each view corresponds to a direction with a unique spatial coordinate;

[0087] b. In each possible view, the ratio of black and white lines is approximately equal, with a ratio of 1:0.9 - 1.1;

[0088] c. The circular ratio of the stripe spacing is greater than 5%;

[0089] d. The stripe directions in adjacent regions should be kept consistent.

[0090] After continuous optimization and comparative testing, this embodiment selects a four-color texture with black, white, green, and blue distributions, where the visualization of the texture pattern and the texture pattern of the simulated particles covered with the texture pattern is as Figure 3 shown.

[0091] 1.2 Generation and annotation of the dataset

[0092] After designing the texture, a dataset was created using the VTK simulation library; in order to accurately generate three-dimensional simulation images in different poses, uniform distribution sampling was adopted; the pose range of the target particles is the global space pose angle of 360°, and quantitative sampling of images with different poses was carried out for each region, with a minimum sampling interval of 3°.

[0093] First, fix the virtual camera, and then rotate the simulated sphere, with an angle of rotation of 3° each time; in order to ensure the independence of the three spatial angles, only one axis is rotated each time, generating a series of particle images with different poses, and creating a synthetic sample library. However, the sample library does not have inherent position and pose information; in order to simplify the annotation process, an automatic annotation method was implemented; during the generation process of the simulation images, a naming convention was established to incorporate the three-axis pose angle information of the images into the file name; then a dedicated label extraction function was developed to parse these file names, extract the corresponding pose information, and convert it into an array, effectively serving as the label for each image; after the above processing, the sample library was converted into a labeled dataset suitable for training; some of the generated sample library and the processed labeled dataset are as Figure 4 shown.

[0094] Step 2: Design of the CNN network model

[0095] 2.1 MobileNet

[0096] Select a lightweight neural network - MobileNet, which significantly reduces the model parameters and computational cost while maintaining high accuracy. Compared with VGG 16, the accuracy of MobileNet on the ImageNet public dataset only drops by 0.9%, while its model size is 32 times smaller than that of VGG 16.

[0097] Different from traditional convolutional neural networks, MobileNet uses depthwise separable convolution method for parameter optimization, separating the steps of feature extraction and feature fusion. First, each channel of the input feature map is separated, and a single convolutional kernel is used to perform a separate convolution operation on each channel. This process is called depthwise convolution. In this process, the depth of each convolutional kernel is fixed at 1, which means that each convolutional kernel only operates on a single channel of the input feature map. This significantly reduces the number of parameters while maintaining the feature extraction ability, because the number of parameters of depthwise convolution is only 1 / N (where N is the number of input channels) of that of standard convolution. However, depthwise convolution cannot expand or compress the feature map along the channel dimension and fails to effectively utilize the feature correlation between different channels at the same spatial position. Therefore, although the computational cost is reduced, the information interaction along the channel dimension is lost. Therefore, after performing depthwise convolution, pointwise convolution is needed to fuse these feature maps. The operation of pointwise convolution is very similar to that of conventional convolution, that is, a 1x1 convolutional kernel is used. The size of its convolutional kernel is 1×1×M, where M is the number of channels of the output information of the previous layer. Therefore, each pointwise convolutional kernel performs a weighted combination of the feature maps of the previous step along the channel direction to generate a new feature map. This step effectively utilizes the feature correlation between different channels at the same spatial position, reducing the parameters without losing information. This form of convolution significantly reduces the number of parameters and computational cost while maintaining the model performance, making it more suitable for mobile devices and embedded devices. Figure 5 shows the processes of standard convolution and depthwise separable convolution, where Figure 5 (a) represents standard convolution and (b) represents depthwise separable convolution.

[0098] Structurally, the MobileNet network adopts depthwise separable convolution modules. First, it uses a feature extraction layer similar to traditional CNNs, replacing the standard convolution module with depthwise convolution. Subsequently, feature fusion is performed through pointwise convolution, normalization layer, and activation function layer. Figure 6 This paper compares the standard convolution block structure and the depthwise separable convolution block structure. The present invention uses the MobileNet model pre-trained on the ImageNet dataset as the base network. ImageNet is a large dataset containing more than 14 million images and 20,000 categories. Its pre-trained model learns more general parameters and exhibits excellent generalization ability, making it applicable to various visual tasks beyond specific domains. Such a large-scale dataset enables the network to learn more general parameters. Using the model pre-trained on this dataset significantly improves the performance and generalization ability of the model, especially suitable for specific tasks with limited data. In addition, since the pre-trained model has learned a large number of image features, applying it to new tasks can accelerate the convergence speed of the model, saving a large amount of time and computational resources.

[0099] 2.2 Model Optimization

[0100] The original MobileNet model was initially used for image classification and was later modified to achieve image regression prediction. To improve the feature extraction ability, structural adjustments and further enhancements were made. The top network layer of the original model, which was responsible for probability output, was removed. The remaining components include: depthwise separable convolutional layers, BN layers, and ReLU layers, which serve as the feature extraction part of the improved model. The input size of the model was set to 224×224×3, and the adjusted output size became 7×7×1024. This structural adjustment maintained the powerful feature extraction ability of the MobileNet network and provided a basis for subsequent customized network structures.

[0101] To further enhance the model's feature representation ability, a SE (Squeeze and Excitation) module was added after the last convolutional block. Conventional convolutional operations mix spatial features and channel features learned through convolution; in contrast, the SE module explicitly models the interdependencies between the channel features of convolutional features. It learns the importance of each channel and weights its features according to the importance of each channel, emphasizing important features and suppressing unimportant features. The SE module consists of three steps: global compression, channel excitation, and scale transformation, as Figure 7 shown. First, the global compression operation reduces the input features from H×W×C to a 1×1×C vector, which can be regarded as the statistical information of global features. The channel excitation step uses two fully connected layers to generate channel attention weights, where the first fully connected layer uses a scaling parameter ratio r to reduce the features, reducing the number of channels from C to C / r, thereby reducing the number of channels and computational costs, and the second fully connected layer restores the number of channels to C and uses the Sigmoid activation function to obtain the weights of each channel. Finally, the scale transformation operation multiplies the weight values by the two-dimensional matrix corresponding to the channels of the original feature map to obtain the final output features after weight adjustment. In this way, after feature extraction, the data is separated from the mixture of spatial features and channel features, enabling the model to adaptively learn the importance of each channel feature and weight the features according to their importance, thereby improving the model's representation ability.

[0102] After the SE module, a Pyramid Pooling Module (PPM) is introduced to further enhance the model's understanding of multi-scale information. The module includes the following steps: Multi-scale partitioning: The input feature map is evenly partitioned into sub-regions of different scales, such as 1x1, 2x2, 3x3, 6x6. Dimensionality reduction processing: Each sub-region is downsampled using a 1x1 convolution to reduce the number of channels, computational complexity, and the number of parameters. Upsampling: The downsampled sub-region feature maps are upsampled to the same size as the original input feature map. Feature fusion: The original input feature map is concatenated with the upsampled sub-region feature maps to form the final feature representation. Through the PPM module, the input feature map extracts information at different scales. The model can capture context information at various scales, thereby enhancing its understanding of the target and improving the model's generalization ability. The Softmax output layer is replaced by a Multi-Layer Perceptron (MLP) model for pose estimation as the prediction output of the network. MLP is a feedforward neural network that contains one or more hidden layers and an output layer, and all layers are fully connected. The MLP of the present invention is designed to be refined and contains three fully connected layers with 1024, 512, and 256 neurons respectively, and the GeLU is used as the activation function in all fully connected layers of this structure. This is a function that combines the tanh function and its approximation, rather than the commonly used ReLU function. This activation function shows a curve biased towards the y-axis when approaching 0 in the left half region, and outputs x itself in the right half region. The GeLU activation function can effectively alleviate the vanishing gradient problem and is superior to other activation functions; the last layer is a fully connected layer with a linear activation, and the output dimension is 3, corresponding to the three predicted spatial pose angles. This MLP model can receive inputs of any shape, and through the recursive calculation of the three fully connected layers, finally outputs a prediction tensor with a batch size of 3 for predicting the spatial pose of the target. In the entire model, the MLP serves as the regression part and performs the spatial pose regression prediction of the input data. Using these methods, a deep learning model is constructed that can effectively predict the spatial pose of target particles. The network architecture is as Figure 8 shown.

[0103] 2.3 Color quantization

[0104] Although the human eye can perceive a vast range of colors from 10,000 to 400 million, typical natural images usually contain far fewer colors than this, often less than 2 million, which is only a small fraction of the colors perceptible to the human eye. Additionally, the red, green, and blue (RGB) components exhibit a high degree of correlation in natural images. For example, the red and green components are usually positively correlated, resulting in the color distribution mainly along the diagonal of the RGB space. Therefore, through color quantization algorithms, natural images can be represented using relatively fewer colors, thereby achieving image compression and efficiency improvement. Although the particle surface pattern in the present invention consists of four colors and the training image set is synthetically generated, images captured by cameras in the real world may have color deviations or noise due to factors such as lighting and shadows, leading to a deviation of the color distribution from the training image set. Color quantization is applied to real-world images to mitigate the impact of colors through color reduction.

[0105] Utilizing the unique advantages of data clustering algorithms in the color quantization application of images and graphics, the present invention employs the K-Means clustering algorithm for color quantization. K-Means is a typical partition-based clustering algorithm belonging to the family of unsupervised learning algorithms. Its concept is very simple: for a given data set, an initial number of clusters (k) and the centroids of the initial clusters are predefined. The Euclidean distance is used to measure the similarity between data objects and the clustering centers. The Euclidean distance between the data object X and the clustering centroid Ci is calculated as follows:

[0106]

[0107] where X is the data object, Ci is the centroid of the i-th cluster, and m is the dimension of the data object; Xj and Cij are the j-th attribute values of X and Ci respectively. Since the similarity is inversely proportional to the distance between data objects, a higher similarity corresponds to a smaller distance. Based on the distances between samples, the data set is divided into multiple clusters. The positions of the clustering centroids are continuously updated according to the similarity between data objects and the clustering centroids. Through iteration, the sum of squared errors (SSE) within the clusters gradually decreases. When the SSE no longer changes or the objective function converges, the clustering terminates, and the final result is obtained. The SSE of the entire data set is calculated as follows:

[0108]

[0109] where k is the number of clusters, and the smaller the SSE, the better the clustering performance.

[0110] In the context of color quantization, the K-means algorithm first selects a predefined number k of color clusters. For each pixel in the image, it calculates the Euclidean distance between the color of the pixel and the centers of the k clusters to which it belongs, and assigns the pixel to the nearest cluster. The average value calculates the values of the three color channels of the pixels in each cluster as the new cluster centers. These two steps are repeated until the cluster centers no longer change or the maximum number of iterations is reached. In this way, a large number of colors in the original image can be reduced to a predefined smaller number of colors, achieving effective quantization while also showing adaptability to image noise and local variations. Comparative experiments show that the Lab color space is more sensitive to color changes than the commonly used RGB color space and can better display the color information in the image. Therefore, the image is first converted from RGB to the Lab color space, then color quantization is performed in the LAB color space, and finally it is converted back to RGB. Figure 9 The color quantization results of the sample image under different numbers of colors are shown, including the LAB and RGB color spaces (the upper row is the LAB color space and the lower row is the RGB color space). It can be seen that when the number of colors is relatively small, the quantization effect in the LAB color space (upper right corner) can better distinguish color blocks and is less affected by changes in lighting conditions, thus more truly reflecting the original color distribution in the image. In contrast, the RGB color space is more sensitive to changes in ambient light, which may lead to inaccurate colors or quantization of some parts of the surface of the target particles.

[0111] Step 3. Experimental Results and Discussion

[0112] 3.1 Model Performance Comparison

[0113] The performance of the model proposed in the present invention is compared with other convolutional neural network architectures, such as the original MobileNet, ResNet50, DenseNet, and VGG16. All models are initialized with pre-trained weights on ImageNet, and the top fully connected layers are replaced to make them suitable for the dataset. An experimental dataset of 1728 images is selected, with an angle range of 0° - 36° (the angle range in this experiment is all 0 - 36°, and the number of 0 - 360° is too large, so partial angles are used for experimental verification), and the angle interval is 3°. This dataset is then divided into 1600 training images and 128 validation images. All models are trained with the same hyperparameter configuration, including batch size, number of epochs, and learning rate. The evaluation metrics used are mean squared error (MSE) and mean absolute error (MAE), and training is stopped when the MSE on the validation set converges. Table 1 shows the performance metrics of different models on the test set, including MSE, MAE, average error for each axis (Err_X, Err_Y, Err_Z), the number of model parameters, and training time.

[0114] Table 1 Performance of Different Networks

[0115]

[0116] The MSE of MobileNetPlus is 0.041, slightly higher than that of ResNet50 (0.022) and VGG16 (0.010), but significantly lower than that of DenseNet (0.239), despite having a similar number of model parameters. Similarly, the MAE of MobileNetPlus (0.169) also falls between ResNet50 and VGG16 and is significantly lower than that of DenseNet. In terms of the errors in three dimensions, MobileNetPlus shows relatively balanced performance, without any obvious weaknesses, and performs relatively well in Err_Y and Err_Z. Although the parameters are slightly higher than those of MobileNet, its training time is only 3 minutes and 16 seconds, significantly less than that of other models, taking less than half or even one-third of the training time of ResNet50, DenseNet, and VGG16. Compared with the original MobileNet, the improved MobileNetPlus performs well in terms of MSE, MAE, Err_X, Err_Y, and Err_Z, with the error reduced by nearly half. Although the number of parameters has increased slightly, the training time is comparable to that of the original MobileNet. Through comparative analysis, MobileNetPlus demonstrates competitiveness in multiple error metrics. It achieves good prediction performance and training efficiency while maintaining a relatively low model complexity, making it competitive in various application scenarios.

[0117] 3.2 Comparison of Generalization Performance

[0118] To further evaluate the generalization ability of the model and eliminate the possibility of overfitting to a specific dataset, further simulation experiments were conducted. In this experiment, a new image dataset was generated. This dataset shares the same angular range (36° spatial angular range) as the dataset used above, but has randomly generated specific spatial pose angles to ensure no regularity. A total of 100 simulated images were generated, and some examples are shown as Figure 10 shown.

[0119] The models trained in the previous section were then tested on this simulated image set. Figure 11Shows the spatial attitude angle prediction errors of different models along three directions (Err_X, Err_Y, Err_Z) and the average error. It can be clearly seen from the figure that the mobile network and the improved mobile network + show excellent performance. Although the mobileNet-based networks may lag behind ResNet and VGG in terms of performance during the training phase, they are superior to other models on the simulated image set, with average errors of 0.349° and 0.308° respectively, and even achieving the lowest average error among all models. The low generalization error of the mobileNet-based networks indicates their robust generalization ability on unseen data. In addition, the error distribution in the three directions is relatively balanced, without significant deviation in any specific direction, which indicates that the network's predictions for all directions are stable. Considering that the real images captured by the camera are used in practical applications, which are significantly different from the training set images, the generalization performance of the model is crucial. The excellent performance of the mobileNet-based networks indicates that they can resist overfitting, making them more suitable for predicting real-world images captured by the camera. In addition, after adding the SE and PPM modules, compared with the original MobileNet, the error is significantly reduced, exceeding the performance of the original network, and even superior to other models in some aspects. This proves the effectiveness of optimizing and improving the model. The following part will evaluate the performance of the model in terms of actual prediction ability.

[0120] Example:

[0121] 1. Experimental system:

[0122] The experiment was conducted in a controlled environment. By fixing the particles on a three-axis (X - Y - Z) angular displacement stage and using an industrial CMOS camera connected to a personal computer to capture images. To ensure that the background color does not affect the image quality, a white polyvinyl chloride sheet was placed between the particles and the angular displacement stage, and a structured acrylic plate with holes was added to further reinforce the particles. Initially, the attitude of the particles was adjusted to align with the 0° - 0° - 0° attitude in the sample library. Then, the angular displacement stage was used to rotate the particles to capture images in different attitudes. A total of 100 images were collected, and the attitude changes, attitude distribution range, and coordinate distribution pattern were all consistent with the sample library. Figure 12 Shows a photo of the entire system setup.

[0123] 2. Image normalization processing

[0124] Predicting directly using images captured by a camera usually yields suboptimal results because the captured images do not match the size of the images used in model training. Model training typically uses images standardized to 224×224 pixels, while real-world images often have larger sizes, and the target area (i.e., the target particle) may differ in size and the proportion of the image it occupies compared to the training data. For example, Figure 13 In (a), the differences in the sizes of images captured by the camera are shown. These images are 2048×1536 pixels in size, while the model requires images of 224×224 pixels. Additionally, in the red image captured by the camera, the particle may only occupy a small part of the image, while in the training image, it occupies a larger part.

[0125] To address this difference, the images were preprocessed, and the target area was extracted using a contour detection algorithm. Subsequently, the target area was resized to 224×224 pixels to meet the input requirements of the neural network. The specific steps are as follows: 1. Image preprocessing. The images were enhanced using the Multi-Scale Retinex with Color Restoration (MSRCR) algorithm. This algorithm adjusted the ratio of the color channels in the original image, highlighting the information in darker areas and effectively solving the problem of color distortion. The processed images showed enhanced local contrast and brightness, being closer to the real scene and thus presenting a more realistic visual effect. 2. Target area localization. The position of the particle in the image was located using the Hough circle transform algorithm. This algorithm can effectively identify circular objects in the image and return the center coordinates and radius of the circular object. 3. Profile extraction and target area extraction. Then, a contour extraction algorithm was used to determine the outermost contour of the particle. Subsequently, the target area was extracted. The image was extracted using fitting and masking algorithms. The fitting ellipse function was used to fit an ellipse to the extracted contour and generate a mask to isolate the target area. 4. Zooming in and out. Finally, the extracted target area was enlarged to 224×224 pixels. Figure 13 In (b), the results after image normalization are shown.

[0126] 3. Model Training and Comparison

[0127] The performance application scenarios of the evaluation model in practice were trained based on the MobileNet network respectively. The training set included 1,600 images with an angle range of 0 - 36° and an angle interval of 3°. A validation set of 128 images with the same distribution was also used. The training parameters were set as follows: 20 epochs, a batch size of 16, and an initial learning rate of 0.0006. A composite loss function was adopted. The mean squared error (MSE) was used when the difference between the predicted value and the true value (y - x) was less than the preset threshold. When the difference (y - x) was equal to or greater than the threshold, the Huber loss function was used, which is more robust to outliers. The weight decay was set to 10^-5. The mean absolute error (MAE) was selected as the evaluation metric for model training. The AdamW optimizer and the cosine annealing scheduler were selected to effectively prevent the model from falling into local optima. The minimum learning rate threshold was set to 0.000006. Figure 14 in (a), Figure 14 in (b) shows the performance curves based on the MobileNet network. The graph shows the training and validation losses as well as the prediction errors. As shown, in the initial training stage, the losses and errors of both networks decreased significantly. In the mid-training stage, the model gradually converged, and the losses and errors continued to decrease. In the late training stage, the losses and prediction errors of the network tended to stabilize. The validation curve gradually converged with the training curve, indicating that the model found a local optimum in the late stage and was not overfitted. Figure 14 in (a) and Figure 14 The comparison between (a) and (b) shows that the validation curve of the original MobileNet showed more obvious fluctuations during the training process. In addition, there was still an obvious gap between the validation curve and the training curve in the late training stage. However, the loss and MAE curves of the optimized MobileNetPlus model had less fluctuations, and the validation curve was closer to the training curve. This indicates that the optimized network can be trained and converge more robustly, achieving a good balance between training performance and validation data. After training was completed, the performance of the model on the validation set was evaluated, and the error distribution was visualized. Figure 15 in (a) shows the validation error distribution of the original MobileNet model. The error distribution ranges from 0° to 0.7°, with uneven distribution and significant outlier deviations. In contrast, Figure 15Figure (b) shows the error distribution of the optimized MobileNetPlus model, which is mainly concentrated between 0° and 0.4°, and is more concentrated between 0.1° and 0.3°. Quantitative analysis shows that the average errors of the original network on each spatial axis are 0.205°, 0.296° and 0.200° respectively. In contrast, the average errors of the optimized network on the three axes are 0.145°, 0.186° and 0.124° respectively, indicating that the average error on each axis has decreased and there are no significant outliers. This further proves the improvement of the regression prediction accuracy after network optimization. Then, the images taken by the single-shot camera are used for pose prediction. The trained network is used to estimate the spatial pose of the preprocessed test set. Table 2 shows the prediction results of the two networks on the test set, where the pose error data on each spatial axis represents the average error. Table 2 shows that the original MobileNet model achieved a relatively low average error of about 2.8° in the pose estimation of real images, but the accuracy is still average. The optimized MobileNetPlus model shows significant improvement, with an average error of about 1.6° on the three axes. This is nearly 60% less than the prediction result of the original model. It is worth noting that the error on the Y-axis (about 1.3°) is lower than the errors on the X-axis and Z-axis in the real image prediction, which contradicts the error distribution observed during the training set validation. This difference indicates that the diversity and complexity of real images may affect the prediction results of the model.

[0128] Table 2 Average error on each spatial axis, and the digits in the following table are in degrees (°).

[0129]

[0130] 4. Image Data Augmentation

[0131] Although the model has been optimized, the prediction accuracy is still insufficient. Inspection shows that some images show significant prediction errors, resulting in outliers, which have a negative impact on the overall prediction accuracy. Through the comparative analysis of the test samples and the image database, two key problems are found: there are differences in the poses captured between some test samples and the database, which hinder the accurate pose estimation of the network; insufficient image preprocessing leads to black borders or incomplete images, misleading the network and resulting in significant prediction errors. Figure 16 Shows two preprocessed test images in the database and their corresponding images. Figure 16 The bottom of the image in (a) has a black border, while Figure 16 the upper right corner of (b) is missing a contour. These defects result in significant prediction errors: Figure 16 In Figure (a), they are 1.6°, 5.3° and 4.6° respectively, Figure 16In the middle figure (b), they are 3.5°, 2.5°, and 3.5° respectively, exceeding the simulation error. These problems may lead to significant attitude prediction biases, thus affecting the overall prediction accuracy. The limited diversity of the training set hinders the model's ability to adapt to the complex and diverse nature of real-world images.

[0132] The model was retrained using image data augmentation techniques to address this challenge. Data augmentation techniques include applying a series of random transformations with different focuses to the training data, generating more training samples to enhance the model's generalization ability and prevent overfitting. This technique plays a crucial role in the training of deep learning networks. There are many data augmentation techniques, including methods based on image transformation and methods based on feature space. In recent years, methods for image expansion using neural networks, such as generative adversarial networks (GANs), have been widely applied. Position transformations such as flipping and rotating are excluded because the position information and texture need to be maintained. The data augmentation techniques adopted in this study include grayscale conversion, randomly cropping 20% of the edge area of the image, randomly adjusting the saturation and brightness, normalizing the image data, and adding random noise. The specific effects of data augmentation are as Figure 17 shown. By performing these operations, the original training images underwent five data augmentation iterations, effectively expanding the training set to six times its original size and significantly increasing the number of samples.

[0133] To maintain the same training parameters, the data augmentation technique was implemented and the model was retrained. The prediction results for the test set before and after data augmentation are as Figure 18 shown. The results show that after data augmentation, the model prediction accuracy has been significantly improved. The errors on the three axes have decreased significantly, with the average error dropping from approximately 1.7° to 0.45°, a reduction of more than 1.2°. This improvement in accuracy is substantial. In addition, the enhanced error distribution is more uniform, and no axis shows an overly high prediction error.

[0134] 5. Image Color Quantization

[0135] Different from the simulated images in the dataset, real-world images captured by cameras usually contain rich color information and may be distorted due to factors such as lighting, exposure, and shadows. To reduce the impact of color variations on prediction accuracy, a color quantization algorithm was applied to the normalized test images. The images in the test dataset were quantized to 4, 16, 32, and 64 colors respectively. Figure 19Shows the color quantization results and their three-dimensional color distributions before and after image processing. Although the surface texture of the target particles consists of black, white, green, and blue, the captured image has color distortion due to insufficient light and other environmental factors, causing the particles to appear green and yellow under camera sampling. The color quantization part corrects the color distortion and restores the distorted yellow to green, as Figure 19 shown, achieving local image restoration. After pose prediction on the test set, quantized to different numbers of colors, a radar chart is used to compare the average errors under different numbers of colors, as Figure 20 shown. This chart compares the average errors along different axes and the overall average error.

[0136] The results show that when the number of colors is set to 4, the overall error is the lowest and the prediction performance is the best. It is worth noting that on the X-axis and Z-axis, the errors under 4-color quantization are significantly lower than those of the original image and other color quantization results. The differences for 1.957 1.265 1.708 0.512 0.371 0.464 0 0.51 1.52 2.5 Error_X Error_Y Error_Z without enhancement error (°) are approximately 0.05°. The trend of the error distribution is also observed. If the number of colors is less, the error distribution will remain balanced and generally decrease. As the number of colors increases, the error distribution gradually approaches that of the original image, i.e., the errors along the X-axis and Z-axis are larger, and the error along the Y-axis is smaller. Eventually, the error distribution is almost indistinguishable from that of the original image. This phenomenon indicates that better processing results can be obtained by minimizing the number of colors while ensuring that the color information of the image is not lost. The fewer the colors, the more uniform the color distribution of the image, making the prediction error distribution more balanced. As the number of colors increases, the image gradually restores its fidelity to the real image, and the error distribution gradually converges to that of the original image.

[0137] The experimental results show that when the number of colors is 4, the overall prediction error is the smallest, resulting in the best prediction performance. The average errors along the X, Y, and Z axes are 0.445°, 0.431°, and 0.405° respectively, and the overall average error is reduced to 0.424°.

[0138] 6. Algorithm Comparison

[0139] The algorithm proposed in the present invention was compared with three traditional image matching algorithms: Scale-Invariant Feature Transform (SIFT), Oriented FAST and Rotated BRIEF (ORB), and Structural Similarity Index Measure (SSIM) using perceptual hashing matching, adopting the template matching method. To improve the matching accuracy, a comprehensive database containing 50,653 images was constructed, covering all poses in the range of 0 - 36°, with a pose interval of only 1°. The target image was matched with the set of images in the database with the highest similarity, and the corresponding three-dimensional spatial angles were determined to estimate the pose of the particle. The results obtained by applying the above algorithms to the test set images were compared with the prediction results of the proposed algorithm. The results are as Figure 21 shown. The results show that the proposed prediction algorithm based on the CNN network has significant advantages. The improved model shows an average prediction time of about 20 milliseconds per image. Compared with the traditional matching algorithms that need to traverse the entire database, this indicates a substantial improvement in speed. In addition, despite the significant speed increase, the average 3D pose error of the model remains stable at about 0.42°. Compared with the traditional matching algorithms, this indicates a considerable improvement in accuracy, as the traditional matching algorithms show significant deviations, especially along the z-axis, with an average error exceeding 10 degrees. Compared with the matching algorithms, the CNN-based prediction method has superior performance in both prediction accuracy and efficiency, achieving high accuracy in a short time.

[0140] A key aspect of this method is the use of the virtual dataset creation method, which simplifies the construction of the dataset, improves efficiency, and provides a high-quality data basis for subsequent network training. The model was further optimized by combining image data augmentation and color quantization algorithms. The optimized algorithm was used to estimate the pose of spherical particles, achieving fast and accurate prediction for each image. The estimation time is only 20 milliseconds, and the average pose error is 0.424°. Compared with the traditional image matching algorithms, this algorithm has significant improvements in both speed and accuracy. This innovative algorithm provides new insights and methods for the research goal of spatial pose measurement.

Claims

1. A method for estimating target visual posture based on convolutional neural network, characterized in that: The steps include: Step 1) Dataset establishment: Spherical virtual particles are used as target objects to generate a virtual sample library, design textures on the surface of virtual particles, and simulate three-dimensional textured spheres; by maintaining a fixed focal length and sampling position, images of virtual particles are captured in different postures; and image information is annotated; Step 2) Data preprocessing: Preprocess the image and use the contour detection algorithm to extract the target area, and adjust the target area to 224×224 pixels to meet the input requirements of the neural network model; Step 3) Neural network model design: Select the neural network model MobileNet, remove the top network of the network model, introduce the SE module and the pyramid pool module PPM to enhance the feature extraction ability of the network model; The feature vectors are integrated using global average pooling (GAP), and the composite multi-layer perceptron (MLP) model is used to replace the Softmax layer model to construct the MobileNetPlus network model. Step 4) Model training: The data set is divided into training set and test set in proportion. The model training parameters are set including cycle, batch size, initial learning rate, and minimum learning rate threshold. A composite loss function is used. When the difference between the predicted value and the true value is less than the preset threshold, the mean square error (MSE) is used as the loss function. When the difference between the predicted value and the true value is equal to or greater than the preset threshold, the Huber loss function is used, and the weight decay coefficient is set. The mean absolute error (MAE) is selected as the evaluation indicator for model training. The AdamW optimizer and cosine annealing scheduler are selected to prevent the model from falling into local optimality. Step 5) Model retraining: Retrain the neural network model using image data augmentation on the training set images. Image data augmentation includes grayscale conversion, edge cropping, random adjustment of saturation and brightness, image data standardization, and addition of random noise. Step 6) Color quantization: In order to reduce the color deviation or noise of the image caused by lighting and shadow factors, which causes the color distribution to deviate from the training image set, the color gamut transformation and K-Means clustering algorithm are used to reprocess the image using image color quantization; Step 7) Pose estimation: The processed test set is input into the retrained neural network model to estimate the spatial posture of the target particles.

2. The method for estimating target visual posture based on convolutional neural network according to claim 1, characterized in that: The specific process of step 1) is as follows: 1.1) Texture design: Based on the texture features, a particle surface texture pattern is designed using Python's VTK simulation library. The design follows the following rules: a. Each view corresponds to a unique direction of spatial coordinates; b. In each view, the ratio of black and white lines is 1:0.9-1.1; c. The annular ratio of the fringe spacing is greater than 5%; d. The stripe directions in adjacent areas should remain consistent; 1.2) Dataset generation and annotation: After the texture design is completed, the VTK simulation library is used to create a data set. The attitude range of the target particles is 360° in the global space attitude angle. Images with different attitudes are quantitatively sampled for each area with a sampling interval of 2°-3°. A series of particle images with different attitudes are generated to create a synthetic sample library. In the process of simulating image generation, a file naming rule is established to encode the three-axis attitude angle information of the image into the file name; the corresponding attitude parameters are extracted by parsing the file name and converted into numerical vectors as the annotation data for each image.

3. The method for estimating target visual posture based on convolutional neural network according to claim 1, characterized in that: The specific process of step 2) is as follows: 3.1) Model selection: The MobileNet model uses a deep separable convolution method for parameter optimization, separating the steps of feature extraction and feature fusion: First, separate each channel of the input feature map and use a single convolution kernel to perform a separate convolution operation on each channel. The depth of each convolution kernel is fixed to 1, that is, each convolution kernel only operates on a single channel of the input feature map. After performing the depthwise convolution, perform point-by-point convolution to fuse the feature maps. The point-by-point convolution operation uses a 1x1 convolution kernel, and the size of the convolution kernel is 1×1×M, where M is the number of channels of the previous layer's output information. Each point-by-point convolution kernel performs a weighted combination of the feature map of the previous step along the channel direction to generate a new feature map. 3.2) Model optimization: Remove some of the top network layers of the original MobileNet model, and retain the depthwise separable convolutional layer, BN layer, and ReLU layer as the feature extraction part of the improved model; The SE module is embedded after the last depth-separable convolution stack structure of the MobileNet model. The SE module consists of three steps: global compression, channel excitation, and scale transformation. First, the global compression operation compresses the input features from the dimensions of H×W×C to a channel descriptor of 1×1×C. The channel excitation step uses two fully connected layers to generate channel attention weights. The first fully connected layer uses a scaling parameter ratio r to reduce the features, reducing the number of channels from C to C / r, thereby reducing the number of channels and computational costs. The second-level fully connected layer restores the channel dimension to the original input channel number C, and generates normalized channel attention weights through the Sigmoid function. The scale transformation operation multiplies the weight value with the two-dimensional matrix corresponding to the original feature map channel to obtain the final output feature after weight adjustment. After the SE module, the pyramid pooling module PPM is introduced. PPM includes multi-scale division: the input feature map is evenly divided into sub-regions of different scales; dimensionality reduction processing: each sub-region is downsampled using 1x1 convolution to reduce the number of channels, reduce computational complexity and the number of parameters; upsampling: the downsampled sub-region feature map is upsampled to the same size as the original input feature map; feature fusion: the original input feature map is spliced ​​with the upsampled sub-region feature map to form the final feature representation; The Softmax output layer is replaced by a multi-layer perceptron (MLP) model for posture estimation as the predicted output of the network; the MLP contains one or more hidden layers and an output layer, all layers are fully connected layers, and all fully connected layers use GeLU as the activation function.

4. The method for estimating target visual posture based on convolutional neural network according to claim 1, characterized in that: The specific process of step 6) is as follows: 6.1) For a given data set, an initial number of clusters k and the centroids of the initial clusters are predefined; the Euclidean distance is used to measure the similarity between data objects and cluster centers; the Euclidean distance between a data object X and a cluster centroid Ci is calculated as follows: Among them, X is the data object, Ci is the centroid of the i-th cluster, and m is the dimension of the data object; Xj and Cij are the j-th attribute values ​​of X and Ci respectively; since the similarity is inversely proportional to the distance between the data objects, a higher similarity corresponds to a smaller distance; based on the distance between samples, the data set is divided into multiple clusters, and the position of the cluster centroid is continuously updated according to the similarity between the data object and the cluster centroid. Through iteration, the square error SSE within the cluster gradually decreases. When the SSE no longer changes or the objective function converges, the clustering is terminated and the final result is obtained; the SSE of the entire data set is calculated as follows: Where k is the number of clusters, the smaller the SSE, the better the clustering performance; 6.2) According to the K-means algorithm, first select the predefined k color clusters; for each pixel in the image, calculate the Euclidean distance between the pixel's color and the centers of the k clusters to which it belongs, and assign the pixel to the nearest cluster; average the three color channel values ​​of the pixel in each cluster as the new cluster center; 6.3) Repeat steps 6.1)-6.2) until the cluster center no longer changes or the maximum number of iterations is reached; 6.4) Convert the image from RGB to Lab color space, then perform color quantization in LAB color space, and finally convert back to RGB.

5. The method for estimating target visual posture based on convolutional neural network according to claim 1, characterized in that: The step 2) is as follows: 2.1) Image preprocessing: The image is enhanced using the multi-scale retinal restoration MSRCR algorithm; 2.2) Target area positioning: The position of particles in the image is located using the Hough circle transform algorithm; 2.3) Profile extraction and target region extraction: The outermost contour of the particle is determined using a contour extraction algorithm, and then the target region is extracted; the image is extracted using a fitting and masking algorithm; an ellipse is fitted to the extracted contour using a fitting ellipse function, and a mask is generated to isolate the target region; 2.4) Zoom in and out: The extracted target area is enlarged to 224×224 pixels.