Global-to-Local Keypoint Localization Method and Device Based on Fusion Attention
The hybrid global-to-local keypoint localization method enhances precision and robustness by integrating global and local feature attention, addressing issues of lighting sensitivity and texture reliance in existing methods.
Patent Information
- Application Number
- CN202411553760.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-03
AI Technical Summary
The existing key point positioning method has large positioning errors in weak texture areas, and ignores the importance of global features to weak texture areas, resulting in a reduced positioning accuracy.
The global to local key point positioning method based on fusion attention is adopted. By constructing a global regression model and a local regression model, combining multi-scale feature extraction and fusion attention modules, key points are accurately positioned stage by stage.
It improves the accuracy and robustness of key point positioning, reduces the number of parameters, enhances feature extraction capabilities, and improves the accuracy and convergence speed of the model.
Smart Images

Figure CN119399491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of picture key - point localization and deep learning technology, in particular to a global - to - local key - point localization method and device based on fused attention. This method uses a deep neural network model to calculate the fused attention of global features and local patch features, and realizes the precise global - to - local stage - by - stage localization of key - points, so as to improve the accuracy and robustness of picture key - point localization. Background Technique
[0002] The key - point localization task refers to identifying some specific and significant points in a given image, which are usually feature points with important information about the image content. For example, in face recognition, the key - points may be the positions of eyes, nose, and mouth, etc.; in object detection, the key - points can be the positions of object edges, corners, etc. The deep - learning method for the key - point localization task uses a deep neural network to extract structural and semantic information from the image, analyze and process its complex features, and thus accurately predict the coordinates of each key - point. It is a key application branch that combines image processing and machine learning in the field of computer vision, and is widely used in scenarios such as human pose estimation and object recognition, providing important technical support for the automation and intelligence of picture - related applications.
[0003] In the existing automatic key - point localization methods, the method using image gradients and second - order derivative matrices is sensitive to image illumination changes and noise, and the method based on feature - point descriptors and matching algorithms cannot flexibly handle large - scale or dynamically changing scenarios. Deep - learning methods usually directly regress the key - point coordinates from image features through heat - map regression or coordinate - regression models. However, direct regression overly relies on the texture features around the key - points, resulting in large key - point localization errors in regions with weak texture features. In addition, when using the attention mechanism to learn the relative relationships between different key - points, the key role of global features in the localization of weak - texture regions is often ignored, thus reducing the accuracy of key - point localization.
[0004] Therefore, in order to reduce the key - point localization error and considering the problem that the local visual features in weak - texture regions are not sufficient to support the direct and accurate localization of picture key - point coordinates, the present invention proposes a global - to - local key - point localization method and device based on fused attention. Summary of the Invention
[0005] The present invention provides a global - to - local key - point localization method and device based on fused attention for the problem existing in the prior art that the local visual features in weak - texture regions are not sufficient to support the direct and accurate localization of key - point coordinates.
[0006] To achieve the above - mentioned invention purpose, the present invention provides the following technical solutions:
[0007] Step 1: Collect images with a camera, annotate their key points, and preprocess the image data to construct a key point localization dataset;
[0008] Step 2: Construct a global regression model; the global regression model mainly includes a backbone network and a heatmap decoding module; the global regression model is constructed by connecting the backbone network and the heatmap decoding module in series; the backbone network extracts the image structure information and deep semantic information in the image representation; the heatmap decoding module generates a probability distribution heatmap for each key point position to complete the regression from the image representation to the heatmap;
[0009] Step 3: Conduct end-to-end supervised training on the global regression model;
[0010] Step 4: Construct a local regression model based on the global regression model; the local regression model mainly includes a local block feature encoding module, a global feature encoding module, and a coordinate decoding module; the local block feature encoding module encodes the local block features through multi-scale image representations according to the two-dimensional spatial position of the forward key point coordinates; the global feature encoding module fuses the multi-scale image representations to extract the global feature representation of the image; the coordinate decoding module uses the local block features to guide the generation of the key point encoding vector from the global features and regresses the encoding vector into a local offset;
[0011] Step 5: Conduct end-to-end joint supervised training on the local regression model and the global regression model;
[0012] Step 6: Input the preprocessed required localization image into the trained global regression model, and the output of the local regression model is the predicted key point coordinates.
[0013] As a preferred solution of the present invention, a global-to-local key point localization method based on fused attention is characterized in that the following processes are included in the step 1:
[0014] S11: Resize the image to a size of n×n and normalize the pixel values of the image;
[0015] S12: With the help of a pre-trained public image bounding box detection model, roughly detect and crop the image to exclude most of the irrelevant background information, and at the same time retain the margin background of k pixel values outside the detection box to ensure that the image information will not be missing due to cropping;
[0016] S13: Calculate the new key point coordinates according to the cropping bounding box in step S12, and at the same time generate a Gaussian heatmap with each key point coordinate of each image as the center by using the two-dimensional Gaussian function;
[0017] S14: Divide the processed dataset into a training set, a test set, and a validation set.
[0018] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fused attention is characterized in that in step 2, the backbone network includes a multi-scale feature extraction module and a feature alignment module; the multi-scale feature extraction module is the feature extraction part of the HRNet network with multiple output branches of different scales, and is composed of a two-dimensional convolutional layer and a ReLu activation layer; the feature alignment module completes feature alignment through interpolation operations;
[0019] Specifically, the picture is preprocessed to obtain three-channel picture data, and the multi-scale feature extraction module extracts multi-scale image information in the picture data by loading a pre-trained HRNet model; the feature alignment module takes the multi-scale image information as input, generates interpolation to supplement pixel values by using neighboring information, and generates spatially aligned image structure information and deep semantic information to better support subsequent heat map decoding and coordinate decoding tasks.
[0020] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fused attention is characterized in that in step 2, the heat map decoding module includes a feature mapping module and a heat map prediction head module; the feature mapping module is composed of a two-dimensional transposed convolutional layer, a batch normalization layer, a ReLu activation function, and a Dropout layer in series; the heat map prediction head module is composed of a two-dimensional convolutional layer;
[0021] Specifically, the feature mapping module fuses and maps image features of different scales to a unified initial heat map representation, and also provides basic features for the subsequent global feature extraction module; the heat map prediction head module fuses the initial heat map representation, maps the possible position of each key point in the global image to a separate probability distribution map, and the position with the largest activation value in each probability distribution map is the predicted key point, which serves as the initial local block center of the subsequent local regression model.
[0022] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fused attention is characterized in that in step 3, the supervised training includes: using the key point heat map in the key point localization dataset as label information and the processed picture as a sample; performing n rounds of iterative training, and saving the model parameters with the smallest loss function value on the test set; the backbone network and the heat map decoding module after training are used as the initial parameter values for subsequent joint training; the loss function used in the supervised training is the mean square error loss.
[0023] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fused attention is characterized in that in step 4, the global feature encoding module includes a feature extraction module and a feature compression module; the global feature extraction module is composed of a large kernel two-dimensional convolutional layer, a batch normalization layer, and a ReLu activation function in series; the feature compression module is composed of an adaptive pooling layer and a linear mapping layer in series;
[0024] Specifically, the global feature extraction module uses large-kernel convolutions to cover a larger area of the image, aggregates features within a larger range, thereby enhancing the understanding of the overall structure and long-range dependencies, capturing more comprehensive overall image information, and extracting the global representation of the image; the feature compression module integrates the global representation and performs dimensionality reduction to ensure that the most important global information in the image is retained, enhances the model's understanding of the global context, and improves the ability to focus on specific regions or features, effectively reducing the size of the feature map and compressing the two-dimensional feature map into a one-dimensional global feature vector to provide global features for subsequent fusion attention calculation.
[0025] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fusion attention is characterized in that the local block feature encoding module in step 4 includes a multi-scale feature fusion module and a block encoding module; the multi-scale feature fusion module consists of a two-dimensional convolutional layer, a batch normalization layer, and a ReLu function; the block encoding module samples and weights the local block pixel values for fusion to generate local block features;
[0026] Specifically, the multi-scale feature fusion module fuses the multi-scale image features aligned by channel attention weighted fusion as the initial feature map of the local block features; the block encoding module uniformly samples n×n points on the fused initial feature map within a square box with a side length of d pixel values centered on the local block to form the local block LP, and calculates the weighted sum of each local block pixel value as the current local block feature Code, specifically:
[0027]
[0028] where i represents the channel number of the local block; LP i represents the i-th channel feature map of the local block; C i represents the encoded value of the i-th channel feature map of the local block, and each channel of the local block is encoded with a real number; ω i,j,k is the pixel value weight at the coordinate (j,k) of the i-th feature map of the local block.
[0029] As a preferred embodiment of the present invention, a global-to-local key point localization method based on fusion attention is characterized in that the coordinate decoding module in step 4 includes a fusion attention module and an offset prediction module; the coordinate decoding module is alternately stacked by multiple fusion attention modules and offset prediction modules; the fusion attention module consists of a self-attention layer, a cross-attention layer, a forward propagation layer, and three two-dimensional position encoding vectors; the offset prediction module consists of 1 single linear mapping layer;
[0030] Specifically, the fusion attention module calculates the self-correlation of the global image features as the self-attention weights through the self-attention layer, and aggregates the global image feature representation. The specific operations are as follows:
[0031]
[0032] Among them, self-attention represents the self-attention weights of the global features; Q t-1 represents the input of the global image features of the current coordinate decoding module; is the query weight matrix of the self-attention module; is the key weight matrix of the self-attention module; is the value weight matrix of the self-attention module; is the feature dimension of the self-attention module; The Softmax function maps the element values to between 0 and 1; P is the global feature position encoding vector; Q S is the aggregated feature; LayerNorm is the layer normalization layer;
[0033] Then, the correlation between the aggregated feature and the local block feature is calculated as the cross-attention weight, and the target block feature is weighted and fused to guide the global feature to approach the local block feature. The specific operations are as follows:
[0034]
[0035] Among them, M t-1 is the input of the local block feature of the current coordinate decoding module; [·] represents the concatenation operation; is the value weight matrix of the self-attention module; is the query weight matrix of the cross-attention module; is the key weight matrix of the cross-attention module; is the value weight matrix of the cross-attention module; is the feature dimension of the cross-attention module; cross-attention represents the cross-attention weight; P′ is the local block feature position encoding vector; R M is the coordinate encoding embedding vector; R is the coordinate encoding vector; Q M is the output after weighting by the cross-attention module;
[0036] The globally weighted feature and the locally weighted feature share the same forward propagation network, reducing the number of parameters while transmitting and updating the global feature and the local block feature. The specific operations are as follows:
[0037]
[0038] Among them, FFN is the forward propagation layer; Dropuout is the Dropuout layer; Q tIs the global feature input for the current coordinate decoding module; M t Is the local block feature output of the current coordinate decoding module;
[0039] The offset prediction module maps the global feature to the key point coordinate offset Offset through linear transformation t , so as to obtain the key point coordinates. The specific operations are as follows:
[0040]
[0041] Among them, OffsetPredictor is the offset prediction module; A t-1 Is the key point coordinate of the pre-input; A t Is the key point coordinate output by the coordinate decoding module.
[0042] As a preferred solution of the present invention, a global-to-local key point localization method based on fused attention is characterized in that the end-to-end joint supervised training in step 5 is divided into three stages, and the training process is as follows:
[0043] S91: In the first stage, the position with the maximum activation of the key point heat map generated by the pre-trained global regression model is selected as the initial local block center;
[0044] S92: Generate local block features by sampling and encoding from the aligned multi-scale image features according to the local block center;
[0045] S93: Input the local block features and the global feature vector into the coordinate decoding module to obtain the predicted key point coordinates of this stage;
[0046] S94: In the second and third stages, while the side length d of the local block is halved, the predicted coordinates of the local regression model in the previous stage are used as the initial local block center, and step S92 is repeated;
[0047] S95: Collect the heat map output by the global regression model and the key point coordinates output by the local regression model in each stage, and use the joint loss function L to optimize the model parameters. Save the model parameters that minimize the value of the joint loss function L on the test set. The joint loss function L used is:
[0048]
[0049] Among them, stage i Represents the i-th stage of the joint supervised training; L H Is the global regression model loss; Is the loss of the local regression model in the i-th stage, using the standardized mean error; λ H Is the weight coefficient of the global regression model loss function, is the weight coefficient of the loss of the i-th stage of the local regression model.
[0050] A global-to-local key point positioning device based on fused attention, characterized in that the device includes: an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory corresponds to instructions specified by at least one processor at the same time, and the instructions are executed by the at least one executor so that the at least one processor can execute the method according to any one of claims 1 to 9.
[0051] (1) The present invention uses the HRNet network to extract multi-scale visual features contained in the picture, greatly reducing the loss of original image information; at the same time, the proposed fused attention module deeply fuses the global features and local block features through attention weights, maps the global features to the local block feature space, and then uses the fused global features to predict the offset of the key point in the local block, greatly enhancing the positioning accuracy of the key point in the local block.
[0052] (2) The present invention globally locates the key points on the entire picture through the global regression model, and on this basis, constructs a local regression model to achieve more accurate positioning of the key point coordinates in the local block. The local block size gradually decreases in three stages to achieve precise positioning stage by stage, greatly enhancing the positioning accuracy and robustness; by pre-training the global regression model and jointly training it with the local regression model on this basis, compared with separate training and direct joint training, this method speeds up the convergence rate of the local regression module and further enhances the accuracy of the model.
[0053] (3) The global feature extraction module proposed by the present invention acts on the basic features in the middle of the heatmap decoding module. Under the guidance of the heatmap, the basic features retain global information while focusing on the key point positioning task. Compared with directly extracting global features from multi-scale image information, the number of parameters is greatly reduced and the feature extraction ability is enhanced; at the same time, the global features participate in the prediction of the local block offset, ensuring the consistency of the global positioning of the heatmap and the target of the local block offset regression task. Description of the Drawings
[0054] Figure 1 It is a schematic flow chart of a global-to-local key point positioning method based on fused attention according to Embodiment 1 of the present invention.
[0055] Figure 2 It is a schematic diagram of a global regression model in a global-to-local key point positioning method based on fused attention according to Embodiment 1 of the present invention.
[0056] Figure 3Schematic diagram of the global feature encoding module of a global-to-local key point localization method based on fused attention according to Embodiment 1 of the present invention.
[0057] Figure 4 Schematic diagram of the local block feature encoding module of a global-to-local key point localization method based on fused attention according to Embodiment 1 of the present invention.
[0058] Figure 5 Schematic diagram of the coordinate decoding module in a global-to-local key point localization method based on fused attention according to Embodiment 1 of the present invention.
[0059] Figure 6 Schematic diagram of the local regression model in a global-to-local key point localization method based on fused attention according to Embodiment 1 of the present invention.
[0060] Figure 7 Schematic diagram of a global-to-local key point localization device based on fused attention according to Embodiment 2 of the present invention. Detailed implementation manners
[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] Embodiment 1
[0063] As Figure 1 shown, a global-to-local key point localization method based on fused attention includes the following steps:
[0064] S1: Collect pictures with a camera and label their acupoints.
[0065] S2: Preprocess the picture data to construct an acupoint localization data set.
[0066] S3: Construct a global regression model, which includes a backbone network and a heatmap decoding module; the global regression model is constructed by connecting the backbone network and the heatmap decoding module in series.
[0067] S4: Input the preprocessed picture data into the global regression model for forward propagation to obtain a predicted heatmap, train and optimize the weights of the global regression model through the mean square error loss function, and finally obtain the optimal weights.
[0068] S5: Construct a local regression model based on the global regression model, where the local regression model includes a local block feature encoding module, a global feature extraction module, and a coordinate decoding module.
[0069] S6: Train and optimize the parameter weights of the local regression model through the combined loss function, while fine-tuning the parameter weights of the global regression model, and save the optimal weights;
[0070] S7: Load the optimal weights into the global regression model and input the preprocessed image to be located, and the corresponding coordinates of the acupoints can be obtained.
[0071] Further, the preprocessing in step S2 includes the following specific steps:
[0072] S21: Resize the image to a size of 256×256 and normalize the pixel values of the image;
[0073] S22: With the help of the FaceBoxesV2 face bounding box detection model pre-trained on 12,880 face images in the Wider Face subset, roughly detect and crop the face image, exclude most of the irrelevant background information, and at the same time retain a 5-pixel margin background outside the detection box to ensure that the face information will not be missing due to cropping;
[0074] S23: Calculate the new acupoint coordinates according to the cropping bounding box in step S22, and at the same time use the two-dimensional Gaussian function to generate a Gaussian heat map with each acupoint coordinate of each face image as the center;
[0075] S24: Divide the processed data set into a training set, a test set and a validation set.
[0076] Further, as Figure 2 shown, the backbone network includes a multi-scale feature extraction module and a feature alignment module; the multi-scale feature extraction module is the feature extraction part of the HRNetV2-18 network with 4 different scale output branches, mainly composed of a two-dimensional convolutional layer and a ReLu activation layer; the feature alignment module completes feature alignment through interpolation operations;
[0077] Specifically, the preprocessed image obtains three-channel image data. The multi-scale feature extraction module extracts multi-scale image information in the image data by loading the HRNetV2-18 model pre-trained on the COCO 2017 data set; the feature alignment module takes the multi-scale image information as input, generates interpolation to supplement pixel values by using adjacent information, and generates spatially aligned image structure information and deep semantic information to better support subsequent heat map decoding and coordinate decoding tasks.
[0078] Further, as Figure 2 shown, the heat map decoding module includes a feature mapping module and a heat map prediction head module; the feature mapping module is composed of a two-dimensional transposed convolutional layer, a batch normalization layer, a ReLu activation function and a Dropout layer in series; the heat map regression module is composed of a two-dimensional convolutional layer;
[0079] Specifically, the feature mapping module fuses image features of different scales and maps them to a unified initial heatmap representation, enabling the pixels at each position of the heatmap to reflect the feature intensity and position of the relevant region in the image. It not only accurately reproduces the initial image features but also provides basic features for subsequent global feature extraction. The heatmap prediction head module fuses the initial heatmap representation and maps the possible positions of each acupoint globally in the image into a separate probability distribution map. The position with the maximum activation value in each probability distribution map is the predicted coordinate of each acupoint, serving as the initial local block center for the subsequent local regression model.
[0080] Furthermore, as Figure 2 shown, the supervised training in step 3 includes: using the acupoint heatmap in the face acupoint localization dataset as label information and the processed pictures as samples; performing n rounds of iterative training and saving the model parameters with the minimum loss function value on the test set; using the backbone network and the heatmap decoding module after training as the initial parameter values for subsequent joint training; and using the mean squared error loss as the loss function for the supervised training.
[0081] Furthermore, as Figure 3 shown, the global feature encoding module includes a feature extraction module and a feature compression module. The global feature extraction module is composed of a large kernel two-dimensional convolutional layer, a batch normalization layer, and a ReLu activation function in series. The feature compression module is composed of an adaptive pooling layer and a linear mapping layer in series.
[0082] Specifically, the global feature extraction module uses large kernel convolutions to cover a larger area of the image, aggregates features within a larger range, thereby enhancing the understanding of the overall structure and long-term dependencies, capturing more comprehensive overall image information, and extracting the global representation of the image. The feature compression module integrates the global representation and performs dimensionality reduction to ensure that the most important global information in the image is retained, enhances the model's understanding of the global context, and improves the ability to focus on specific regions or features, effectively reducing the size of the feature map and compressing the two-dimensional feature map into a one-dimensional global feature vector, providing global features for subsequent fusion attention calculation.
[0083] Furthermore, as Figure 4 shown, the local block feature encoding module includes a multi-scale feature fusion module and a block encoding module. The multi-scale feature fusion module is composed of a two-dimensional convolutional layer, a batch normalization layer, and a ReLu function. The block encoding module samples and weights the pixel values of the local block to generate local block features.
[0084] Specifically, the multi-scale feature fusion module fuses the multi-scale image features after alignment through channel attention weighted fusion as the initial feature map of the local block features; the block encoding module uniformly samples n×n points on the fused initial feature map within a square box with a side length of d pixel values centered on the local block to form the local block LP, and calculates the weighted sum of each local block pixel value as the current local block feature Code, specifically:
[0085]
[0086] where i represents the channel number of the local block; LP i represents the i-th channel feature map of the local block; C i represents the encoding value of the i-th channel feature map of the local block, and each channel of the local block is encoded with a real number; ω i,j,k is the pixel value weight at the coordinate (j,k) of the i-th feature map of the local block.
[0087] Furthermore, as Figure 5 shown, the coordinate decoding module includes a fusion attention module and an offset prediction module; the coordinate decoding module is composed of multiple fusion attention modules and offset prediction modules stacked alternately; the fusion attention module consists of a self-attention layer, a cross-attention layer, a forward propagation layer, and three two-dimensional position encoding vectors; the offset prediction module consists of 1 single linear mapping layer;
[0088] Specifically, the fusion attention module calculates the self-correlation of the global image features as the self-attention weight through the self-attention layer, and aggregates the global image feature representation. The specific operations are as follows:
[0089]
[0090] where self-attention represents the self-attention weight of the global features; Q t-1 represents the input of the global image features of the current coordinate decoding module; is the query weight matrix of the self-attention module; is the key weight matrix of the self-attention module; is the value weight matrix of the self-attention module; is the feature dimension of the self-attention module; the Softmax function maps the element values to between 0 and 1; P is the global feature position encoding vector; Q S is the aggregated feature; LayerNorm is the layer normalization layer;
[0091] Then, the correlation between the aggregated feature and the local block feature is calculated as the cross-attention weight, and the target block feature is weighted and fused to guide the global feature to approach the local block feature. The specific operations are as follows:
[0092]
[0093] Among them, M t-1 is the local block feature input of the current coordinate decoding module; [·] represents the concatenation operation; is the query weight matrix of the cross-attention module; is the key weight matrix of the cross-attention module; is the value weight matrix of the cross-attention module; is the feature dimension of the cross-attention module; cross-attention represents the cross-attention weight; P′ is the local block feature position encoding vector; R M is the coordinate encoding embedding vector; R is the coordinate encoding vector; Q M is the output after weighting by the cross-attention module;
[0094] The global weighted feature and the local weighted feature share the same forward propagation network, reducing the number of parameters while transmitting and updating the global feature and the local block feature. The specific operations are as follows:
[0095]
[0096] Among them, FFN is the forward propagation layer; Dropuout is the Dropuout layer; Q t is the global feature input of the current coordinate decoding module; M t is the local block feature output of the current coordinate decoding module;
[0097] The offset prediction module maps the global feature to the key point coordinate offset Offset t through linear transformation, so as to obtain the key point coordinates. The specific operations are as follows:
[0098]
[0099] Among them, OffsetPredictor is the offset prediction module; A t-1 is the key point coordinate of the pre-input; A t is the key point coordinate output by the coordinate decoding module.
[0100] Furthermore, as Figure 6 shown, the global regression model and the local regression model adopt an end-to-end joint supervision training method. The training is divided into three stages. The training process is as follows:
[0101] S91: In the first stage, the position with the maximum activation of the acupoint heat map generated by the pre-trained global regression model is selected as the initial local block center;
[0102] S92: Generate local block features by sampling and encoding from the aligned multi-scale image features according to the local block center;
[0103] S93: Input the local block features and the global feature vector into the coordinate decoding module to obtain the predicted acupoint coordinates at this stage;
[0104] S94: In the second and third stages, while halving the side length d of the local block, use the predicted coordinates of the local regression model in the previous stage as the initial local block center of this stage, and repeat step S92;
[0105] S95: Collect the heatmap output by the global regression model and the acupoint coordinates output by the local regression model at each stage, and optimize the model parameters using the joint loss function L. Save the model parameters that minimize the value of the joint loss function L on the test set. The joint loss function L used is:
[0106]
[0107] where stage i represents the i-th stage of the joint supervised training; L H is the loss of the global regression model; is the loss of the local regression model at the i-th stage, using the standardized mean error; λ H is the weight coefficient of the loss function of the global regression model, is the weight coefficient of the loss of the local regression model at the i-th stage.
[0108] Embodiment 2
[0109] As Figure 7 shown, the global-to-local key point localization device based on fused attention includes an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory simultaneously corresponds to instructions specified by at least one processor, and the instructions are executed by the at least one executor. The input-output interface includes a display, a keyboard, a mouse, and a USB interface for completing data interaction operations; the power supply can be an external power supply or a rechargeable battery to provide electrical energy for the electronic device.
[0110] Those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as a removable storage device, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0111] When the above-mentioned integrated units of the present invention are implemented in the form of software functional units and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.
[0112] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A global-to-local key point localization method based on fused attention, characterized in that, Including the following steps: Step 1: Use a camera to collect pictures and label their key points, preprocess the picture data, and construct a key point localization data set; Step 2: Construct a global regression model; the global regression model includes a backbone network and a heatmap decoding module; The global regression model is constructed by concatenating the backbone network and the heatmap decoding module; the backbone network extracts the image structure information and deep semantic information in the image representation; the heatmap decoding module generates a probability distribution heatmap of the position of each key point to complete the regression from the image representation to the heatmap; the backbone network includes a multi-scale feature extraction module and a feature alignment module; The multi-scale feature extraction module is the feature extraction part of the HRNet network with multiple output branches of different scales, and is composed of a two-dimensional convolutional layer and a ReLu activation function; the feature alignment module completes feature alignment through interpolation operations; the heatmap decoding module includes a feature mapping module and a heatmap prediction head module; The feature mapping module is composed of a two-dimensional transposed convolutional layer, a batch normalization layer, a ReLu activation function, and a Dropout layer in series; the heatmap prediction head module is composed of a two-dimensional convolutional layer; Step 3: Perform end-to-end supervised training on the global regression model; Step 4: Construct a local regression model based on the global regression model; the local regression model includes a local block feature encoding module, a global feature encoding module, and a coordinate decoding module; The local block feature encoding module encodes the local block features through multi-scale image representations according to the two-dimensional spatial position of the forward key point coordinates; The global feature encoding module fuses multi-scale image representations and extracts the global feature representation of the image; The coordinate decoding module uses the local block features to guide the global features to generate key point encoding vectors and regresses the encoding vectors into local offsets; The global feature encoding module includes a feature extraction module and a feature compression module; The global feature extraction module is composed of a large kernel two-dimensional convolutional layer, a batch normalization layer, and a ReLu activation function in series; the feature compression module is composed of an adaptive pooling layer and a linear mapping layer in series; The local block feature encoding module includes a multi-scale feature fusion module and a block encoding module; the multi-scale feature fusion module is composed of a two-dimensional convolutional layer, a batch normalization layer, and a ReLu function; the block encoding module samples and weighted fuses the local block pixel values to generate local block features; The coordinate decoding module includes a fusion attention module and an offset prediction module; The coordinate decoding module is composed of multiple fusion attention modules and offset prediction modules stacked alternately; the fusion attention module is composed of a self-attention layer, a cross-attention layer, a forward propagation layer, and three two-dimensional position encoding vectors; the offset prediction module is composed of 1 single linear mapping layer; Step 5: Perform end-to-end joint supervised training on the local regression model and the global regression model; Step 6: Input the preprocessed required localization picture into the trained global regression model, and the output of the local regression model is the predicted key point coordinates.
2. The global-to-local key point localization method based on fusion attention according to claim 1, wherein The preprocessing in the said Step 1 includes the following processes: S11: Adjust the picture to a size of and normalize the pixel values of the picture; S12: With the help of a pre-trained public image bounding box detection model, roughly detect and crop the image to exclude most of the irrelevant background information, while retaining the margin background of k pixel values outside the detection box to ensure that the image information will not be missing due to cropping; S13: Calculate the new key point coordinates according to the cropping bounding box in step S12, and at the same time use the two-dimensional Gaussian function to generate a Gaussian heat map centered on each key point coordinate of each image; S14: Divide the processed data set into a training set, a test set and a validation set.
3. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The picture is preprocessed to obtain three-channel picture data, and the multi-scale feature extraction module extracts multi-scale image information in the picture data by loading a pre-trained HRNet model; The feature alignment module takes the multi-scale image information as input, generates interpolation to supplement pixel values by using adjacent information, and generates spatially aligned image structure information and deep semantic information to better support subsequent heat map decoding and coordinate decoding tasks.
4. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The feature mapping module fuses image features of different scales and maps them to a unified initial heat map representation, and also provides basic features for the subsequent global feature extraction module; the heat map prediction head module fuses the initial heat map representation and maps the position of each key point globally in the image to a separate probability distribution map. The position with the largest activation value in each probability distribution map is the predicted key point, which serves as the initial local block center of the subsequent local regression model.
5. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The supervised training in step 3 includes: using the key point heat map in the key point localization data set as label information and the processed picture as a sample; performing n rounds of iterative training, and saving the model parameters with the smallest loss function value on the test set; the backbone network and the heat map decoding module after training are used as the initial parameter values for subsequent joint training; the loss function used in the supervised training is the mean square error loss.
6. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The global feature extraction module uses large kernel convolution to cover a larger area of the image, aggregates features in a larger range, and extracts the global representation of the image; the feature compression module integrates the global representation and performs dimensionality reduction to ensure that the most important global information in the image is retained, effectively reducing the size of the feature map, and compressing the two-dimensional feature map into a one-dimensional feature vector to provide global features for subsequent fused attention calculation.
7. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The multi-scale feature fusion module fuses the multi-scale image features after alignment through channel attention weighted fusion as the initial feature map of the local block features; the block encoding module evenly samples point pixels on the fused initial feature map within a square box with a side length of d pixel values centered on the local block to form the local block , and calculates the weighted sum of the pixel values of each local block as the current local block feature Code, specifically: ; Among them, i represents the channel number of the local block; represents the i-th channel feature map of the local block; represents the encoding value of the i-th channel feature map of the local block, and each channel of the local block is encoded with a real number; is the coordinate of the i-th feature map of the local block pixel value weight at the position.
8. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The fused attention module calculates the self-correlation of the global features of the image as the self-attention weight through the self-attention layer, and aggregates the global feature representation of the image. The specific operations are as follows: ; Among them, self-attention represents the self-attention weight of global features; represents the input of the global image features of the current coordinate decoding module; is the query weight matrix of the self-attention module; is the key weight matrix of the self-attention module; is the value weight matrix of the self-attention module; is the feature dimension of the self-attention module; the Softmax function maps the element values between 0 and 1; is the global feature position encoding vector; is the aggregated feature; LayerNorm is the layer normalization layer; Then, the correlation between the aggregated feature and the local block feature is calculated as the cross-attention weight, and the target block feature is weighted and fused to guide the global feature to approach the local block feature. The specific operations are as follows: ; Among them, is the local block feature input of the current coordinate decoding module; represents the concatenation operation; is the value weight matrix of the self-attention module; is the query weight matrix of the cross-attention module; is the key weight matrix of the cross-attention module; is the value weight matrix of the cross-attention module; is the feature dimension of the cross-attention module; cross-attention represents the cross-attention weight; is the local block feature position encoding vector; is the coordinate encoding embedding vector; R is the coordinate encoding vector; is the output after weighting by the cross-attention module; The globally weighted feature and the locally weighted feature share the same forward propagation network, reducing the number of parameters while transmitting and updating the global feature and the local block feature. The specific operations are as follows: ; Among them, is the forward propagation layer; is layer; is the global feature input of the current coordinate decoding module; is the local block feature output of the current coordinate decoding module; The offset prediction module maps the global features to the key point coordinate offsets through linear transformation , so as to obtain the key point coordinates. The specific operations are as follows: ; Among them, OffsetPredictor is an offset prediction module; is the key point coordinate of the pre-input; is the key point coordinate output by the coordinate decoding module.
9. A global-to-local key point localization method based on fused attention according to claim 1, characterized in that The end-to-end joint supervised training in step 5 is divided into three stages. The training process is as follows: S91: In the first stage, the position with the maximum activation of the key-point heatmap generated by the pre-trained global regression model is selected as the initial local block center; S92: According to the local block center, local block features are sampled and encoded from the aligned multi-scale image features; S93: The local block feature and the global feature are input into the coordinate decoding module to obtain the predicted key-point coordinates in this stage; S94: In the second and third stages, while the side length d of the local block is halved, the predicted coordinates of the local regression model in the previous stage are used as the initial local block center, and step S92 is repeated; S95: The heatmap output by the global regression model and the key-point coordinates output by the local regression model in each stage are collected, and the joint loss function L is used to optimize the model parameters. The model parameters that minimize the value of the joint loss function L on the test set are saved. The joint loss function L used is: ; Among them, represents the i-th stage of joint supervised training; is the loss of the global regression model; is the loss of the local regression model in the i-th stage, using the standardized mean error; is the weight coefficient of the loss function of the global regression model, is the weight coefficient of the loss of the local regression model in the i-th stage.
10. An apparatus using a global-to-local key point localization method based on fused attention as described in claim 1, characterized in that, The device includes: an input device, an output device, a power supply, at least one processor, and a memory communicatively connected to the processor; the memory simultaneously corresponds to instructions specified by at least one processor, and the instructions are executed by the at least one executor so that the at least one processor can execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Face alignment method and system based on local-global fusion attention
CN116665276A
Lightweight pedestrian re-identification method based on double-branch fusion attention mechanism
CN118116029A