Real-time 2D key point detection method and device based on Mama
By introducing a Mamba-based network structure in two-dimensional key point detection, combining the context modeling module and the two-dimensional selective scanning module, the problem of difficulty in taking into account accuracy and real-time in the existing technology is solved, and efficient and accurate key point detection is achieved.
Patent Information
- Application Number
- CN202411949509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The existing two-dimensional key point detection methods are difficult to balance the accuracy and real-time. The CNN-based method lacks structural information modeling capabilities, while the Transformer-based method is difficult to meet the needs of real-time scenarios due to its high computational complexity.
A real-time 2D key point detection method based on Mamba is proposed. By combining convolutional neural network and Mamba network structure, the context modeling module and the two-dimensional selective scanning module are used to improve the positioning accuracy and calculation efficiency of key points.
It realizes an effective balance between real-time and detection accuracy, improves the positioning accuracy of key points, optimizes computing efficiency, is suitable for resource-constrained scenarios, and is suitable for embedded devices and real-time applications.
Smart Images

Figure CN119942638A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of two-dimensional key point detection, and in particular to a real-time 2D key point detection method and device based on Mamba. Background Art
[0002] Two-dimensional keypoint detection has made significant progress in recent years, mainly due to the development of deep learning technology. Among the mainstream methods, methods based on convolutional neural networks (CNN) and Transformer dominate.
[0003] CNN-based methods are the earliest mainstream methods used for 2D key point detection. These methods can be divided into regression-based and heat map-based methods. Regression-based methods attempt to directly predict the 2D coordinates of key points, such as the DeepPose and IEF methods proposed by some scholars. These methods have certain advantages in efficiency by directly learning the mapping relationship from images to coordinates. However, directly predicting coordinates usually loses the structural relationship between key points, which affects the detection accuracy.
[0004] In contrast, heatmap-based methods are widely used because they preserve the structural information of the image and simplify the training process by predicting the heatmap of each key point. For example, HRNet and SBL proposed by some research teams represent a high level of heatmap-based methods. These methods achieve a high level of accuracy by gradually refining the heatmap generation process. However, these methods usually require large computing resources, which poses a challenge to real-time applications.
[0005] With the widespread application of Transformer in the field of computer vision, many scholars have also tried to introduce it into the task of two-dimensional key point detection. For example, TFPose proposed by some scholars is one of the earliest works to apply Transformer to pose estimation, modeling the long-distance dependencies of key points through the global attention mechanism. Subsequently, some teams developed PRTR, combining cascaded transformers with regression methods to improve the accuracy of pose estimation. In addition, some studies further optimized the structure of Transformer, such as ViTPose, which improved the feature representation ability and accuracy of the model by introducing visual Transformer. However, the quadratic computational complexity of Transformer makes it have a high demand on hardware resources, which limits its widespread application in real-time scenarios.
[0006] In the task of real-time two-dimensional key point detection, researchers have proposed a variety of optimization strategies to reduce the computational complexity of the model. Some methods use pruning technology to remove redundant layers or irrelevant modules to improve the efficiency of model reasoning; other methods use lightweight backbone networks, such as EfficientNet and YOLO, to quickly extract key point features. For example, methods such as RTMPose and RTMO proposed by some teams have made significant progress in the accuracy and efficiency of real-time detection. However, these methods often require a trade-off between accuracy and efficiency and cannot take both into account at the same time.
[0007] In general, existing two-dimensional key point detection methods have their own advantages in terms of accuracy and real-time performance, but they still have shortcomings. The CNN-based method lacks the ability to model structural information, while the Transformer-based method has high computational complexity and is difficult to meet the needs of real-time scenarios. Therefore, how to design an efficient and accurate key point detection network has become an urgent problem to be solved in this field. Summary of the invention
[0008] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0009] To this end, the first objective of this application is to propose a real-time 2D key point detection method based on Mamba.
[0010] The second objective of this application is to propose a real-time 2D key point detection device based on Mamba.
[0011] The third objective of the present application is to provide an electronic device.
[0012] A fourth objective of the present application is to provide a computer-readable storage medium.
[0013] A fifth object of the present application is to provide a computer program product.
[0014] To achieve the above objectives, the first embodiment of the present application proposes a real-time 2D key point detection method based on Mamba, comprising:
[0015] The input image is received through the Stem network based on the convolutional neural network, and the initial posture features are extracted to obtain preliminary features;
[0016] Processing the initial features using a Mamba-based encoder to output multi-level posture features, wherein the encoder includes multiple stages, each stage consisting of a context modeling module, a two-dimensional selective scanning module, and a normalization layer;
[0017] The multi-level posture features output by the encoder are upsampled into key point heat maps using a decoder, and the heat map of each key point represents the key point position of the target instance.
[0018] Optionally, the Stem network is used to:
[0019] A convolution layer with a convolution kernel size of 7, a batch normalization layer, and a ReLu activation function layer are used to initially capture the low-level features in the input image. The expression is:
[0020]
[0021] Two deep convolution layers with a convolution kernel size of 3, a batch normalization layer, and a ReLU activation function layer are used to remove redundant information in the original input image and further extract posture-related features. The initial features extracted by the Stem network are It is expressed as:
[0022]
[0023] Among them, X I is the input image, H and W are the height and width of the input image respectively, d s is the feature dimension, Conv(·) represents the processing of the convolutional layer, DWCnv(·) represents the processing of the deep convolutional layer, BN(·) represents the processing of the batch normalization layer, and ReLU(·) represents the processing of the ReLU activation function.
[0024] Optionally, each stage of the encoder selects a different feature downsampling rate, and the processing process of each stage is:
[0025]
[0026] in, is the input of the i-th stage. If i=1, then is the output of the i-th stage; LN(·) represents the processing of the normalization layer; SS2D(·) represents the processing of the two-dimensional selective scanning module. The i-th stage has N i A two-dimensional selective scanning module; CMM(·) represents the processing process of the context modeling module.
[0027] Optionally, the context modeling module comprises a deep convolutional layer, a patch embedding layer and a linear layer, wherein:
[0028] The deep convolutional layer is used to capture the dependencies between image patches within the receptive field and downsample the input features;
[0029] The patch embedding layer is used to perform patch division and encoding on the features extracted by the deep convolutional layer;
[0030] The linear layer is used to map the context features obtained by the patch embedding layer into the representation space of the matching Mamba patch input;
[0031] The calculation process of the context modeling module is:
[0032]
[0033] in, Represents the context features learned by the context modeling module; Embed(·) represents the processing process of the patch embedding layer, LN(·) represents the processing process of the linear layer, and DWConv(·) represents the processing process of the deep convolutional layer.
[0034] Optionally, the two-dimensional selective scanning module is used to:
[0035] The image patches are rearranged through sequential scanning strategies in four directions, the information between the image patches is integrated, and the attention weights are calculated through context features to activate the key image areas. Each scanning strategy obtains the adaptive matrix and parameter matrix of the state space model through a linear layer, and obtains the discretized parameter matrix through a discretization process.
[0036] The state at the current moment is calculated using the discretized parameter matrix and the state at the previous moment;
[0037] The calculation process of the two-dimensional selective scanning module is:
[0038]
[0039] Among them, LN(·) represents the processing of the linear layer, Δ is the adaptive matrix, and is the state matrix, Discretize(·) represents the discretization process, p k Indicates the current state output, p k-1 Indicates the state output at the previous moment, Represents the output characteristics of the two-dimensional selective scanning module Features extracted by the context modeling module.
[0040] Optionally, the decoder includes two deconvolution layers with a kernel size of 4, and the calculation formula of the decoder is:
[0041]
[0042] Wherein, H is the key point heat map, is the multi-level posture feature output by the encoder, K is the number of key points, and Decoder(·) represents the decoding process of the decoder.
[0043] Optionally, also include:
[0044] The detection method is optimized using the mean square error loss function, and its loss function is expressed as:
[0045]
[0046] Among them, M is the total number of training samples, H i and Represent the predicted key point heat map and true value of the i-th sample respectively.
[0047] To achieve the above-mentioned purpose, the second aspect of the present application proposes a real-time 2D key point detection device based on Mamba, comprising:
[0048] A feature extraction module is used to receive an input image through a Stem network based on a convolutional neural network, and extract initial posture features to obtain preliminary features;
[0049] An encoding module, configured to process the initial features using a Mamba encoder and output multi-level posture features, wherein the encoder includes multiple stages, each of which is composed of a context modeling module, a two-dimensional selective scanning module, and a normalization layer;
[0050] A decoding module is used to use a decoder to upsample the multi-level posture features output by the encoder into a key point heat map, and the heat map of each key point represents the key point position of the target instance.
[0051] To achieve the above-mentioned purpose, the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0052] The memory stores computer-executable instructions;
[0053] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.
[0054] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a computer-readable storage medium, in which computer-readable storage medium is stored computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method as described in any one of the first aspects.
[0055] To achieve the above-mentioned purpose, the fifth aspect of the present application proposes a computer program product, which implements any method in the first aspect when executed by a processor.
[0056] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:
[0057] This application achieves an effective balance between real-time performance and detection accuracy by introducing a lightweight Mamba network structure; by combining the context modeling module and the two-dimensional selective scanning module, the positioning accuracy of key points is improved while optimizing the computational efficiency. Compared with traditional methods, this application also enhances the robustness of the model by retaining the structural constraints between key points, adapting to complex scenarios and task requirements, and providing an innovative solution for the field of key point detection. This application can run efficiently in resource-constrained scenarios and is suitable for embedded devices and real-time applications.
[0058] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0060] Figure 1 A flowchart of a real-time 2D key point detection method based on Mamba provided in an embodiment of the present application;
[0061] Figure 2 A flow chart of a real-time 2D key point detection method based on Mamba provided in an embodiment of the present application;
[0062] Figure 3 A schematic structural diagram of a Mamba-based real-time 2D key point detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0063] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0064] In view of the problems existing in the prior art, the present application provides a real-time 2D key point detection method based on Mamba. Figure 1 A flowchart of a real-time 2D key point detection method based on Mamba provided in an embodiment of the present application is shown in FIG. Figure 2 A process framework diagram of a real-time 2D key point detection method based on Mamba provided in an embodiment of the present application.
[0065] In addition, in order to facilitate the subsequent description of the method proposed in this application, it is named PoseMamba2D. PoseMamba2D adopts a top-down paradigm and is a simple and effective pose estimation framework.
[0066] like Figure 1 As shown, the method comprises the following steps:
[0067] Step 101, receiving an input image through a Stem network based on a convolutional neural network, and extracting initial posture features to obtain preliminary features.
[0068] In the embodiment of the present application, the main function of the Stem network is to filter out irrelevant information (such as background, color, etc.) from the input image and extract initial posture features as input for the subsequent encoder.
[0069] In order to achieve the above functions, refer to Figure 2 The Stem network proposed in the embodiment of the present application is specially designed with an efficient multi-layer structure to extract initial posture-related features from the input image while filtering out a large amount of redundant information (such as irrelevant features such as background and color). The network mainly consists of two modules:
[0070] The first module uses a normal convolution layer with a convolution kernel size of 7, combined with a batch normalization layer (BatchNormalization) and a ReLU activation function layer. Its main task is to capture low-level features of the image, such as edges, contours, and other basic texture information. By using a larger convolution kernel, a larger receptive field can be covered, thereby capturing more global information.
[0071] The second module consists of two deep convolutional layers with a kernel size of 3, which are also combined with a batch normalization layer and a ReLU activation function layer. The deep convolutional layer focuses on further processing the low-level features extracted in the first stage, removing redundant information in the image, such as background and color interference, through local convolution operations. At the same time, this module enhances the extraction of posture-related features to ensure that the generated initial features not only contain rich edge and contour information, but also highlight the significant features related to the key points of the human body or animal.
[0072] In specific implementation, the calculation process of the Stem network is:
[0073]
[0074] Among them, X Iis the input image, H and W are the height and width of the input image respectively, d s is the feature dimension, Conv(·) represents the processing of the convolutional layer, DWCnv(·) represents the processing of the deep convolutional layer, BN(·) represents the processing of the batch normalization layer, and ReLU(·) represents the processing of the ReLU activation function.
[0075] The entire Stem network uses this staged processing method to significantly reduce the computational cost while ensuring the quality of extracted features, providing efficient and accurate input features for the subsequent Mamba encoder. This structure can effectively reduce the interference of background and irrelevant factors on key point detection, while ensuring the retention of low-level features and the expression of highly relevant features, laying a solid foundation for subsequent key point detection.
[0076] Step 102, using a Mamba-based encoder to process the initial features and output multi-level posture features, wherein the encoder includes multiple stages, each of which consists of a context modeling module, a two-dimensional selective scanning module and a normalization layer.
[0077] In the embodiment of the present application, the encoder includes multiple stages, each of which is composed of a context modeling module, an N i The context modeling module captures the dependencies between image patches, extracts multi-scale features and optimizes the expression of local information; the two-dimensional selective scanning module further models the interactions between all image patches and integrates global information; the normalization layer is used to adjust the distribution of output features to improve model stability and training effect.
[0078] In this step, the encoder uses the main image features As input, it outputs the extracted multi-level posture features as the input of the subsequent simple decoder.
[0079] In the embodiment of the present application, each stage of the encoder selects a different feature downsampling rate, and the processing process of each stage is:
[0080]
[0081] in, is the input of the i-th stage. If i=1, then is the output of the i-th stage, which contains the feature enhancement of context information and patch interaction; LN(·) represents the processing of the normalization layer, which is responsible for adjusting the distribution of features; SS2D(·) represents the processing of the two-dimensional selective scanning module, which captures global dependencies through multiple scans. There are N in the i-th stage. iA two-dimensional selective scanning module; CMM(·) represents the processing process of the context modeling module.
[0082] It can be understood that each stage of the encoder selects different feature downsampling rates according to the task requirements to capture multi-scale posture features. In one possible embodiment, the encoder uses three stages, corresponding to three feature downsampling rates {4, 8, 16}, for extracting detail features, extracting medium-scale features, and capturing global features. This design can effectively extract features at different scales, thereby effectively covering local and global posture features, and ensuring the detection capability of the model in complex scenes.
[0083] The structure and function of the context modeling module and the two-dimensional selective scanning module are described in detail below.
[0084] It should be noted that the traditional Mamba module recursively extracts features from image patches during the modeling process while maintaining a constant feature dimension, which leads to two inherent limitations of Mamba. First, it is difficult for Mamba to establish long-term dependencies between non-adjacent and distant patches. Second, its ability to extract multi-scale features is limited. To this end, an embodiment of the present application proposes a context modeling module (CMM) to assist Mamba in extracting multi-scale image features and enhance information interaction between different image patches.
[0085] In order to balance the size and performance of the model, in this application, the context modeling module is designed as a simple but functionally effective module, which mainly includes the following three parts: deep convolutional layer, patch embedding layer and linear layer.
[0086] First, a deep convolutional layer is used to perform convolution operations along the dimensions of the image patch to capture the dependencies between image patches within the receptive field. This process can not only effectively integrate local context information, but also downsample the input features to expand the receptive field and provide richer context information for subsequent modules.
[0087] Next, the patch embedding layer further processes the features extracted by the deep convolutional layer. The main function of the patch embedding layer is to convert the input feature representation into a more compact and meaningful feature representation through patch partitioning and patch encoding operations. The partitioning process divides the input features into multiple image patch units, while the encoding process further extracts and integrates these patches through additional convolution operations to capture higher-level feature expressions.
[0088] Finally, the context features processed by the patch embedding layer are mapped through a linear layer. The main function of the linear layer is to map the context features to a new representation space that matches the input feature space of the Mamba module, thereby providing a consistent format and high-quality input feature representation for the calculation of subsequent modules.
[0089] Through the above three steps, the context modeling module can effectively model the contextual relationship between patches and lay a solid feature foundation for the operation of the Mamba module. Therefore, the feature processing process of the context modeling module can be formally expressed as:
[0090]
[0091] in, is the input feature; DWConv(·) represents the deep convolution, which is used to capture the dependencies between patches; Embed(·) represents the patch embedding layer, which is used to further extract the features between patches; LN(·) represents the normalization layer, which is used to adjust the feature distribution; Represents the context features learned by the context modeling module.
[0092] The final output features are obtained through the context modeling module After extraction, the embodiment of the present application further refines the features through a two-dimensional selective scanning module. Specifically, the two-dimensional selective scanning module rearranges the features extracted by the context modeling module through a sequential scanning strategy in four directions, further integrates the information relationship between patches, and calculates attention weights through context features to activate image areas related to key tasks. The core goal of this processing process is to enhance the global perception ability and local detail interaction ability of features, and to ensure that on the basis of the global context features extracted by the context modeling module, the distinguishability and adaptability of the features are further enhanced, thereby providing more efficient and accurate input features for subsequent key point detection and feature decoding modules.
[0093] In the embodiment of the present application, the two-dimensional selective scanning module first rearranges the input image features. Specifically, a sequential scanning strategy in four directions (such as from left to right, from top to bottom, diagonal, etc.) is used to sequentially arrange the input image patches to ensure that the dependencies between adjacent patches are captured through the rearrangement process. This process can not only strengthen local feature interactions, but also provide a basis for global feature fusion.
[0094] After the rearrangement is completed, three linear layers (LN) are used to further map the rearranged features. The main function of the linear layer is to transform the original context features into a unified feature space to ensure that the features generated after scanning in different directions are consistent. This feature mapping provides a standardized input for subsequent discretization processing and state update.
[0095] Next, the mapped features are discretized through the discretization function to generate a discretized state matrix and a discretized parameter matrix. The purpose of discretization is to convert the continuous feature space into discrete features that are more suitable for state updating, so that subsequent modules can perform efficient calculations.
[0096] Next, based on the generated discretized state matrix and parameter matrix, combined with the state information of the previous moment, the state update mechanism is used to calculate the state of the current moment. Specifically, the state of the current moment is the linear transformation result of the discretized state matrix and the state of the previous moment. This process captures the contextual dependency in the time series features through a clear state recursive relationship and gradually strengthens the information between features.
[0097] Finally, after completing the state update, the updated state vector p k The input features are fused through a linear layer. The core of the fusion process is to integrate the dynamically updated state information with the original input features to generate the final output features. The fused features not only contain the rearranged context information, but also combine the adaptively calculated states, which significantly improves the model's ability to identify key areas.
[0098] Therefore, the feature processing process of the two-dimensional selective scanning module can be formally expressed as:
[0099]
[0100] Among them, LN(·) represents the processing of the linear layer, Δ is the adaptive matrix, and is the parameter matrix of the state space model, Discretize(·) represents the discretization process, p k Indicates the current state output, p k-1 Indicates the state output at the previous moment, Represents the output features of the two-dimensional selective scanning module.
[0101] Through the above processing, the 2D selective scanning module can effectively integrate the information interaction between patches and activate key areas through context feature calculation. This method not only enhances the correlation between features, but also further improves the accuracy of feature representation through state update and feature fusion, laying the foundation for efficient reasoning of subsequent model modules.
[0102] Step 103: Use a decoder to upsample the multi-level posture features output by the encoder into a key point heat map, where the heat map of each key point represents the key point position of the target instance.
[0103] The main function of the decoder is to convert the multi-level posture features output by the encoder into a key point heat map. The heat map expresses the position distribution of each key point, retains the structure of the posture information, and provides support for the prediction of subsequent key point coordinates.
[0104] In the embodiment of the present application, the decoder gradually upsamples the features through two deconvolution layers with a kernel size of 4, and then maps the upsampled features into a key point heat map. The decoding process can be expressed as:
[0105]
[0106] Among them, H is the key point heat map, is the multi-level posture feature output by the encoder, K is the number of key points, and Decoder(·) represents the decoding process of the decoder.
[0107] The deconvolution layer can gradually restore the spatial resolution of the feature map, making the output heat map closer to the scale of the original input image; through the linear mapping of the decoder, the multi-level posture features of the encoder are mapped to the probability distribution of each key point.
[0108] This heatmap-based representation preserves the spatial relationship between key points and avoids the errors that may be introduced by directly regressing coordinates. Through the above design, the decoder achieves an efficient conversion from pose features to key point heatmaps, which can provide high-quality key point predictions while maintaining the simplicity of the model.
[0109] In addition, in order to simplify the training process of the detection model of the present application and ensure the performance of the model in the two-dimensional key point detection task. The embodiment of the present application adopts the commonly used mean square error loss function for training. The mean square error loss function measures the prediction accuracy of the model by calculating the square difference between the predicted value and the true value. Its loss function is expressed as:
[0110]
[0111] Among them, M is the total number of training samples, H i and Represent the predicted key point heat map and true value of the i-th sample respectively.
[0112] The results show that by using the mean square error loss function, PoseMamba2D can achieve efficient training without complex supervision and achieve remarkable performance in the two-dimensional key point detection task, demonstrating the superiority and simplicity of the method.
[0113] In order to verify the effectiveness of PoseMamba2D, the present application embodiment also conducted a large number of experiments on human and animal pose estimation datasets. The main goal of the experiment is to evaluate the performance (accuracy) and reasoning efficiency (speed) of PoseMamba2D in key point detection, and compare it with the current mainstream methods to highlight the advantages of PoseMamba2D.
[0114] The results show that:
[0115] (1) On the COCO dataset, PoseMamba2D achieves 77.3% AP, demonstrating its pose estimation capability in complex scenes.
[0116] (2) On an NVIDIA GTX4090 GPU, PoseMamba2D achieves an inference speed of 1492 FPS, far exceeding mainstream models, fully demonstrating its application potential in real-time tasks.
[0117] (3) On the MPII dataset, PoseMamba2D achieved better results than ViTPose, demonstrating its advancedness in the field of human pose estimation.
[0118] (4) On the AP-10K animal dataset, PoseMamba2D achieved competitive results while saving 85% of parameters compared to ViTPose, further demonstrating the efficiency of its lightweight design.
[0119] Through the above experiments and analysis, PoseMamba2D demonstrates excellent performance in two-dimensional key point detection tasks, taking into account accuracy, speed and resource optimization, and provides a powerful solution for practical applications.
[0120] In order to implement the above embodiment, the present application also proposes a real-time 2D key point detection device based on Mamba. Figure 3 A schematic diagram of the structure of a real-time 2D key point detection device based on Mamba provided in an embodiment of the present application. Figure 3 As shown, the device comprises:
[0121] The feature extraction module 100 is used to receive the input image through the Stem network based on the convolutional neural network, and extract the initial posture features to obtain preliminary features;
[0122] An encoding module 200, used to process the initial features using the Mamba encoder and output multi-level posture features, wherein the encoder includes multiple stages, each of which is composed of a context modeling module, a two-dimensional selective scanning module and a normalization layer;
[0123] The decoding module 300 is used to use a decoder to upsample the multi-level posture features output by the encoder into a key point heat map, and the heat map of each key point represents the key point position of the target instance.
[0124] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0125] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0126] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0127] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0128] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.
[0129] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.
[0130] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0131] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0132] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0134] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0135] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0136] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0137] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
[0138] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.
[0139] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A real-time 2D key point detection method based on Mamba, characterized in that: The following steps are involved: The input image is received through the Stem network based on the convolutional neural network, and the initial posture features are extracted to obtain preliminary features; Processing the initial features using a Mamba-based encoder to output multi-level posture features, wherein the encoder includes multiple stages, each stage consisting of a context modeling module, a two-dimensional selective scanning module, and a normalization layer; The multi-level posture features output by the encoder are upsampled into key point heat maps using a decoder, and the heat map of each key point represents the key point position of the target instance.
2. The method according to claim 1, characterized in that The Stem network is used to: A convolution layer with a convolution kernel size of 7, a batch normalization layer, and a ReLu activation function layer are used to initially capture the low-level features in the input image. The expression is: Two deep convolution layers with a convolution kernel size of 3, a batch normalization layer, and a ReLU activation function layer are used to remove redundant information in the original input image and further extract posture-related features. The initial features extracted by the Stem network are It is expressed as: Among them, X I is the input image, H and W are the height and width of the input image respectively, d s is the feature dimension, Conv(·) represents the processing of the convolutional layer, DWConv(·) represents the processing of the deep convolutional layer, BN(·) represents the processing of the batch normalization layer, and ReLU(·) represents the processing of the ReLU activation function.
3. The method according to claim 2, characterized in that Each stage of the encoder selects a different feature downsampling rate, and the processing process of each stage is: in, is the input of the i-th stage. If i=1, then is the output of the i-th stage; LN(·) represents the processing of the normalization layer; SS2D(·) represents the processing of the two-dimensional selective scanning module. The i-th stage has N i A two-dimensional selective scanning module; CMM(·) represents the processing process of the context modeling module.
4. The method according to claim 3, characterized in that The context modeling module contains a deep convolutional layer, a patch embedding layer and a linear layer, where: The deep convolutional layer is used to capture the dependencies between image patches within the receptive field and downsample the input features; The patch embedding layer is used to perform patch division and encoding on the features extracted by the deep convolutional layer; The linear layer is used to map the context features obtained by the patch embedding layer into the representation space of the matching Mamba patch input; The calculation process of the context modeling module is: in, i∈{1,2,3} represents the context features learned by the context modeling module; Embed(·) represents the processing process of the patch embedding layer, LN(·) represents the processing process of the linear layer, and DWConv(·) represents the processing process of the deep convolutional layer.
5. The method according to claim 4, characterized in that The two-dimensional selective scanning module is used for: The image patches are rearranged through sequential scanning strategies in four directions, the information between the image patches is integrated, and the attention weights are calculated through context features to activate the key image areas. Each scanning strategy obtains the adaptive matrix and parameter matrix of the state space model through a linear layer, and obtains the discretized parameter matrix through a discretization process. The state at the current moment is calculated using the discretized parameter matrix and the state at the previous moment; The calculation process of the two-dimensional selective scanning module is: Among them, LN(·) represents the processing of the linear layer, Δ is the adaptive matrix, and is the state matrix, Discretize(·) represents the discretization process, p k Indicates the current state output, p k-1 Indicates the state output at the previous moment, represents the output characteristics of the two-dimensional selective scanning module, Features extracted by the context modeling module.
6. The method according to claim 5, characterized in that The decoder includes two deconvolution layers with a kernel size of 4. The calculation formula of the decoder is: Wherein, H is the key point heat map, is the multi-level posture feature output by the encoder, K is the number of key points, and Decoder(·) represents the decoding process of the decoder.
7. The method according to claim 6, characterized in that Also includes: The detection method is optimized using the mean square error loss function, and its loss function is expressed as: Among them, M is the total number of training samples, H i and Represent the predicted key point heat map and true value of the i-th sample respectively.
8. A real-time 2D key point detection device based on Mamba, characterized in that: include: A feature extraction module is used to receive an input image through a Stem network based on a convolutional neural network, and extract initial posture features to obtain preliminary features; An encoding module, configured to process the initial features using a Mamba encoder and output multi-level posture features, wherein the encoder includes multiple stages, each of which is composed of a context modeling module, a two-dimensional selective scanning module, and a normalization layer; A decoding module is used to use a decoder to upsample the multi-level posture features output by the encoder into a key point heat map, and the heat map of each key point represents the key point position of the target instance.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Two-dimensional human body posture estimation method and system based on lightweight multi-branch network
CN110969124A