Lightweight attitude recognition model construction method and device and attitude recognition method
By building a lightweight attitude recognition model, optimizing the network structure to achieve a balance of accuracy and speed on the drone platform, the problem of difficulty in deploying neural networks on miniaturized devices in the prior art is solved, and is suitable for violation detection of intelligent security systems.
Patent Information
- Application Number
- CN202510328264.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The existing human posture recognition neural network cannot be effectively deployed on miniaturized devices such as drones, making it difficult to balance accuracy and operation speed.
Build a lightweight pose recognition model, define the pose recognition model, deep search space and fusion selection space, use evolutionary algorithms to search, optimize the network structure to reduce redundancy, and use segmented evolutionary search method to improve search efficiency.
It has achieved both accuracy and operating speed on miniaturized platforms such as drones, and is suitable for violation detection in intelligent security systems.
Smart Images

Figure CN120260124A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision recognition, and specifically relates to a method and device for constructing a lightweight pose recognition model and a pose recognition method. Background Art
[0002] Human pose recognition, as an upstream task of behavior recognition, aims to locate the coordinates of key parts such as the head, shoulders, and hips of ground targets in an image, so as to construct a skeleton sequence to provide input data sources for behavior recognition. In recent years, intelligent security has been increasingly applied in public places such as railway stations and airports to protect the lives and property of the people. Pose recognition, as the most important part of the intelligent security system, can be used for violation detection in various dynamic scenarios, such as detecting potential threat behaviors in intelligent security for key areas.
[0003] Existing neural networks for human pose recognition cannot be effectively deployed on small devices such as drones. Therefore, it is necessary to lightweight the neural network for human pose recognition and achieve a balance between accuracy and running speed on the drone platform.
[0004] Currently, for neural networks used for lightweight human pose recognition, such as the patent "A Method and System for Multi-Person Human Pose Estimation Based on Knowledge Distillation" with the patent number "CN 115187660A", it discloses the "GhostPoseNet" network, which uses knowledge distillation technology to generate joint offsets using the student heatmap and the target joint heatmap of the data label, and dynamically adjusts the knowledge transfer from the teacher network to the student network to construct a lightweight network. Although this method significantly reduces the number of network parameters, the network accuracy also drops severely. In the literature with the literature number "10.48550 / arXiv.2104.06403", it discloses the "Lite-HRNet" network, which uses the channel shuffle module Shuffle Block to replace the 3×3 convolution with a high computational cost in the network. Although it significantly reduces the complexity of the model, the loss in model accuracy is relatively serious. In the literature with the literature number "10.48550 / arXiv.2204.10762", it discloses the "Dite-HRNet" network, which uses dynamic segmentation convolution and adaptive context modeling to design two lightweight modules, namely dynamic multi-scale context and dynamic global context. Compared with previous methods, the accuracy has been improved, but it is still not ideal.
[0005] In summary, for how to deploy a neural network for human pose recognition with both accuracy and running speed on small platforms such as drones, it is still an urgent technical problem to be solved at present. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a method for constructing a lightweight pose recognition model, an apparatus, and a pose recognition method.
[0007] The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0008] In a first aspect, the present invention provides a method for constructing a lightweight pose recognition model, including:
[0009] Define a pose recognition model; the pose recognition model includes: an input layer, a conversion layer, a deep stacking module, and an output layer; the input layer is used to send an image to be recognized for pose into the conversion layer; the conversion layer includes a plurality of feature extraction branches, and the plurality of feature extraction branches are used to extract feature maps of different scales from the image; the deep stacking module includes a plurality of cascaded deep search modules; each deep search module includes a search layer and a fusion layer; the search layer includes a plurality of feature extraction units corresponding one by one to the plurality of feature extraction branches, and is used to perform feature extraction on the plurality of feature maps input to the search layer, and the fusion layer is used to perform feature fusion on the feature maps extracted by the plurality of feature extraction units and output to the next-level deep search module; the output layer is used to output a pose recognition result according to the feature map output by the last-level fusion layer;
[0010] Define a deep search space; the deep search space defines the cascaded number space of the deep stacking module in the pose recognition model;
[0011] Define a fusion selection space according to the deep search space; the fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select fusion objects;
[0012] Under the deep search space and the fusion selection space, search for a lightweight pose recognition model for pose recognition.
[0013] Optionally, the searching for a lightweight pose recognition model for pose recognition under the deep search space and the fusion selection space includes:
[0014] Under the deep search space and the fusion selection space, use an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0015] Optionally, the using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition under the deep search space and the fusion selection space includes:
[0016] Under the depth search space and the fusion selection space, by giving the pose recognition model a fixed hybrid operator, an evolutionary algorithm is used to perform channel structure optimization search to obtain a number of channel structure optimized subnets; wherein, the channel structure optimization search is used to determine the cascade number of the depth stacking module and the selection strategy of each fusion layer for the fusion object through search.
[0017] Define a hybrid operator search space. Under the hybrid operator search space, using the number of channel structure optimized subnets as the search set, an evolutionary algorithm is used to perform operator optimization search to obtain a lightweight pose recognition model for pose recognition; wherein, the operator optimization search is used to determine the operator type adopted by the lightweight pose recognition model through search.
[0018] Optionally, before using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, the method further includes: under the depth search space and the fusion selection space, performing supernet weight pre-training through a single-path uniform sampling strategy, so as to use the pre-trained supernet weights to search for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm.
[0019] Optionally, after completing the supernet weight pre-training and before using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, the method further includes:
[0020] Under the depth search space and the fusion selection space, perform subnet random sampling search according to the pre-trained weights to obtain a number of subnets, which are used as the parent subnets when subsequently using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0021] In a second aspect, the present invention provides a device for constructing a lightweight pose recognition model, the device includes:
[0022] A first definition module, used to define a pose recognition model; the pose recognition model includes: an input layer, a conversion layer, a depth stacking module, and an output layer; the input layer is used to send an image to be recognized for pose into the conversion layer; the conversion layer includes a plurality of feature extraction branches, and the plurality of feature extraction branches are used to extract feature maps of different scales from the image; the depth stacking module includes a plurality of cascaded depth search modules; each depth search module includes a search layer and a fusion layer; the search layer includes a plurality of feature extraction units corresponding one-to-one to the plurality of feature extraction branches, and is used to perform feature extraction on the plurality of feature maps input to the search layer, and the fusion layer is used to perform feature fusion on the feature maps extracted by the plurality of feature extraction units and output to the next-level depth search module; the output layer is used to predict a pose recognition result according to the feature map output by the last fusion layer.
[0023] A second definition module for defining a depth search space; the depth search space defines the cascade number space of the depth stacking modules in the pose recognition model;
[0024] A third definition module for defining a fusion selection space according to the depth search space; the fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select fusion objects;
[0025] A network search module for searching for a lightweight pose recognition model for pose recognition under the depth search space and the fusion selection space.
[0026] Optionally, the network search module is specifically configured to:
[0027] Under the depth search space and the fusion selection space, use an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0028] Optionally, the network search module includes: a first search unit and a second search unit;
[0029] The first search unit is configured to, under the depth search space and the fusion selection space, perform channel structure optimization search by giving the pose recognition model a fixed mixing operator and using an evolutionary algorithm to obtain a number of channel structure optimized subnets; wherein, the channel structure optimization search is used to determine the cascade number of the depth stacking modules and the selection strategy of each fusion layer for the fusion objects through search;
[0030] The second search unit is configured to define a mixing operator search space, and under the mixing operator search space, use the number of channel structure optimized subnets as a search set and use an evolutionary algorithm to perform operator optimization search to obtain a lightweight pose recognition model for pose recognition; wherein, the operator optimization search is used to determine the operator type adopted by the lightweight pose recognition model through search.
[0031] Optionally, the device further includes: a super network weight pre-training module;
[0032] The super network weight pre-training module is configured to, before using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, perform super network weight pre-training through a single-path uniform sampling strategy under the depth search space and the fusion selection space, so that the search module uses an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition according to the pre-trained super network weights.
[0033] In a third aspect, the present invention provides a pose recognition method, including:
[0034] Obtain an image for which pose recognition is to be performed;
[0035] Input the image into a pre-trained lightweight pose recognition model for pose recognition, so that the lightweight pose recognition model outputs a pose recognition result for the image;
[0036] Wherein, the lightweight pose recognition model is constructed by using the construction method of any one of the above-mentioned lightweight pose recognition models.
[0037] The construction method of the lightweight pose recognition model provided by the present invention defines a pose recognition model as the backbone network, so that based on this backbone network, a deep search space that can effectively reduce redundancy can be defined, and the computational complexity of the algorithm can be reduced to a certain extent. On the basis of the deep search space, a fusion selection space is defined, and a search space that can effectively suppress the large expansion of the network parameter quantity and the floating-point number of operations caused by the large stacking of depths, and avoid the network learning redundant information to reduce the network accuracy can be constructed. Thus, under the deep search space and the fusion selection space, the lightweight pose recognition model searched for pose recognition can achieve the performance of both accuracy and running speed, and thus can be deployed on small platforms such as unmanned aerial vehicles.
[0038] The present invention further adopts a segmented evolutionary search method. By giving the pose recognition model a fixed mixing operator, first use the evolutionary algorithm to perform channel structure optimization search, and then under the mixing operator search space, use the channel structure optimized subnet obtained by the channel structure optimization search as the search set, and further use the evolutionary algorithm to perform operator optimization search to obtain a lightweight pose recognition model for pose recognition. This segmented evolutionary search method makes the network randomly mutate and the selected subnets will synchronously obtain higher accuracy in the subsequent search process, making it easier for the evolutionary algorithm to find the optimal solution in the huge search space, improving the algorithm efficiency, and at the same time ensuring that the searched network has sufficient accuracy.
[0039] The following will further elaborate on the present invention in conjunction with the accompanying drawings. Description of the Drawings
[0040] Figure 1 is a flowchart of a construction method of a lightweight pose recognition model provided by the present invention;
[0041] Figure 2 is a schematic diagram of the pose recognition model in the present invention;
[0042] Figure 3 is a schematic diagram of the deep search module in the present invention;
[0043] Figure 4An exemplary search process for a search lightweight pose recognition model used in the present invention is shown;
[0044] Figure 5 Exemplarily shown is the training accuracy (Train Acc) of the present invention as the number of search iterations increases. Detailed Description of the Invention
[0045] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0046] In order to be able to deploy a neural network for human pose recognition with both accuracy and running speed on a small platform such as a drone, an embodiment of the present invention provides a method for constructing a lightweight pose recognition model, as Figure 1 shown, the method includes the following steps:
[0047] S10. Define a pose recognition model; the pose recognition model includes: an input layer, a conversion layer, a deep stacking module, and an output layer; the deep stacking module includes a plurality of cascaded deep search modules; each deep search module includes a search layer and a fusion layer.
[0048] Referring to Figure 2 , the input layer is used to send an image to be recognized for pose into the conversion layer. The conversion layer includes a plurality of feature extraction branches (the number is represented by n), and these feature extraction branches are used to extract feature maps of different scales from the input image.
[0049] It can be understood that the actual function of the input layer is to convert the image to be recognized for pose into sizes respectively matching the input ends of the respective feature extraction branches of the conversion layer, so that these feature extraction branches can extract feature maps from the image. Each feature extraction branch of the conversion layer may include a convolutional neural network and a 1×1 convolutional compression channel; among them, the feature maps extracted by the convolutional neural networks of different feature extraction branches have different sizes, and these feature maps are input into the deep search module through a 1×1 convolutional compression channel.
[0050] Continuing to refer to Figure 2 , each search layer in each deep search module includes a plurality of feature extraction units corresponding one-to-one to the plurality of feature extraction branches of the conversion layer, and these feature extraction units are used to further extract features from the feature maps input to this search layer, so as to send the further extracted feature maps into the fusion layer of the same deep search module to achieve feature fusion.
[0051] In the present invention, for what type of operator is used to implement the feature extraction unit of the search layer, it can be maliciously determined through subsequent neural network search steps, and specific reference is made to the subsequent steps.
[0052] The fusion layer is used to perform feature fusion on the feature maps extracted by multiple feature extraction units of the search layer connected thereto and output the result to the next-level deep search module.
[0053] See Figure 3 , in the deep search module, its search layer ( Figure 3 represented by Transform Layer in Figure 3 ) includes n feature extraction units, and its fusion layer ( Figure 3 represented by Fuse Layer in
[0054] ) includes m fusion blocks for implementing specific fusion operations, where m ≤ n.
[0055] Specifically, for each fusion layer, all the feature maps input to this fusion layer can be pairwise fused, or some of them can be selected for pairwise fusion. If the feature map output by a certain feature extraction unit needs to be fused with the feature map output by another feature extraction unit, then these two feature units need to be correspondingly connected to the same fusion block, and the role of this fusion block is to fuse the feature maps extracted by these two feature extraction units and output the result backward. If the feature map of a certain feature extraction unit does not need to be fused with other feature maps, then this feature extraction unit is not connected to a fusion block, but is directly connected to the search layer of the next deep search module. In this way, the feature maps output from the fusion layer still include n, some of which are the fused feature maps output by the fusion blocks in this fusion layer, and some are the feature maps directly output by a certain feature extraction unit in the search layer connected to this fusion layer without being fused by this fusion layer.
[0056] Continue to see Figure 2 , the last stage of the pose recognition model is the output layer, which is used to output the pose recognition result according to the feature map output by the last-stage fusion layer. This output layer includes multiple task branches, and the feature maps output by the last-stage deep search module are sent into these task branches for final classification output to achieve pose recognition.
[0057] In summary, the pose recognition model defined in the present invention, which focuses on the task of human pose recognition, can separate feature maps of multiple resolutions and interact information at different stages of the network, so as to better capture features of different scales and improve the accuracy of human pose recognition.
[0058] It is understandable that with the cascaded stacking of the depth search module, the number of model parameters will surge. At the same time, too many fusion layers will not only not increase the accuracy of the network, but also cause the network to learn redundant information and exhibit overfitting. Therefore, the present invention performs an optimal search by using a neural network search method, so as to select an appropriate cascaded number on the premise of ensuring accuracy. For details, refer to the following steps.
[0059] S20. Define a depth search space; this depth search space defines the cascaded number space of the depth stacking module in the pose recognition model.
[0060] Here, the depth search space is actually a set of pose recognition models. The pose recognition models in this set are the same except for the cascaded number of the depth stacking module. In addition, this set covers all possible cascaded numbers of the depth stacking module. Note that all possible cascaded numbers here are not infinite, but are covered as much as possible within a certain reasonable range.
[0061] The defined depth search space of the present invention can support searching for the optimal cascaded number of the depth stacking module, that is, support searching for the optimal cascaded depth and obtaining the optimal depth stacking module.
[0062] S30. Define a fusion selection space according to the depth search space; this fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select the fusion object.
[0063] Here, the fusion selection space is actually also a set of pose recognition models. The cascaded number of the pose recognition models in this set is within the constraints of the depth search space, and this set covers all possible selection strategies for each fusion layer to select the fusion object. The selection strategy mentioned here is also the connection relationship between the feature extraction unit and the fusion block mentioned in step S10, as well as the specific selection of whether each depth search module participates in feature fusion. Similarly, it can be understood that all possible selection strategies mentioned here do not mean an exhaustive list of selection strategies, but are covered as much as possible within a certain reasonable range.
[0064] Thus, the fusion selection space defined in step S30 can support searching for the optimal fusion method and fusion channels.
[0065] S40. Search for a lightweight pose recognition model for pose recognition under the depth search space and the fusion selection space.
[0066] Specifically, under the depth search space and the fusion selection space, an evolutionary algorithm is used to search for a lightweight pose recognition model for pose recognition.
[0067] Among them, when using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, all pose recognition models in the deep search space and the fusion selection space are used as the population, and each pose recognition model among them is an individual in the population. The basic idea of evolutionary computation is to optimize the adaptability of the population through operations such as genetic variation, mating, and selection of individuals, so as to continuously evolve a neural network architecture that is more suitable for the problem. Each time the population iterates, multiple subnets are extracted from the defined search space and trained through forward propagation, and the accuracy of the trained subnets is calculated and evaluated. Then, several subnets with higher accuracy are selected, and based on the idea of One-Shot NAS, parents are generated, and offspring are generated by mutating the parents, so as to realize the optimization and update of the population, and finally search for a lightweight pose recognition model with both accuracy and running speed.
[0068] It can be understood that, based on the population-based gradient-free evolutionary algorithm, different from the traditional gradient descent method, the evolutionary algorithm does not require assumptions such as continuity, the existence of derivatives, and unimodality, and can find the global optimal solution with a high probability from discrete, multi-extremum, and noisy high-dimensional problems. In the present invention, through the evolutionary algorithm, the optimal search of the neural network is continuously carried out. In the search process, the network randomly mutates, and the selected subnets will synchronously obtain higher accuracy. Based on this characteristic of the evolutionary algorithm, it shows a trend of overall optimization in the subnet search process. This overall optimization trend can make the evolutionary algorithm easier to find the optimal solution in the huge search space.
[0069] The method for constructing the lightweight pose recognition model provided by the present invention defines the pose recognition model as the backbone network, so that based on this backbone network, a deep search space that can effectively reduce redundancy can be defined, and to a certain extent, the computational amount of the algorithm can be reduced. On the basis of the deep search space, a fusion selection space is defined, and a search space can be constructed that can effectively suppress the large expansion of the network parameter quantity and the floating-point number of operations caused by the large stacking of depths, and avoid the network learning redundant information and reducing the network accuracy. Thus, under the deep search space and the fusion selection space, the searched lightweight pose recognition model for pose recognition can achieve the performance of both accuracy and running speed, and thus can be deployed on small platforms such as unmanned aerial vehicles.
[0070] In addition, considering that the defined search space is large, in order to further improve the search efficiency, the present invention designs a segmented evolutionary search method. Thus, under the deep search space and the fusion selection space, using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition includes:
[0071] (1) Under the deep search space and the fusion selection space, by giving the pose recognition model a fixed mixing operator, an evolutionary algorithm is used to perform an optimal search for the channel structure to obtain several subnets with optimal channel structures; among them, the optimal search for the channel structure is used to determine the cascade number of the deep stacking module and the selection strategy of each fusion layer for the fusion object through search.
[0072] Specifically, when using the evolutionary algorithm to perform an optimal search for the channel structure, by giving the pose recognition model a fixed mixing operator, the mutation positions during the search are locked in the deep search space and the fusion selection space. Since the channel structure has a great influence on the network, the mutation rate from the parent generation to the offspring during the search in step (1) should be set to be small. At the same time, the optimization accuracy is replaced from the average accuracy to a higher required accuracy to ensure that the excellent results after the subnet mutation are due to the change in the channel structure. Since the optimal subnets are clustered in the search space, in this step, the mutated subnets are not further trained, and only the one with the best overall performance among multiple candidate spaces is determined.
[0073] (2) Define the mixing operator search space. Under the mixing operator search space, using the above-mentioned several subnets with optimal channel structures as the search set, an evolutionary algorithm is used to perform an optimal search for the operator to obtain a lightweight pose recognition model for pose recognition; among them, the optimal search for the operator is used to determine the operator type adopted by the lightweight pose recognition model through search.
[0074] It can be understood that after determining the optimal channel structure subnet, at this time, the offspring subnet will only exist in a small area near the entire search space. Therefore, when performing the optimal search for the operator in step (2), even if the mutation rate is adjusted to a larger value, it can ensure that the network only fluctuates within a small range.
[0075] In this step (2), the accuracies of the subnets in the search set are similar. To ensure that each subnet can have a more intuitive performance, the offspring selected in each round are subjected to a small number of rounds of backpropagation to observe the convergence speed of the network. For example, the convergence speed of the network can be observed based on the object keypoint similarity. Finally, the average accuracy is used to find the offspring with the best performance among them.
[0076] Among them, the calculation formula of the object keypoint similarity (OKS) is as follows:
[0077]
[0078] Among them, OKS represents the object keypoint similarity, i is the number of keypoints, d i represents the Euclidean distance between the detected keypoint and the annotated keypoint, v iVisibility flag bits representing the key points of the true annotation, where v i = 0 indicates not annotated, v i = 1 indicates annotated but occluded, v i = 2 indicates annotated and occluded, δ(·) represents the impulse function, s is the scale of the object, k i is a constant for controlling the attenuation of each key point.
[0079] The present invention adopts a segmented evolutionary search method. By giving a fixed hybrid operator to the pose recognition model, first, an evolutionary algorithm is used to search for the optimal channel structure, and then, under the search space of the hybrid operator, with the channel structure optimal subnet obtained by the channel structure optimal search as the search set, an evolutionary algorithm is further used to search for the optimal operator, so as to obtain a lightweight pose recognition model for pose recognition. This segmented evolutionary search method enables the network to randomly mutate and the selected subnets to synchronously obtain higher accuracy in the subsequent search process, making it easier for the evolutionary algorithm to find the optimal solution in the huge search space, improving the algorithm efficiency, and at the same time ensuring that the searched network has sufficient accuracy.
[0080] In one embodiment, before using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, the method of the present invention further includes: under the deep search space and the fusion selection space, performing hypernetwork weight pre-training through a single-path uniform sampling strategy, so as to use the pre-trained hypernetwork weights to search for a lightweight pose recognition model for pose recognition by the evolutionary algorithm.
[0081] It can be understood that the hypernetwork is a shared weight model based on the deep search space and the fusion selection space, which contains a large number of possible sub-networks. By pre-training the weights of the hypernetwork, it can provide a good initialization for the subsequent architecture search, thereby accelerating the architecture search process and reducing the search time and computational cost.
[0082] In one embodiment, after completing the hypernetwork weight pre-training and before using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, the method of the present invention further includes: under the deep search space and the fusion selection space, performing subnet random sampling search according to the pre-trained hypernetwork weights to obtain a number of sub-networks as the parent sub-networks when subsequently using the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0083] Specifically, after completing the hypernetwork weight pre-training, subnet random sampling will be performed for a period of time. Since the efficient subnets are mostly distributed adjacently in the search space, showing a heatmap-like shape, when using a number of sub-networks obtained by subnet random sampling search as the optimal parents of the evolutionary algorithm for evolutionary search, the situation of excluding potential optimal subnets from the network structure searched by the evolutionary algorithm will not occur.
[0084] It is understandable that, under the deep search space and the fusion selection space, random sampling search of subnets according to the pre-trained hypernetwork weights can avoid exhaustive search of all possible subnets under the deep search space and the fusion selection space, thereby reducing the computational cost and time overhead.
[0085] In a specific example, refer to Figure 4 , hypernetwork weight pre-training can be performed first, then random sampling search of subnets can be carried out according to the pre-trained hypernetwork weights, and then a segmented evolutionary algorithm can be used for neural network search. Figure 5 shows the training accuracy (Train Acc) of the present invention as the number of search iterations increases. Then, the finally searched subnet is retrained, and the trained subnet is used as the pose recognition model constructed by the present invention to compare the performance with multiple existing models to verify the effectiveness of the present invention. The specific experiments are described as follows:
[0086] Table 1 shows the performance comparison between the pose recognition model constructed and trained by the method of the present invention and multiple existing models. Among them, the pose recognition model constructed by the present invention includes 4 feature extraction branches, which respectively extract four feature maps with the number of channels being 32, 64, 128, and 256. Each feature map is compressed in channels through a 1×1 convolution. The specific structure and parameter settings of the depth stacking module are obtained through search. The finally searched number of stages of the depth stacking module is 44. The experimental configurations and experimental results of the remaining parts are shown in Table 1:
[0087] Table 1
[0088]
[0089] In Table 1, AP is the average precision of the model, AR is the average recall rate, and AP50, AP75, APM, and APL are all indicators further subdivided based on AP. AP50 refers to the average precision when the IoU (Intersection over Union) threshold is 0.5, AP75 refers to the average precision when the IoU threshold is 0.75, APM refers to the average precision for medium-scale targets, and APL refers to the average precision for large-scale targets.
[0090] As can be seen from Table 1, when the pose recognition model designed by the present invention is compared with traditional models such as Dite-HRNet-30, Lite-HRnet-30, and GhostPoseNet that adopt the same input layer and conversion layer, the performance indicators AP and AR of the pose recognition model searched by the present invention are respectively improved by 1.8 to 14.9 percentage points and 1.6 to 13.7 percentage points, indicating that better performance can still be maintained under relatively lightweight settings. At the same time, a higher accuracy is obtained on the higher-precision indicator AP50, proving that the lightweight human pose recognition model of the present invention is more effective in recognizing small targets.
[0091] Moreover, the method of the present invention is simple, effective, and has good robustness, and is more suitable for actual application scenarios, especially has better use effects in the human pose recognition task of the intelligent security system.
[0092] The method provided by the present invention can be applied to electronic devices. Specifically, the electronic device can be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc. There is no limitation here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.
[0093] Corresponding to the above method for constructing a lightweight pose recognition model, an embodiment of the present invention also provides a device for constructing a lightweight pose recognition model, and the device includes:
[0094] A first definition module, configured to define a pose recognition model; the pose recognition model includes: an input layer, a conversion layer, a depth stacking module, and an output layer; the input layer is configured to send an image to be recognized for pose into the conversion layer; the conversion layer includes a plurality of feature extraction branches, and the plurality of feature extraction branches are configured to extract feature maps of different scales from the image; the depth stacking module includes a plurality of cascaded depth search modules; each depth search module includes a search layer and a fusion layer; the search layer includes a plurality of feature extraction units corresponding one by one to the plurality of feature extraction branches, and is configured to perform feature extraction on the plurality of feature maps input to the search layer, and the fusion layer is configured to perform feature fusion on the feature maps extracted by the plurality of feature extraction units and output to the next-level depth search module; the output layer is configured to predict a pose recognition result according to the feature map output by the last-level fusion layer;
[0095] A second definition module, configured to define a depth search space; the depth search space defines the cascaded number space of the depth stacking module in the pose recognition model;
[0096] A third definition module, configured to define a fusion selection space according to the depth search space; the fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select a fusion object;
[0097] A network search module, configured to search for a lightweight pose recognition model for pose recognition in a deep search space and a fusion selection space.
[0098] Optionally, the network search module is specifically configured to:
[0099] In a deep search space and a fusion selection space, use an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0100] Optionally, the network search module includes: a first search unit and a second search unit;
[0101] The first search unit is configured to, in a deep search space and a fusion selection space, by giving a fixed mixing operator to the pose recognition model, perform evolutionary algorithm-based channel structure optimization search to obtain a number of channel structure optimized subnets; wherein, the channel structure optimization search is used to determine the cascade number of the deep stacking module and the selection strategy of each fusion layer for the fusion object through search.
[0102] The second search unit is configured to define a mixing operator search space, and in the mixing operator search space, use the number of channel structure optimized subnets as the search set, perform operator optimization search using an evolutionary algorithm to obtain a lightweight pose recognition model for pose recognition; wherein, the operator optimization search is used to determine the operator type adopted by the lightweight pose recognition model through search.
[0103] Optionally, the above device further includes: a supernet weight pre-training module;
[0104] The supernet weight pre-training module is configured to, before using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, perform supernet weight pre-training through a single-path uniform sampling strategy in a deep search space and a fusion selection space, so that the search module uses the pre-trained supernet weights to search for a lightweight pose recognition model for pose recognition using an evolutionary algorithm.
[0105] Optionally, the above device further includes: a subnet random sampling search module;
[0106] The subnet random sampling search module is configured to, after completing the supernet weight pre-training and before using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, perform subnet random sampling search in a deep search space and a fusion selection space according to the pre-trained weights to obtain a number of subnets as the parent subnets when subsequently using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition.
[0107] Based on the same inventive concept, the present invention also provides a pose recognition method, including the following steps:
[0108] Step 1, obtain an image to be subjected to pose recognition;
[0109] Step 2: Input the acquired image into a pre-trained lightweight pose recognition model for pose recognition, so that the lightweight pose recognition model outputs a pose recognition result for the image;
[0110] Among them, the lightweight pose recognition model is constructed by using any one of the lightweight pose recognition model construction methods provided above and has been pre-trained.
[0111] It should be noted that for the apparatus / pose recognition method embodiments, since they are basically similar to the embodiments of the lightweight pose recognition model construction method, the description is relatively simple. For the relevant parts, refer to the description of the embodiments of the lightweight pose recognition model construction method.
[0112] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention.
[0113] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0114] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed present invention, those skilled in the art can understand and implement other changes of the disclosed embodiments by viewing the accompanying drawings and the disclosed content. In the description of the present invention, the term "including" does not exclude other components or steps, the term "one" or "a" does not exclude the case of multiple, and the meaning of "multiple" is two or more, unless otherwise clearly specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0115] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for constructing a lightweight gesture recognition model, characterized in that, Including: Defining a pose recognition model; The pose recognition model includes: an input layer, a conversion layer, a deep stacking module, and an output layer; the input layer is used to send an image to be recognized for pose into the conversion layer; the conversion layer includes a plurality of feature extraction branches, and the plurality of feature extraction branches are used to extract feature maps of different scales from the image; the deep stacking module includes a plurality of cascaded deep search modules; each deep search module includes a search layer and a fusion layer; the search layer includes a plurality of feature extraction units corresponding one-to-one to the plurality of feature extraction branches, and is used to perform feature extraction on the plurality of feature maps input to the search layer, and the fusion layer is used to perform feature fusion on the feature maps extracted by the plurality of feature extraction units and output to the next-level deep search module; the output layer is used to output a pose recognition result according to the feature map output by the last-level fusion layer; Defining a deep search space; the deep search space defines the cascading number space of the deep stacking module in the pose recognition model; Defining a fusion selection space according to the deep search space; the fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select fusion objects; Searching for a lightweight pose recognition model for pose recognition under the deep search space and the fusion selection space.
2. The method for constructing a lightweight pose recognition model according to claim 1, wherein, The searching for a lightweight pose recognition model for pose recognition under the deep search space and the fusion selection space includes: Searching for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm under the deep search space and the fusion selection space.
3. The method for constructing a lightweight pose recognition model according to claim 2, wherein The searching for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm under the deep search space and the fusion selection space includes: Under the deep search space and the fusion selection space, by giving the pose recognition model a fixed mixing operator, using an evolutionary algorithm to perform channel structure optimization search to obtain a number of channel structure optimized subnets; wherein, the channel structure optimization search is used to determine the cascading number of the deep stacking module and the selection strategy of each fusion layer for fusion objects through search; Defining a mixing operator search space, and under the mixing operator search space, using the number of channel structure optimized subnets as a search set, and using an evolutionary algorithm to perform operator optimization search to obtain a lightweight pose recognition model for pose recognition; wherein, the operator optimization search is used to determine the operator type adopted by the lightweight pose recognition model through search.
4. The method for constructing a lightweight pose recognition model according to claim 3, characterized in that Before searching for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm, the method further includes: under the deep search space and the fusion selection space, performing supernet weight pre-training through a single-path uniform sampling strategy, so as to search for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm according to the pre-trained supernet weights.
5. The method for constructing a lightweight pose recognition model according to claim 4, wherein After completing the supernet weight pre-training and before searching for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm, the method further includes: Under the depth search space and the fusion selection space, subnet random sampling search is performed according to pre-trained weights to obtain a number of sub-networks, which are used as the parent sub-networks when the evolutionary algorithm is subsequently used to search for a lightweight pose recognition model for pose recognition.
6. An apparatus for constructing a lightweight pose recognition model, characterized in that The device includes: A first definition module for defining a pose recognition model; the pose recognition model includes an input layer, a conversion layer, a depth stacking module, and an output layer; the input layer is used to send an image to be recognized for pose into the conversion layer; the conversion layer includes a plurality of feature extraction branches, and the plurality of feature extraction branches are used to extract feature maps of different scales from the image; the depth stacking module includes a plurality of cascaded depth search modules; each depth search module includes a search layer and a fusion layer; the search layer includes a plurality of feature extraction units corresponding one-to-one to the plurality of feature extraction branches, and is used to perform feature extraction on the plurality of feature maps input to the search layer, and the fusion layer is used to perform feature fusion on the feature maps extracted by the plurality of feature extraction units and output to the next-level depth search module; the output layer is used to predict a pose recognition result according to the feature map output by the last-level fusion layer; A second definition module for defining a depth search space; the depth search space defines the cascade number space of the depth stacking module in the pose recognition model; A third definition module for defining a fusion selection space according to the depth search space; the fusion selection space defines the selection strategy space for the fusion layer in the pose recognition model to select a fusion object; A network search module for searching for a lightweight pose recognition model for pose recognition under the depth search space and the fusion selection space.
7. The apparatus for constructing a lightweight pose recognition model according to claim 6, wherein The network search module is specifically used for: Searching for a lightweight pose recognition model for pose recognition by using an evolutionary algorithm under the depth search space and the fusion selection space.
8. The apparatus for constructing a lightweight pose recognition model according to claim 7, wherein The network search module includes: a first search unit and a second search unit; The first search unit is used to perform channel structure optimization search by using an evolutionary algorithm by giving the pose recognition model a fixed mixing operator under the depth search space and the fusion selection space to obtain a number of channel structure optimized sub-networks; wherein, the channel structure optimization search is used to determine the cascade number of the depth stacking module and the selection strategy of each fusion layer for the fusion object by search; The second search unit is used to define a mixing operator search space, and perform operator optimization search by using an evolutionary algorithm with the number of channel structure optimized sub-networks as the search set under the mixing operator search space to obtain a lightweight pose recognition model for pose recognition; wherein, the operator optimization search is used to determine the operator type adopted by the lightweight pose recognition model by search.
9. The apparatus for constructing a lightweight pose recognition model according to claim 8, wherein The device further includes: a super-network weight pre-training module; The supernet weight pre-training module is used to perform supernet weight pre-training through a single-path uniform sampling strategy under the deep search space and the fusion selection space before using an evolutionary algorithm to search for a lightweight pose recognition model for pose recognition, so that the search module uses the evolutionary algorithm to search for a lightweight pose recognition model for pose recognition according to the pre-trained supernet weights.
10. A gesture recognition method, characterized in that, It includes: Obtain an image to be recognized for pose; Input the image into the pre-trained lightweight pose recognition model for pose recognition, so that the lightweight pose recognition model outputs a pose recognition result for the image; Among them, the lightweight pose recognition model is constructed by using the construction method of the lightweight pose recognition model according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-person body posture estimation method and system based on knowledge distillation
CN115187660A