Estimation method and device for multiple human body postures, electronic equipment and medium
Through a single-stage multi-person posture estimation method and explicit prompt learning technology, the coordinates of key points of multiple human bodies are directly predicted, and the problem of insufficient timeliness and usability in the existing technology is solved, achieving a more efficient and accurate posture estimation effect.
Patent Information
- Application Number
- CN202510123031.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
The existing multi-human pose estimation technology has poor timeliness and availability, and lacks effective cross-task coordination and information sharing.
The single-stage multi-person posture estimation method is used to directly predict the coordinates of key points of multiple human bodies through a unified network structure, and the prompt information generated by object detection is integrated into the pose estimation process using explicit prompt learning and lightweight adapters.
It improves the timeliness and usability of human pose estimation, enhances the model's processing ability for complex scenarios and multi-task requirements, and achieves more accurate and efficient multi-task collaborative processing.
Smart Images

Figure CN120047971A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multi-human posture estimation, and in particular to a multi-human posture estimation method, device, electronic equipment and medium. Background Art
[0002] Pose Estimation is an important research direction in the field of computer vision. Its goal is to identify and locate the key points of the human body from images or videos, and then infer the human body's posture and movement. This technology has a wide range of applications in motion analysis, virtual reality, augmented reality, behavior recognition, and human-computer interaction. The existing technology uses a multi-stage method to predict the posture of multiple human bodies by gradually refining the key points. That is to say, the existing technology first locates the human body and then performs posture estimation, which leads to poor timeliness and availability of human posture estimation. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a method, device, electronic device and medium for estimating multiple human body postures, so as to solve the above-mentioned problems existing in the prior art and improve the timeliness and usability of human body posture estimation.
[0004] In a first aspect, a method for estimating multiple human body postures is provided, and the method may include:
[0005] Acquire image data to be processed; the image data includes a plurality of human objects;
[0006] The image data is processed using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that incorporates prompt learning information.
[0007] In a possible implementation, the detection head of the pose estimation model includes a bounding box generation unit, a lightweight adapter, and a pose estimation unit connected in sequence;
[0008] The bounding box generating unit processes the image data to obtain a bounding box coordinate prompt of a human body in the image data;
[0009] The lightweight adapter processes the bounding box coordinate hint to obtain a hint embedding vector;
[0010] The posture estimation unit processes the prompt embedding vector to obtain the posture estimation feature.
[0011] In a possible implementation, the bounding box generation unit includes a category detector and a bounding box detector;
[0012] The image data is processed using the category detector and the bounding box detector to obtain category probabilities of target categories and bounding box coordinate prompts.
[0013] In a possible implementation, obtaining a bounding box coordinate hint in the image data includes:
[0014] Normalizing the acquired bounding box features to obtain a target ratio value relative to the image data;
[0015] Performing restoration processing on the target ratio value to obtain an absolute pixel value;
[0016] The absolute pixel value is mapped to the image data to obtain the bounding box coordinate prompt.
[0017] In a possible implementation, the lightweight adapter includes a first linear layer, a SiLU activation function, and a second linear layer connected in sequence;
[0018] Using the first linear layer to perform dimension-upgrading processing on the bounding box coordinate prompt to obtain an initial coordinate prompt;
[0019] Using the SiLU activation function to perform nonlinear processing on the initial coordinate prompt to obtain an initial embedding vector;
[0020] The second linear layer is used to perform dimensionality reduction processing on the initial embedding vector to obtain the hint embedding vector.
[0021] In a possible implementation, the posture estimation unit includes at least 2 CM modules;
[0022] For any CM module, the CM module includes 2 1×1 convolutional layers, 1 3×3 separable convolutional layer and 1 residual layer.
[0023] In a second aspect, a device for estimating multiple human postures is provided, which may include:
[0024] An acquisition unit, used for acquiring image data to be processed; the image data includes a plurality of human objects;
[0025] A processing unit is used to process the image data using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that integrates prompt learning information.
[0026] In a possible implementation, the pose estimation model includes a bounding box generation unit, a lightweight adapter, and a pose estimation unit connected in sequence;
[0027] The bounding box generating unit processes the image data to obtain a bounding box coordinate prompt of a human body in the image data;
[0028] The lightweight adapter processes the bounding box coordinate hint to obtain a hint embedding vector;
[0029] The posture estimation unit processes the prompt embedding vector to obtain the posture estimation feature.
[0030] In a third aspect, an electronic device is provided, the electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0031] Memory, used to store computer programs;
[0032] The processor is used to implement any method step described in the first aspect when executing the program stored in the memory.
[0033] In a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any method step described in the first aspect is implemented.
[0034] The present application provides a method for estimating multiple human poses, which is simply single-stage multi-human pose estimation (Single-Stage Human Pose Estimation), which can be specifically understood as a type of computer vision method that aims to directly estimate the coordinates of multiple human key points from input image data without the need for complex multi-step processing. The multi-stage method of the prior art usually decomposes the pose estimation task into multiple steps, such as first detecting the human body area and then locating the key points in the detected area. Compared with the multi-stage method, the method of the present application directly predicts the coordinates of multiple human key points on an image, and uses a unified network structure and a single forward propagation to complete the entire process from image input to key point output, which simplifies the process and improves efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 A system architecture diagram of a method for estimating multiple human body postures provided in an embodiment of the present application;
[0037] Figure 2 A schematic diagram of a flow chart of a method for estimating multiple human body postures provided in an embodiment of the present application;
[0038] Figure 3 A schematic diagram of the structure of a posture estimation model detection head provided in an embodiment of the present application;
[0039] Figure 4 A schematic diagram comparing the prior art provided in the embodiment of the present application with a method for estimating multiple human body postures;
[0040] Figure 5 A schematic diagram of the structure of a device for estimating multiple human postures provided in an embodiment of the present application;
[0041] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0043] For ease of understanding, the terms involved in the embodiments of the present application are explained below:
[0044] Explicit Prompt Learning: It is a machine learning technique, especially in the field of natural language processing, that involves guiding the model to generate or improve its response through explicit prompts or questions. In the field of images, explicit prompt learning is a technique that uses explicit prompts or instructions to guide computer vision models. This method can help the model better understand the tasks and goals, thereby improving its performance on specific tasks. Prompt learning usually includes the following concepts:
[0045] 1) Explicit Prompt: refers to the use of clear and specific instructions or descriptions in the model input to guide the model's behavior.
[0046] 2) Task Adaptation: This is done by adjusting the model’s input or prompts so that the model can better complete a specific task. This usually involves using explicit prompts to guide the model to focus on specific aspects of the task.
[0047] Pose Estimation: It is an important field in computer vision, which aims to determine the position and posture of an object or human body in space. Pose estimation is divided into pose estimation (identifying and locating the position of key points of a human body or object in an image on a two-dimensional plane), 3D pose estimation (identifying and locating key points of an object or human body in three-dimensional space, involving depth information), and multi-human pose estimation (simultaneously identifying and tracking the postures of multiple individuals in an image).
[0048] The method for estimating multiple human postures provided in the embodiment of the present application can be applied in Figure 1 In the system architecture shown in Figure 1 As shown, the system may include: a server and a terminal. The server may be a physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal may be a user equipment (UE) such as a mobile phone, smart phone, laptop, digital broadcast receiver, personal digital assistant (PDA), tablet computer (PAD), handheld device, vehicle-mounted device, wearable device, computing device or other processing device connected to a wireless modem, mobile station (MS), mobile terminal, etc. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0049] The terminal is used to obtain the image data to be processed and send the image data to the server.
[0050] The server is used to receive image data to execute a method for estimating multiple human body postures provided in the present application.
[0051] In the task of multi-human pose estimation, it is usually necessary to process the poses of multiple human bodies at the same time and solve problems such as body occlusion and overlap. Research in this field mainly focuses on how to effectively separate and identify different human instances and achieve high-precision key point positioning in complex scenes. Although human pose estimation technology has made significant progress, it still faces the following problems in practical applications:
[0052] 1) Task independence: Existing pose estimation models usually run independently when processing each task. Object detection and pose estimation are usually processed separately, with the object detection model responsible for locating the target area, while the pose estimation model predicts key points within the detected area. This separate processing method limits the model's coordination and information sharing across tasks.
[0053] 2) Lack of context awareness: Since there is no effective interaction between the object detection and pose estimation models, these models often lack a comprehensive understanding of the context. In complex scenes, people may occlude each other, and the background may be complex and changeable. The model cannot fully utilize the context information provided by the object detection stage to optimize the pose estimation results.
[0054] 3) Insufficient information transfer: Existing methods do not transfer information sufficiently during the transition from object detection to pose estimation. There is no effective coupling between the bounding box information generated in the object detection stage and the key point prediction in the pose estimation stage, which may lead to the performance degradation of the model when dealing with occlusion, overlap or complex background.
[0055] 4) Model limitations: Most existing models focus on the optimization of a single task, while ignoring the complementarity across tasks. For example, the object detection model mainly focuses on detection accuracy, while the pose estimation model focuses on the accuracy of key points. The lack of joint optimization and information fusion of these two tasks limits the performance of the model in practical applications.
[0056] Therefore, the purpose of the embodiments of the present application is to provide a method for estimating multiple human body postures, so as to solve the above-mentioned problems existing in the prior art and to improve the timeliness and usability of human body posture estimation.
[0057] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.
[0058] Figure 2 The following is a flow chart of a method for estimating multiple human postures provided in an embodiment of the present application. Figure 2 As shown, the method may include:
[0059] S210: Obtain image data to be processed.
[0060] The image data includes multiple human objects.
[0061] It should be noted that the acquired data may also be video data, and the video data is subjected to frame decomposition processing to obtain a plurality of image data.
[0062] S220: Use the trained posture estimation model to process the image data to obtain posture estimation features of each human body in the image data.
[0063] Specific, combined Figure 3 As shown, the detection head of the pose estimation model includes a bounding box generation unit, a lightweight adapter, and a pose estimation unit which are connected in sequence.
[0064] Wherein, A, a bounding box generation unit processes the image data to obtain a bounding box coordinate prompt of a human body in the image data;
[0065] The bounding box generation unit includes a category detector and a bounding box detector.
[0066] The category detector and bounding box detector are used to process the image data respectively to obtain the category probability of the target category and the bounding box coordinate prompts.
[0067] Before obtaining the bounding box coordinate prompt in the image data, the bounding box features are also obtained. The acquisition process includes:
[0068] While generating the bounding box, the bounding box detector extracts the bounding box features from the image data; specifically, the input image data is processed and the bounding box detection information of each human body is output. The bounding box detection information usually includes the coordinates of the upper left corner, the coordinates of the lower right corner, the height and width of the bounding box. While generating the bounding box, the feature information within each bounding box is extracted. Usually, it includes the pixel value, texture, shape and other information within the bounding box. These feature information will be organized into a feature map (Feature Map), and each point in the feature map corresponds to a local area within the bounding box. The bounding box feature is composed of certain specific points in the feature map (usually the center point of the bounding box) and the size and position information of the bounding box. The context information and the bounding box detection information in the feature map together constitute the bounding box feature.
[0069] The bounding box detector processes the image data to obtain the bounding box coordinate prompts in the image data, including:
[0070] Normalize the acquired bounding box features to obtain the target ratio value relative to the image data;
[0071] The target ratio value is restored to obtain the absolute pixel value;
[0072] Map absolute pixel values to image data to get bounding box coordinate hints.
[0073] This step can also be understood as: the bounding box generation unit divides the detection task into two branches: one for target category prediction and the other for bounding box regression; that is, through the category detector and the bounding box detector, the bounding box coordinate prompts and category probabilities of the predicted target are output respectively.
[0074] That is, the bounding box detection information is decoded (Bounding box decoding), normalized and the coordinate mapping is restored to the absolute coordinates of the original image, so that the bounding box coordinates prompt (Bounding box coordinates prompt, BP) is obtained.
[0075] Assume F BF ∈ 4×H×W Represents the bounding box coordinate hint, H and W represent the height and width of the feature map (image data), (x center ,y center ) represents the center point coordinates of each feature point of the bounding box parameters, vox ,w box Represents the height and width of the bounding box respectively. The following formula shows the decoding process of the bounding box coordinate hint:
[0076]
[0077] Among them, It is the coordinate of the upper left corner, and rb is the coordinate of the lower right corner. image and W image are the height and width of the original image data respectively. After the above formula is used to transform the bounding box features, the bounding box coordinate hint F is obtained. BP ∈ 4×H×W .
[0078] These bounding boxes not only locate the position of the object in the image, but also convey rich contextual information, such as the scale, orientation, and relative position of the object. They can provide a clear contextual background for the pose estimation unit, which helps to guide and optimize the reasoning process of the pose estimation task. In this process, the back propagation of the gradient is truncated (Stop gradient), because the bounding box parameters are only used as prompt information, and the reverse calculation of the pose estimation loss cannot be performed here.
[0079] B. The lightweight adapter processes the bounding box coordinate hint to obtain the hint embedding vector;
[0080] Specifically, the lightweight adapter includes a first linear layer, a SiLU activation function, and a second linear layer connected in sequence;
[0081] The first linear layer is used to perform dimension-upgrading on the bounding box coordinate hint to obtain the initial coordinate hint.
[0082] The SiLU activation function is used to perform nonlinear processing on the initial coordinate prompts to obtain the initial embedding vector;
[0083] The second linear layer is used to reduce the dimension of the initial embedding vector to obtain the hint embedding vector.
[0084] The process can be understood as follows: the adapter uses the first linear layer to increase the dimension of the input bounding box coordinate prompt BP, effectively extracting the core elements of the prompt information. Next, through the SiLU activation function, nonlinear characteristics are introduced to increase the model's ability to capture complex patterns, so that the embedded vector can better represent the diversity and complexity of the original prompt information. Finally, after the second linear layer performs a dimensionality reduction operation, the features after dimensionality increase are expanded back to a dimension suitable for model processing.
[0085] The lightweight adapter is designed to efficiently integrate the prompt information generated by the object detection and embed it into the pose estimation process. The adapter achieves the dimensionality upscaling and downscaling of the prompt information to generate the prompt embedding vector by combining two linear layers (the first linear layer and the second linear layer) and the SiLU activation function.
[0086] In summary, the processing of lightweight adapter can be expressed as:
[0087] Adaptor=Linear(SiLU(Linear(F BP )))
[0088] Among them, the first linear layer will be the initial coordinate hint F BP The dimension is increased to Ck, and the second linear layer reduces the dimension from Ck to C, which is aligned with the feature dimension of the posture estimation task in the posture estimation unit.
[0089] In this approach, the lightweight adapter retains the key features of the prompt information while generating a refined and efficient prompt embedding vector through the process of dimensionality increase and decrease, providing a more adaptive input for the downstream tasks of the model, making the entire prompt learning process more flexible, accurate and efficient, to ensure that the explicit prompt information can be seamlessly compatible with the posture estimation unit. This design not only effectively reduces the computational complexity, but also ensures the accurate transmission and effective use of explicit prompts, thereby achieving flexible adaptation of detection information in different scenarios and efficient task processing.
[0090] It should be noted that in order to enhance the robustness of the pose estimation unit and the effective transfer of cross-task information, N lightweight adapters are used to provide explicit prompts to the pose estimation units at different stages.
[0091] C. The pose estimation unit processes the hint embedding vector to obtain the pose estimation feature.
[0092] Among them, continue to combine Figure 3 As shown in the figure, the pose estimation unit includes multiple CM (Conv MLP) modules; that is, the pose estimation unit has N stages to receive these N explicit hint embedding vectors, and the pose estimation unit is composed of multiple lightweight ConvMLP (CM) modules, which play a key role in the entire architecture. Each CM module consists of 2 1×1 convolutional layers, 1 3×3 separable convolutional layer and 1 residual layer.
[0093] Among them, the 1×1 convolution layer is mainly used for compression and expansion of feature channels, which improves computational efficiency while retaining rich feature information.
[0094] The 3×3 separable convolution layer greatly reduces the amount of computation by decomposing the standard convolution into depth convolution and point convolution, while maintaining sensitivity to local spatial features, enabling the model to capture the detailed features of the human body more accurately.
[0095] The introduction of the residual layer helps alleviate the gradient vanishing problem in deep network training and ensures that the model can effectively learn features at a deep level.
[0096] In summary, the processing of the attitude estimation unit can be expressed as:
[0097]
[0098] in, represents the hint embedding vector, represents the output characteristics after passing through the CM unit, represents the features after the prompt embedding vector and CM unit are fused at each stage, F pose ∈ C×H×W Represents the pose estimation feature.
[0099] It is understandable that the posture estimation features of the corresponding human body in the image data can be represented by the relative positions of the key points of each human body.
[0100] The CM unit can not only learn the local spatial features of the human body in the image, but also effectively integrate the contextual information and bounding box prompt information in the target detection task. In addition, the pose estimation detection head based on explicit prompt learning integrates the information in the target detection task into the pose estimation task, making the entire pose estimation model more coordinated and efficient when processing multiple tasks simultaneously. This not only enhances the accuracy of pose estimation, but also improves the robustness and adaptability of the pose estimation model in complex environments, thus providing strong support for multi-task processing in practical application scenarios.
[0101] In some embodiments, the posture estimation model of the present application is based on YOLO-Pose, and proposes an explicit prompt learning posture estimation detection head, uses the bounding box predicted by the target detection task as an explicit prompt, and designs a lightweight adapter to optimize the explicit prompt process, so as to achieve efficient task adaptation in the posture estimation task, realize information interaction and collaborative processing of the two tasks, and use a single-stage multi-human posture estimation method to simultaneously complete the dual tasks of human target detection and human posture estimation, and can quickly process complex image and video data in a variety of scenarios. Figure 4 The comparison between the existing technology and the proposed method in terms of process processing and framework novelty is demonstrated, which fully demonstrates the superiority of the proposed method in the task of pose estimation.
[0102] In some examples, the configuration of the training set of the posture estimation model during the training process includes: collecting image or video data containing human posture information and annotating key points, which can be done manually or with the assistance of automatic annotation tools, or using open source datasets such as COCO for experiments. The collected data is preprocessed, and effective data filtering and useless data are screened out. Finally, the corresponding format of the annotation file is converted to obtain the training set (training samples and training labels).
[0103] During the training process of the pose estimation model, it is also necessary to define a loss function to measure the gap between the model prediction and the true annotation, so as to adjust the pose estimation model parameters to minimize the prediction error. That is, during the training phase, the classification features (cls.feature), bounding box parameters (box.feature) and the categories and bounding boxes of the true annotations are used for loss calculation, that is, loss calculation: the above-mentioned predicted values (classification features and bounding box features) are compared with the true annotations to determine the difference. Commonly used loss functions include but are not limited to classification loss, such as cross-entropy loss (Cross-EntropyLoss), which is used to measure the difference between the classification prediction and the true category. Positioning loss: such as smooth L1 loss (Smooth L1Loss) or mean square error (MSE), which is used to measure the position deviation between the predicted bounding box and the true bounding box. In many cases, classification loss and positioning loss are considered at the same time, and the final comprehensive loss is obtained by weighted summation.
[0104] The configured initial model is trained using the training set. During the training process, online data augmentation technology is used to improve the generalization and robustness of the model detection performance. In order to ensure the diversity of the training data, the probability factors and expansion coefficients of different data augmentation methods are limited to a certain range. This ensures that the distribution of image data loaded in each round of training is different from the previous round, effectively increasing the diversity of the data. The detailed data augmentation methods and their probability factors and expansion coefficients are shown in Table 1.
[0105] Table 1 - Probability factors and expansion coefficients of different data augmentation methods
[0106]
[0107] For hardware equipment and training strategies, we use Ubuntu 18.04LTS system, GeForce RTX4090Ti graphics card, pytorch2.0.1 network development framework, and Pycharm integrated development environment. We uniformly set the training rounds to 300, the batch size to 4, the input image size to 640×640, and the AdamW optimizer. The initial learning rates of the encoder and decoder are set to 0.002 and 0.01, respectively. We use a linear decay learning rate adjustment strategy, and the weight decay is set to 0.05. During the training process, in order to ensure the accuracy of the bounding box positioning of the explicit prompt information, the Adapter prompt process and prompt embedding are turned on in the last 100 training rounds, because in the early stages of training, the bounding box information of the human body detection positioning is not very accurate, and invalid information is easily introduced.
[0108] Finally, the trained pose estimation model is subjected to inference detection on the image dataset, that is, the performance of the pose estimation model is verified through the validation set. First, the input image data is preprocessed and scaled to the input size expected by the pose estimation model. After that, it is transmitted to the pose estimation model, and the detection head uses appropriate prediction box decoding and pose estimation decoding methods. Finally, the human body positioning bounding box and human body key points are detected at the same time.
[0109] The present application provides a method for estimating multiple human poses, which is simply a single-stage multi-human pose estimation (Single-Stage Human Pose Estimation), which can be specifically understood as a type of computer vision method that aims to directly estimate the coordinates of multiple human key points from input image data without complex multi-step processing. The multi-stage method in the prior art usually decomposes the pose estimation task into multiple steps, such as first detecting the human body area and then locating the key points in the detected area. Compared with the multi-stage method, the single-stage method of the present application directly predicts the coordinates of multiple human key points on an image, and uses a unified network structure to complete the entire process from image input to key point output using a forward propagation, which simplifies the process and improves efficiency. The present application first introduces prompt learning into the field of human pose estimation, and realizes cross-task knowledge sharing and information optimization of the model in multi-task learning by using the bounding box predicted by target detection as explicit prompt information. It expands the scope of application of prompt learning and also provides new research ideas for multi-task learning. Secondly, the bounding box of target detection is used as systematic and structured prompt information to clarify the context of the pose estimation task and improve the accuracy and efficiency of key point positioning. The model's ability to focus on the human body area is significantly improved, and more accurate and intuitive multi-task collaborative processing is achieved, providing new technical support for high-precision pose estimation. Secondly, an efficient lightweight adapter is configured to seamlessly integrate the prompt information generated by target detection into the pose estimation process, reducing the computational complexity and ensuring the effective transmission of explicit prompts. In addition, the explicit prompt learning strategy optimizes the transmission and fusion of feature information between tasks, significantly improving the performance of the model in complex scenes and multi-task requirements, and providing a new technical path for multi-task learning and cross-application. Finally, the simultaneous processing of human target detection and pose estimation tasks is achieved in the same framework, and real-time processing capabilities and efficient computing are achieved through the introduction of explicit prompt learning and lightweight adapters.
[0110] Corresponding to the above method, the embodiment of the present application also provides a device for estimating multiple human postures, such as Figure 5 As shown, the device comprises:
[0111] An acquisition unit 510 is used to acquire image data to be processed; the image data includes a plurality of human objects;
[0112] The processing unit 520 is used to process the image data using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that integrates prompt learning information.
[0113] The functions of each functional unit of a device for estimating multiple human postures provided in the above-mentioned embodiment of the present application can be realized through the above-mentioned method steps. Therefore, the specific working process and beneficial effects of each unit in a device for estimating multiple human postures provided in the embodiment of the present application will not be repeated here.
[0114] The present application also provides an electronic device, such as Figure 6 As shown, it includes a processor 610 , a communication interface 620 , a memory 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other via the communication bus 640 .
[0115] Memory 630, for storing computer programs;
[0116] The processor 610 is used to implement the following steps when executing the program stored in the memory 630:
[0117] Acquire image data to be processed; the image data includes a plurality of human objects;
[0118] The image data is processed using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that incorporates prompt learning information.
[0119] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0120] The communication interface is used for communication between the above electronic device and other devices.
[0121] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0122] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0123] The implementation methods and beneficial effects of the components of the electronic device in the above embodiments to solve the problems can be seen in Figure 2 The various steps in the illustrated embodiment are implemented, therefore, the specific working process and beneficial effects of the electronic device provided by the embodiment of the present application are not repeated here.
[0124] In another embodiment provided in the present application, a computer-readable storage medium is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes a method for estimating multiple human postures described in any of the above embodiments.
[0125] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute a method for estimating multiple human postures described in any one of the above embodiments.
[0126] Those skilled in the art will appreciate that the embodiments in the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt a complete hardware embodiment, a complete software embodiment, or a form of an embodiment combining software and hardware. Moreover, the present application may adopt a form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0130] Unless otherwise defined, the technical terms or scientific terms used in this application should be understood by people with ordinary skills in the field to which the present invention belongs. "First", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect", "couple" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0131] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the present application embodiments are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application embodiments.
[0132] Obviously, those skilled in the art can make various changes and modifications to the embodiments in the present application without departing from the spirit and scope of the embodiments in the present application. Thus, if these modifications and variations of the embodiments in the present application are within the scope of the embodiments in the present application and their equivalents, the embodiments in the present application are also intended to include these modifications and variations.
Claims
1. A method for estimating multiple human body postures, characterized in that: The method comprises: Acquire image data to be processed; the image data includes a plurality of human objects; The image data is processed using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that incorporates prompt learning information.
2. The method according to claim 1, characterized in that The detection head of the pose estimation model includes a bounding box generation unit, a lightweight adapter and a pose estimation unit connected in sequence; The bounding box generating unit processes the image data to obtain a bounding box coordinate prompt of a human body in the image data; The lightweight adapter processes the bounding box coordinate hint to obtain a hint embedding vector; The posture estimation unit processes the prompt embedding vector to obtain the posture estimation feature.
3. The method according to claim 2, characterized in that The bounding box generation unit includes a category detector and a bounding box detector; The image data is processed using the category detector and the bounding box detector to obtain category probabilities of target categories and bounding box coordinate prompts.
4. The method according to claim 3, characterized in that Obtaining a bounding box coordinate hint in the image data includes: Normalizing the acquired bounding box features to obtain a target ratio value relative to the image data; Performing restoration processing on the target ratio value to obtain an absolute pixel value; The absolute pixel value is mapped to the image data to obtain the bounding box coordinate prompt.
5. The method according to claim 3, characterized in that The lightweight adapter includes a first linear layer, a SiLU activation function, and a second linear layer connected in sequence; Using the first linear layer to perform dimension-upgrading processing on the bounding box coordinate prompt to obtain an initial coordinate prompt; Using the SiLU activation function to perform nonlinear processing on the initial coordinate prompt to obtain an initial embedding vector; The second linear layer is used to perform dimensionality reduction processing on the initial embedding vector to obtain the hint embedding vector.
6. The method according to claim 5, characterized in that The posture estimation unit includes at least 2 CM modules; For any CM module, the CM module includes 2 1×1 convolutional layers, 1 3×3 separable convolutional layer and 1 residual layer.
7. A device for estimating multiple human body postures, characterized in that: The device comprises: An acquisition unit, used for acquiring image data to be processed; the image data includes a plurality of human objects; A processing unit is used to process the image data using a trained posture estimation model to obtain posture estimation features of each human body in the image data; the posture estimation model is a model that integrates prompt learning information.
8. The device according to claim 7, characterized in that The pose estimation model includes a bounding box generation unit, a lightweight adapter, and a pose estimation unit connected in sequence; The bounding box generating unit processes the image data to obtain a bounding box coordinate prompt of a human body in the image data; The lightweight adapter processes the bounding box coordinate hint to obtain a hint embedding vector; The posture estimation unit processes the prompt embedding vector to obtain the posture estimation feature.
9. An electronic device, characterized in that: The electronic device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 6 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 6 are implemented.