Non-vision-field human body posture modeling method and device based on deep learning network
Through the non-sighted human posture modeling method based on deep learning network, the three-dimensional point cloud data and pose recognition model are used to generate an accurate three-dimensional model, which solves the problem that traditional technology cannot capture object information outside the field of view and large errors in pose modeling, and achieves high-precision non-sighted human posture modeling.
Patent Information
- Application Number
- CN202510275026.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional optical imaging technology cannot capture information about objects located outside the field of view, obscured or hidden, and the posture recognition and modeling process are independent and there are large errors.
A non-sighted human posture modeling method based on deep learning network is adopted. By obtaining the three-dimensional point cloud data of the target non-sighted data, dimensionality reduction processing is performed to obtain the point cloud feature map, input the feature map to the pose recognition model to predict joint data, and a target three-dimensional model is generated based on the prediction data.
It effectively reduces errors and inconsistencies in the modeling process, improves the accuracy of posture modeling, and improves the practicality of non-sight technology.
Smart Images

Figure CN120198587A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of optical non-line-of-sight imaging technology, and more particularly to a method, apparatus, and device for non-line-of-sight human pose modeling based on a deep learning network. Background Art
[0002] NLOS (Non-Line-of-Sight Imaging, NLOS, optical non-line-of-sight imaging technology) is a breakthrough imaging technology that can break through the field-of-view limitation of traditional optical imaging and obtain information about targets outside the line of sight. Light usually travels in a straight line, and traditional optical imaging technologies rely on the condition that the object is within the field of view of the observation device to perform effective imaging. However, for those objects located outside the field of view, blocked, or hidden, traditional imaging methods cannot capture any information about them, while optical non-line-of-sight imaging technology can reconstruct the images of objects hidden outside the line of sight through the reflection signals of the intermediate surface.
[0003] In the related technologies in the field of optical non-line-of-sight imaging technology, pose recognition and modeling are often two independent processes. After pose recognition is completed, manual modeling is required, and there are often large errors in manual modeling. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a method, apparatus, and device for non-line-of-sight human pose modeling based on a deep learning network.
[0005] According to a first aspect of the present disclosure, a method for non-line-of-sight human pose modeling based on a deep learning network is provided. Target non-line-of-sight data is obtained, and the target non-line-of-sight data includes target three-dimensional point cloud data representing the pose of the target object behind the occluder. The target three-dimensional point cloud data is subjected to dimensionality reduction processing to obtain a target point cloud feature map. The point cloud feature map is input into a target pose recognition model to obtain predicted joint data, and the predicted joint data includes predicted joint point data and predicted joint connection data. In the case where the predicted object pose determined by the predicted joint point data and the predicted joint connection data is a first preset pose, a target three-dimensional model is generated based on the predicted joint point data and the predicted joint connection data.
[0006] According to an embodiment of the present disclosure, the above-mentioned predicted joint point data includes a plurality of sub-predicted joint points, and the above-mentioned predicted joint connection data includes a plurality of sub-predicted joint connection data. Generating a target three-dimensional model based on the above-mentioned predicted joint point data and the above-mentioned predicted joint connection data includes: for each sub-predicted joint connection data of a non-root joint connection, determining target root joint connection data based on the above-mentioned predicted joint point data and the above-mentioned predicted joint connection data, wherein the distance between the above-mentioned target root joint connection and the predicted torso is less than the distance between the above-mentioned sub-predicted joint connection data and the above-mentioned predicted torso; determining a root rotation angle based on the sub-predicted joint point corresponding to the above-mentioned target root joint connection and the above-mentioned target root joint connection, and the above-mentioned root rotation angle is the included angle between the above-mentioned target root joint connection data and the initial coordinate axes of the initial three-dimensional model; taking the vector direction of the above-mentioned root rotation angle in the above-mentioned initial coordinate axes as the X-axis direction of the above-mentioned sub-predicted joint connection coordinate axes, and determining the rotation angle of the above-mentioned sub-predicted joint connection based on the above-mentioned predicted joint point data and the above-mentioned predicted joint connection data; adjusting the above-mentioned initial three-dimensional model based on all the above-mentioned target root joint connections and the rotation angles of all the above-mentioned sub-predicted joint connections to obtain the above-mentioned target three-dimensional model.
[0007] According to an embodiment of the present disclosure, the above-mentioned dimensionality reduction processing of the above-mentioned target three-dimensional point cloud data to obtain a target point cloud feature map includes: projecting the above-mentioned target three-dimensional point cloud data to obtain a point cloud depth map and a point cloud intensity map, the above-mentioned point cloud depth map characterizing the depth information of each point in the above-mentioned target three-dimensional point cloud data in space, and the above-mentioned point cloud intensity map characterizing the reflection intensity information of each point in the above-mentioned target three-dimensional point cloud data on the above-mentioned target object; extracting feature data from the above-mentioned point cloud depth map to obtain point cloud depth feature data, and extracting feature data from the above-mentioned point cloud intensity map to obtain point cloud intensity feature data; obtaining the above-mentioned point cloud feature map based on the above-mentioned point cloud depth map, the above-mentioned point cloud intensity map, the above-mentioned point cloud depth feature data, and the above-mentioned point cloud intensity feature data.
[0008] According to an embodiment of the present disclosure, the above-mentioned predicted object pose is determined by the following method:
[0009] Determining a target judgment joint point based on the above-mentioned predicted joint point data and the above-mentioned predicted joint connection data; when the above-mentioned target judgment joint point meets the first preset pose determination condition, determining the above-mentioned predicted object pose as the first preset pose; when the above-mentioned target judgment joint point meets the second preset pose determination condition, determining the above-mentioned predicted object pose as the second preset pose.
[0010] According to an embodiment of the present disclosure, the above method further includes: when the predicted object posture is a second preset posture, performing a rotation process on the predicted joint connection data to obtain the rotated predicted joint connection data; and obtaining the target three-dimensional model based on the predicted joint point data and the rotated predicted joint connection data.
[0011] According to an embodiment of the present disclosure, the above object predicted joint data further includes a sub-predicted joint point confidence corresponding to the sub-predicted joint point and a sub-predicted joint connection strength corresponding to the sub-predicted joint connection data. After inputting the above point cloud feature data into the target posture recognition model to obtain the above target object predicted joint data, the method further includes: determining the sub-predicted joint points with the sub-predicted joint point confidence greater than a preset confidence threshold as target sub-predicted joint points; and determining the sub-predicted joint connection data with the sub-predicted joint connection strength greater than a preset strength threshold as target sub-predicted joint connection data.
[0012] According to an embodiment of the present disclosure, the above target posture recognition model is trained by the following method: obtaining sample non-line-of-sight training data, where the sample non-line-of-sight training data includes sample three-dimensional point cloud data and joint annotation information of a sample object. The three-dimensional point cloud data represents the posture of the sample object behind an occluder, and the three-dimensional point cloud data is obtained by performing a non-line-of-sight simulation on a three-dimensional human model corresponding to the sample object. The three-dimensional human model is obtained by performing posture extraction on at least one frame of an image in a video including the sample object; extracting features from the sample three-dimensional point cloud data to obtain a sample point cloud feature map; inputting the sample point cloud feature map into a preset posture recognition model to obtain sample predicted joint data of the sample object; using a preset loss function to obtain a loss value based on the joint annotation information and the sample predicted joint data; and adjusting the preset posture recognition model based on the loss value to obtain the target posture recognition model.
[0013] According to an embodiment of the present disclosure, the above sample prediction joint data includes sample prediction joint point data, and the above sample prediction joint point data includes a plurality of sub-sample prediction joint point data; the above preset pose recognition model includes a plurality of convolutional calculation modules, and each convolutional calculation module includes a convolutional sub-module, a pooling sub-module, a transposed convolutional sub-module, and an activation sub-module. The above method further includes: for each convolutional calculation module, the convolutional sub-module extracts features from the above sample point cloud feature map based on a preset convolutional weight to obtain sub-point cloud feature data, and the above preset convolutional weight is determined based on a preset weight determination algorithm and the preset number of channels and preset convolutional kernel in the convolutional sub-module; the pooling module downsamples the above sub-point cloud features to obtain downsampled sub-point cloud feature data; the transposed convolutional module upsamples the above downsampled sub-point cloud feature data to generate candidate point cloud feature data; the activation module normalizes the above candidate point cloud feature data to obtain a candidate sample prediction joint map, and the sum of the probabilities that each pixel point in the above candidate sample prediction joint map is determined as a sub-sample prediction joint point is 1.
[0014] A second aspect of the present disclosure provides a non-line-of-sight human pose modeling device based on a deep learning network, including: an acquisition module, configured to acquire target non-line-of-sight data, where the above target non-line-of-sight data includes target three-dimensional point cloud data representing the pose of the above target object behind an occluder; a dimensionality reduction module, configured to perform dimensionality reduction processing on the above target three-dimensional point cloud data to obtain a target point cloud feature map; an input module, configured to input the above point cloud feature map into a target pose recognition model to obtain prediction joint data, where the above prediction joint data includes prediction joint point data and prediction joint connection data; a generation module, configured to generate a target three-dimensional model based on the above prediction joint point data and the above prediction joint connection data when the predicted object pose determined by the above prediction joint point data and the above prediction joint connection data is a first preset pose.
[0015] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0016] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0017] A fifth aspect of the present disclosure further provides a computer program product, including a computer program or instruction, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0018] According to the embodiments of the present disclosure, by performing dimensionality reduction processing on the point cloud features to obtain the target point cloud feature map, the amount of calculation is reduced while the generation of the target three-dimensional model is accelerated. By inputting the point cloud feature map into the target pose recognition model for recognition to determine the predicted joint point data and the predicted joint connection data, and finally by judging the predicted object pose, modeling is carried out specifically according to the predicted joint point data and the predicted joint connection data to obtain the target three-dimensional model, which can effectively reduce the errors and inconsistencies caused by the need for manual modeling based on sample-predicted joint point data during the modeling process, improve the accuracy of modeling, and thus enhance the practicality of the non-line-of-sight technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0020] Figure 1 Schematically shows an application scenario diagram of a non-line-of-sight human pose modeling method and device based on a deep learning network according to an embodiment of the present disclosure;
[0021] FIG. 2 schematically shows a flowchart of a non-line-of-sight human pose modeling method based on a deep learning network according to an embodiment of the present disclosure;
[0022] Figure 3 Schematically shows the three-dimensional model effect and the human pose diagram in the actual scenario in a non-line-of-sight integrated system according to an embodiment of the present disclosure;
[0023] Figure 4 Schematically shows a human T-pose diagram according to an embodiment of the present disclosure;
[0024] Figure 5 Schematically shows a schematic diagram of the network structure of a pose recognition model and joint extraction and modeling according to an embodiment of the present disclosure;
[0025] Figure 6 Schematically shows a block diagram of the structure of a non-line-of-sight human pose modeling device based on a deep learning network according to an embodiment of the present disclosure;
[0026] Figure 7 Schematically shows a block diagram of an electronic device suitable for implementing a non-line-of-sight human pose modeling method based on a deep learning network according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0031] NLOS (Non-Line-of-Sight Imaging, optical non-line-of-sight imaging technology) is a breakthrough imaging technology that can break the field-of-view limitation of traditional optical imaging and obtain target information outside the line of sight. Through the reflected signal of the intermediate surface, the optical non-line-of-sight imaging technology can reconstruct the image of an object hidden outside the line of sight, thus achieving the effect of "imaging through walls" or "perspective". This technology expands the application range of the imaging system, enabling important visual information to be obtained even in complex or enclosed environments.
[0032] In the related art, the imaging result of the optical non-line-of-sight imaging technology is often a three-dimensional point cloud map. However, the data volume of the three-dimensional point cloud map is large, resulting in a large workload for human body pose modeling using the three-dimensional point cloud map, and manual modeling is required after pose recognition, and manual modeling often has large errors.
[0033] In view of this, embodiments of the present disclosure provide a non-line-of-sight human pose modeling method based on a deep learning network, including: obtaining target non-line-of-sight data, where the target non-line-of-sight data includes target three-dimensional point cloud data representing the pose of a target object behind an occluder; performing dimensionality reduction processing on the target three-dimensional point cloud data to obtain a target point cloud feature map; inputting the point cloud feature map into a target pose recognition model to obtain predicted joint data, where the predicted joint data includes predicted joint point data and predicted joint connection data; and when the predicted object pose determined by the predicted joint point data and the predicted joint connection data is a first preset pose, generating a target three-dimensional model based on the predicted joint point data and the predicted joint connection data.
[0034] Figure 1 Schematically shows an application scenario diagram of a non-line-of-sight human pose modeling method and device based on a deep learning network according to an embodiment of the present disclosure.
[0035] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0038] The server 105 may be a server providing various services, such as a background management server (only as an example) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0039] It should be noted that the non-line-of-sight human pose modeling method based on a deep learning network provided in the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the non-line-of-sight human pose modeling device based on a deep learning network provided in the embodiments of the present disclosure can generally be arranged in the server 105. The non-line-of-sight human pose modeling method based on a deep learning network provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the non-line-of-sight human pose modeling device based on a deep learning network provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0040] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0041] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 5 the described scenario, and will describe in detail the non-line-of-sight human pose modeling method based on a deep learning network in the embodiments of the present disclosure through
[0042] Figure 2 FIG. schematically shows a flowchart of a non-line-of-sight human pose modeling method based on a deep learning network according to an embodiment of the present disclosure.
[0043] As Figure 2 shown, the non-line-of-sight human pose modeling method based on a deep learning network in this embodiment includes operations S210 to S240.
[0044] In operation S210, target non-line-of-sight data is acquired.
[0045] Among them, the target non-line-of-sight data includes target three-dimensional point cloud data representing the pose of a target object behind an occluder.
[0046] According to an embodiment of the present disclosure, the dimension of the above three-dimensional point cloud data can be 64*64*512.
[0047] In operation S220, dimensionality reduction processing is performed on the target three-dimensional point cloud data to obtain a target point cloud feature map.
[0048] In operation S230, the point cloud feature map is input into a target pose recognition model to obtain predicted joint data.
[0049] Among them, the predicted joint data includes predicted joint point data and predicted joint connection data.
[0050] According to an embodiment of the present disclosure, the above-mentioned target pose recognition model can be divided into two levels. One level is used to predict PCM (Part confidence maps, joint heatmaps), that is, the predicted joint point data in the present disclosure, and the other level is used to predict PAF (Part Affinity Fields, joint affinity fields), that is, the predicted joint connection data in the present disclosure. After separate processing, the predicted joint point data and predicted joint connection data obtained from the two levels can be spliced and output.
[0051] For the sake of easy understanding, the concepts involved in the embodiments of the present disclosure are briefly explained as follows:
[0052] Joint heatmap: It is a set of heatmaps related to each human body key point (such as shoulders, knees), which is used to represent the confidence that each pixel in the image belongs to a specific joint. Each key point of the human body has a corresponding joint heatmap, and the joint heatmap of each key point is a two-dimensional heatmap, where the value of each pixel represents the probability or confidence that this position belongs to a specific joint.
[0053] Joint affinity field: It is used to represent the connection of human body key points (such as the section of the arm from the shoulder to the elbow). Each pair of limb key point connections in the human body has a corresponding pair of joint affinity fields, where the value of each pixel represents a vector pointing to the connection direction and connection strength between two joints. Also, since the connection of joints is bidirectional, each limb key point connection corresponds to a pair of joint affinity field maps.
[0054] According to an embodiment of the present disclosure, the above-mentioned target three-dimensional model can also perform multi-person pose recognition.
[0055] In operation S240, when the predicted object pose determined by the predicted joint point data and the predicted joint connection data is the first preset pose, a target three-dimensional model is generated based on the predicted joint point data and the predicted joint connection data.
[0056] According to an embodiment of the present disclosure, the above-mentioned target three-dimensional model can be, for example, SMPL (Skinned Multi-Person Linear Model, skinned multi-person linear model). The pose parameters of SMPL can be determined according to the predicted joint point data and the predicted joint connection data. SMPL uses 24 joints, and each joint describes the rotation state through three degrees of freedom. Therefore, the SMPL parameters usually consist of 72 elements, that is, 24 joints and 3 rotation degrees of freedom corresponding to each joint.
[0057] Figure 3 Schematically shows the three-dimensional model effect and the human pose diagram in the actual scene in the non-line-of-sight integration system according to an embodiment of the present disclosure.
[0058] As Figure 3 shown, where the target three-dimensional model is on the left side and the posture of the target object behind the occluder is shown in the right figure. It can be seen that the non-line-of-sight human pose modeling method based on the deep learning network proposed by the present disclosure can accurately map the pose data into the target three-dimensional model, that is, the pose recognition is more accurate. And compared with the traditional model conversion method, during the process of model conversion, the human body form is more real and stable.
[0059] According to an embodiment of the present disclosure, the target point cloud feature map is obtained by reducing the dimension of the point cloud features, which reduces the amount of calculation and accelerates the generation of the target three-dimensional model. By inputting the point cloud feature map into the target pose recognition model for recognition, the predicted joint point data and the predicted joint connection data are determined. Finally, by judging the predicted object posture, modeling is carried out according to the predicted joint point data and the predicted joint connection data in a targeted manner to obtain the target three-dimensional model, which can effectively reduce the errors and inconsistencies caused by the need to manually predict the joint point data based on samples during the modeling process, improve the accuracy of modeling, and thus improve the practicability of the non-line-of-sight technology.
[0060] According to an embodiment of the present disclosure, the above-mentioned predicted joint point data includes multiple sub-predicted joint points, and the predicted joint connection data includes multiple sub-predicted joint connection data. Generating the target three-dimensional model based on the predicted joint point data and the predicted joint connection data includes: for each sub-predicted joint connection data of the non-root joint connection, determining the target root joint connection data based on the predicted joint point data and the predicted joint connection data, where the distance between the target root joint connection and the predicted torso is less than the distance between the sub-predicted joint connection data and the predicted torso; determining the root rotation angle based on the sub-predicted joint point corresponding to the target root joint connection and the target root joint connection, and the root rotation angle is the included angle between the target root joint connection data and the initial coordinate axis of the initial three-dimensional model; taking the vector direction of the root rotation angle in the initial coordinate axis as the X-axis direction of the sub-predicted joint connection coordinate axis, and determining the rotation angle of the sub-predicted joint connection based on the predicted joint point data and the predicted joint connection data; adjusting the initial three-dimensional model based on all the target root joint connections and the rotation angles of all the sub-predicted joint connections to obtain the target three-dimensional model.
[0061] According to an embodiment of the present disclosure, the above-mentioned root joint connection is a relative concept, that is, for two connected sub-predicted joint connection data, the sub-predicted joint connection data closer to the center part of the body is determined as the joint connection data. For example, for the forearm between the wrist and the elbow, the corresponding root joint is the upper arm from the elbow to the shoulder.
[0062] Figure 4 Schematically shows a human T - pose diagram according to an embodiment of the present disclosure.
[0063] As Figure 4 shown, taking the SMPL model as an example, all the pose parameters of SMPL corresponding to the pose shown in the figure are 0. When calculating the rotation angle of the forearm, first, the rotation angle of the upper arm needs to be calculated. Then, the target root joint connection data can be determined based on the predicted elbow point, predicted shoulder point, and upper - arm joint connection data in the predicted joint point data and predicted joint connection data. After that, the vector direction of the rotation angle of the upper arm in the initial coordinate is used as the X - axis direction for calculating the forearm coordinate axis (equivalent to making the X - axis coincide with the vector of the upper arm), and the rotation angle of the forearm is further determined based on the predicted wrist point, predicted elbow point, and forearm joint connection data in the measured joint point data and predicted joint connection data, so as to obtain the pose parameters of SMPL. Furthermore, the initial 3D model is adjusted by the rotation angles of all the target root joint connection data and all the sub - predicted joint connection data, thereby obtaining the target 3D model.
[0064] According to an embodiment of the present disclosure, by determining the target root joint connection data corresponding to the sub - predicted joint connection data for the non - root joint connection data, and then using the rotation angle of the target root joint connection data as the X - axis direction of the sub - predicted joint connection coordinate axis, the rotation angle of the sub - joint is further calculated, gradually refining the pose estimation, avoiding error accumulation caused by complex joint structures, thereby improving the accuracy of the generated target 3D model. And the step - by - step estimation can better adapt to target objects in different poses, and it can efficiently adjust the model pose, making the target 3D model closer to the actual pose of the target object.
[0065] According to an embodiment of the present disclosure, the above - mentioned predicted object pose is determined by the following method: determining a target judgment joint point based on the predicted joint point data and predicted joint connection data; when the target judgment joint point meets the first preset pose determination condition, determining the predicted object pose as the first preset pose; when the target judgment joint point meets the second preset pose determination condition, determining the predicted object pose as the second preset pose.
[0066] According to an embodiment of the present disclosure, the above - mentioned first preset pose can be, for example, as Figure 4 shown, the T - pose, and the second preset pose can be the side T - pose, that is, rotating the T - shaped human body 90 degrees relative to the y - axis of the coordinate axis, which is equivalent to the target human body being side - facing the detector behind the occluding object.
[0067] According to an embodiment of the present disclosure, the above-mentioned target judgment joint point can be, for example, the shoulder point. The above-mentioned first preset judgment condition can be, for example, that the difference between the two shoulder points in the Z-axis direction is less than 4, then it can be considered that the target object is facing the detector directly, and the initial model can be set as a T-shaped model; the above-mentioned second preset judgment condition can be, for example, that the difference between the two shoulder points in the Z-axis direction is greater than 4, then it can be considered that the target object is facing the detector sideways, and the initial model can be set as a side T-shaped model.
[0068] According to an embodiment of the present disclosure, the above method further includes: when the predicted object posture is the second preset posture, performing a rotation process on the predicted joint connection data to obtain the rotated predicted joint connection data; based on the predicted joint point data and the rotated predicted joint connection data, obtaining the target three-dimensional model.
[0069] According to an embodiment of the present disclosure, when the initial posture is a side T-shaped model, first, the predicted joint connection data needs to be rotated, which can be rotating 90 degrees around the y-axis of the initial coordinate axis to fit the initial three-dimensional model, and then subsequent calculations are performed according to the calculation method of the above T-shaped model to obtain the target three-dimensional model.
[0070] According to an embodiment of the present disclosure, the target judgment joint point is determined through the predicted joint point data and the predicted joint connection data, and then the predicted object posture of the target object is determined according to the state of the target judgment joint point, and different processes are performed. Targeted processing can be carried out according to different postures of the target object, thereby improving the adaptability and flexibility of the model, enabling the model to be more widely applied to various scenarios; when the predicted object is in the second preset posture, the predicted joint connection data is rotated, and then the rotated predicted joint connection data and the predicted joint point data are fitted to the initial three-dimensional model to ensure the accuracy of subsequent calculations and improve the accuracy of modeling.
[0071] According to an embodiment of the present disclosure, the above object predicted joint data further includes a sub-predicted joint point confidence corresponding to the sub-predicted joint point and a sub-predicted joint connection strength corresponding to the sub-predicted joint connection data. After inputting the point cloud feature data into the target posture recognition model to obtain the target object predicted joint data, it further includes: determining the sub-predicted joint points with the sub-predicted joint point confidence greater than the preset confidence threshold as the target sub-predicted joint points; and determining the sub-predicted joint connection data with the sub-predicted joint connection strength greater than the preset strength threshold as the target sub-predicted joint connection data.
[0072] According to the embodiments of the present disclosure, the preset confidence threshold and the preset intensity threshold can be selected based on test experience. In the present disclosure, the preset confidence threshold can be, for example, 0.3, and the preset intensity threshold can be, for example, 0.4. The pixel values of the sub-predicted joint points below the above-mentioned preset confidence threshold and the sub-predicted joint connection data below the above-mentioned preset intensity threshold can be set to 0, so as to filter out the sub-predicted joint points with low confidence and the sub-predicted joint connection data with low intensity. At the same time, since the embodiments of the present disclosure can identify the actions of multiple people, if there are joints belonging to different people that are wrongly connected, since the intensity of such joints will be smaller than the intensity value of the correct connection of the joints of the same person, the joint connections can be further screened.
[0073] According to the embodiments of the present disclosure, the above-mentioned dimensionality reduction process of the target three-dimensional point cloud data to obtain the target point cloud feature map includes: projecting the target three-dimensional point cloud data to obtain a point cloud depth map and a point cloud intensity map, where the point cloud depth map represents the depth information of each point in the target three-dimensional point cloud data in space, and the point cloud intensity map represents the reflection intensity information of each point in the target three-dimensional point cloud data on the target object; extracting feature data of the point cloud depth from the point cloud depth map, and extracting feature data of the point cloud intensity from the point cloud intensity map; and obtaining the target point cloud feature map based on the point cloud depth map, the point cloud intensity map, the point cloud depth feature data, and the point cloud intensity feature data.
[0074] According to an embodiment of the present disclosure, in the first step, a two-dimensional depth map and a two-dimensional intensity map are obtained by projecting the target three-dimensional point cloud data. The sizes of the aforementioned depth map and intensity map are both 64*64*1. Then, convolution operations are respectively used on the intensity map and the depth map to further extract features. A convolution kernel with a size of 3x3 is used to perform convolution operations on the intensity map and the depth map respectively, and the ReLU (Rectified Linear Unit) activation function is used to obtain a point cloud depth feature map and a point cloud intensity feature map with an output channel of 100 and a size of 64x64. Then, a 2x2 max pooling operation is performed once to reduce the overall size of the above-mentioned point cloud depth feature map and point cloud intensity feature map to 32x32; in the second step, convolution operations are performed again on the pooled feature map. Similarly, a convolution kernel with a size of 3x3 is used, the output channel is still set to 100, and the ReLU activation function is used. Then, a 2x2 max pooling operation is performed again to reduce the feature map size to 16x16; in the third step, convolution operations are performed again on the pooled feature map. A convolution kernel with a size of 3x3 is used, the number of output channels is 100, and the ReLU activation function is applied. The size of the feature map remains 16x16; in the fourth step, a transposed convolution operation is performed. A convolution kernel with a size of 3x3 is used, the stride is 2, the number of output channels is 100, and the feature map size is upsampled to 32x32 through the transposed convolution; in the fifth step, a transposed convolution operation is performed again. A convolution kernel with the same size of 3x3 is used, the stride is 2, the number of output channels is 100, and the feature map size is upsampled again to 64x64, that is, the point cloud depth feature data and the point cloud intensity feature data are obtained.
[0075] According to an embodiment of the present disclosure, through the operations of the above steps, the features extracted by convolution are aggregated through pooling, and finally the image is restored to its original size through transposed convolution. The output channel is set to 100 so that the network can extract 100 different features in the image for identification. At this time, both the point cloud depth feature data and the point cloud intensity feature data are extracted as feature vectors with a dimension of 64*64*100. Then, they are concatenated according to the channel dimension (which can be set as axis=-1 in the code) to obtain a new feature map. At this time, the number of channels of the feature map will be the sum of the channels of the two images, that is, 200. At the same time, the point cloud depth map and the point cloud intensity map are concatenated along the channel dimension to obtain an original input image with a size of (64, 64, 2). The original input image means that the original depth map and intensity map are combined into an image containing two channels. Finally, the original input image and the concatenated feature map are also concatenated along the channel dimension to obtain a point cloud feature map with a size of (64, 64, 202), which contains the original input image and the point cloud depth feature data and point cloud intensity data after feature extraction.
[0076] According to an embodiment of the present disclosure, by projecting the target three-dimensional point cloud data, a point cloud depth map and a point cloud intensity map are obtained. After feature extraction is performed on each of the two maps, they are concatenated in the channel dimension, thereby avoiding the problem of slow calculation existing in directly using a three-dimensional point cloud map, and at the same time effectively utilizing the relatively good depth resolution of non-line-of-sight imaging. Further, by concatenating the point cloud depth map and the point cloud intensity map with the feature map concatenated by the point cloud depth feature data and the point cloud intensity feature data, the basic information of the original image is retained while the high-level features of the extracted depth and intensity maps are included.
[0077] According to an embodiment of the present disclosure, the above-mentioned target pose recognition model is trained by the following method: obtaining sample non-line-of-sight training data, where the sample non-line-of-sight training data includes sample three-dimensional point cloud data and joint annotation information of a sample object. Among them, the three-dimensional point cloud data represents the pose of the sample object behind an occluder, and the three-dimensional point cloud data is obtained by performing non-line-of-sight simulation on a three-dimensional human model corresponding to the sample object. The three-dimensional human model is obtained by performing pose extraction on at least one frame image in a video containing the sample object; performing feature extraction on the sample three-dimensional point cloud data to obtain a sample point cloud feature map; inputting the sample point cloud feature map into a preset pose recognition model to obtain sample predicted joint data of the sample object; using a preset loss function to obtain a loss value based on the joint annotation information and the sample predicted joint data; and adjusting the preset pose recognition model based on the loss value to obtain the target pose recognition model.
[0078] According to an embodiment of the present disclosure, the above-mentioned sample non-line-of-sight training data may include simulation data. Among them, the simulation data is obtained by relevant staff shooting multiple consecutive actions into a video by a shooting device, and then using a pre-trained video task pose recognition model to identify the pose information of the person in each frame. Each frame of the person is output as a file containing a three-dimensional human model by three-dimensional graphic image software. After obtaining the foregoing file, the pose information of each frame of the person is output according to a script, and the file containing the three-dimensional human model is loaded by a non-line-of-sight imaging engine for non-line-of-sight simulation rendering to obtain a non-line-of-sight imaging three-dimensional point cloud map of each frame and the pose information of the person in each frame (i.e., the joint annotation information in the present disclosure). In this step, data augmentation operations can be performed, that is, the file-person pose information pairs of each frame can be adjusted and enhanced, and operations such as translation, rotation, scaling, and depth change can be performed to increase the network generalization ability.
[0079] According to an embodiment of the present disclosure, the above non-line-of-sight training data of the sample may further include real experimental data. The non-line-of-sight imaging three-dimensional point cloud map captured when relevant personnel are in different postures can be obtained through a non-line-of-sight imaging system. At the same time, the corresponding image is captured by a camera device, and with the image as a reference, the human body posture information is marked on the non-line-of-sight imaging three-dimensional point cloud map to obtain the joint annotation information of the sample object. The above real data can also be subjected to operations such as translation, rotation, scaling, and depth change to enhance generalization.
[0080] According to an embodiment of the present disclosure, the above sample predicted joint data includes sample predicted joint point data, and the sample predicted joint point data includes a plurality of sub-sample predicted joint point data; the preset posture recognition model includes a plurality of convolutional calculation modules, and each convolutional calculation module includes a convolutional sub-module, a pooling sub-module, a transposed convolutional sub-module, and an activation sub-module. The method further includes: for each convolutional calculation module, the convolutional sub-module extracts features from the sample point cloud feature map based on a preset convolutional weight to obtain sub-point cloud feature data, and the preset convolutional weight is determined based on a preset channel number and a preset convolutional kernel in the convolutional sub-module by using a preset weight determination algorithm; the pooling module downsamples the sub-point cloud features to obtain the pooled sub-point cloud feature data; the transposed convolutional module upsamples the pooled sub-point cloud feature data to generate candidate point cloud feature data; the activation module normalizes the candidate point cloud feature data to obtain a candidate sample predicted joint map, and the sum of the probabilities that each pixel point in the candidate sample predicted joint map is determined as a sub-sample predicted joint point is 1.
[0081] Figure 5 Schematically shows a schematic diagram of the network structure of the posture recognition model and joint extraction modeling according to an embodiment of the present disclosure.
[0082] As Figure 5As shown, the target 3D point cloud data is first subjected to feature extraction to obtain a sample point cloud feature map, which includes a sample point cloud depth map, a sample point cloud intensity map, and sample point cloud depth feature data. The convolutional neural network in the figure includes multiple convolutional calculation modules. Each convolutional calculation is divided into two layers. One layer is a PCM image convolutional layer for identifying the joint confidence heat map, and the other layer is a PAF image convolutional layer for identifying the joint link heat map. After separately calculating the PCM image and the PAF image, the PCM image and the PAF image are concatenated in the channel dimension and then input into the next convolutional calculation module. Finally, the sample predicted joint data is the result obtained by the last convolutional calculation module. In the present disclosure, for example, it may include 3 convolutional calculation modules, namely the first convolutional calculation module, the second convolutional calculation module, and the third convolutional calculation module. Each convolutional calculation module includes multiple convolutional sub-modules, a pooling sub-module, a transposed convolutional sub-module, and an activation sub-module. Each convolutional calculation module extracts features through a two-dimensional convolutional layer. The corresponding set number of channels is 100, the convolutional kernel size is 3x3, the activation function is ReLU. Max pooling is performed through the pooling sub-module, and then upsampling is performed through the transposed convolutional sub-module to restore the image to the original size. For the prediction of the target predicted joint points (i.e., PCM), a prediction map with 24 channels will be output through the two-dimensional convolutional layer, and the softmax activation function will be used in the activation module to perform normalization processing in the front dimension to ensure that the sum of the pixel values of the entire candidate target predicted joint image is 1, thus meeting the definition of confidence. For the prediction of the target predicted joint connection data (i.e., PAF), a joint affinity field prediction map with 48 channels will be output through the two-dimensional convolutional layer, and the Sigmoid activation function will be used in the activation module to output values. After obtaining the predicted joint point data and the predicted joint connection data, the 3D human key point coordinate information is screened by preset confidence thresholds and preset intensities, and then converted into a target 3D model through the algorithm provided by the present disclosure.
[0083] According to an embodiment of the present disclosure, since the network prediction tends to optimize the results of the first layer, the above-mentioned first convolutional calculation module, second convolutional calculation module, and third convolutional calculation module can be assigned different loss function weights. The weight of the first convolutional calculation module is 1.5, the weight of the second convolutional calculation module is 1, and the weight of the third convolutional calculation module is 0.5.
[0084] According to an embodiment of the present disclosure, for the PCM image convolutional layer, the cross-entropy loss function can be used as the preset loss function; for the PAF image convolutional layer, the L1 loss function can be used as the loss function.
[0085] According to an embodiment of the present disclosure, the above-mentioned convolutional calculation modules can all initialize the weights through the Xavier initialization layer to avoid overfitting.
[0086] According to an embodiment of the present disclosure, a three-dimensional human model is obtained by performing pose extraction on at least one frame of an image in a video containing a sample object, and non-line-of-sight simulation is performed on the three-dimensional human model to obtain sample three-dimensional point cloud data. A variety of data augmentation techniques are applied to the training process, improving the generalization ability of the model. The target pose recognition model trained by the non-line-of-sight human pose modeling method based on a deep learning network proposed in the present disclosure has an accuracy improvement of more than 30% compared with traditional non-line-of-sight imaging techniques under the same training data conditions. Further, the present disclosure performs pose recognition through a convolutional neural network. Through large-scale training, prior information about the human body can be learned from a large amount of sample data, thereby improving the imaging quality, visualization degree of imaging, and versatility of complex human poses.
[0087] According to an embodiment of the present disclosure, distributed GPUs can be used for calculation during training. Two NVIDIA RTX 4090 GPUs can be used for training. The training contains 40,200 data, and the training time is 1 hour and 22 minutes. The training speed is 20 times faster than using a single NVIDIA RTX 2060 Laptop GPU.
[0088] According to an embodiment of the present disclosure, in the prediction stage, due to the task characteristics of non-line-of-sight imaging, the program provided by the present disclosure needs to be able to run on a CPU, and the demand for a GPU cannot be too high. When making a prediction, only the supporting client application developed in C++ needs to be opened to complete the integrated steps of acquisition-reconstruction-modeling. It is experimentally measured that when running the prediction program on a computer equipped with an Intel Core i5-12400 as the CPU, the prediction time for 64*64*512 data is 0.06285285949707031 seconds.
[0089] Based on the above non-line-of-sight human pose modeling method based on a deep learning network, the present disclosure also provides a non-line-of-sight human pose modeling device based on a deep learning network. The following will be combined with Figure 6 to describe this device in detail.
[0090] Figure 6 A structural block diagram of a non-line-of-sight human pose modeling device based on a deep learning network according to an embodiment of the present disclosure is schematically shown.
[0091] As Figure 6 shown, the non-line-of-sight human pose modeling device 600 of this embodiment includes an acquisition module 610, a dimensionality reduction module 620, an input module 630, and a generation module 640.
[0092] The acquisition module 610 is configured to acquire target non-line-of-sight data, where the target non-line-of-sight data includes target three-dimensional point cloud data representing the pose of a target object behind an occluder. In one embodiment, the acquisition module 610 may be configured to perform the operation S210 described above, which will not be elaborated here.
[0093] The dimensionality reduction module 620 is configured to perform dimensionality reduction processing on the target three-dimensional point cloud data to obtain a target point cloud feature map. In one embodiment, the dimensionality reduction module 620 may be configured to perform the operation S220 described above, which will not be elaborated here.
[0094] The input module 630 is configured to input the point cloud feature map into a target pose recognition model to obtain predicted joint data, where the predicted joint data includes predicted joint point data and predicted joint connection data. In one embodiment, the input module 630 may be configured to perform the operation S230 described above, which will not be elaborated here.
[0095] The generation module 640 is configured to generate a target three-dimensional model based on the predicted joint point data and the predicted joint connection data when the predicted object pose determined by the predicted joint point data and the predicted joint connection data is a first preset pose. In one embodiment, the generation module 640 may be configured to perform the operation S240 described above, which will not be elaborated here.
[0096] According to an embodiment of the present disclosure, any multiple of the acquisition module 610, the dimensionality reduction module 620, the input module 630, and the generation module 640 may be combined and implemented in one module, or any one of them may be split into multiple modules. Or, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 610, the dimensionality reduction module 620, the input module 630, and the generation module 640 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of any several of them. Or, at least one of the acquisition module 610, the dimensionality reduction module 620, the input module 630, and the generation module 640 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0097] Figure 7 A block diagram of an electronic device suitable for implementing a non-line-of-sight human pose modeling method based on a deep learning network according to an embodiment of the present disclosure is schematically shown.
[0098] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0099] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0100] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A driver 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 710 as needed so that a computer program read therefrom can be installed into the storage section 708 as needed.
[0101] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.
[0102] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.
[0103] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the non-line-of-sight human pose modeling method based on a deep learning network provided by the embodiments of the present disclosure.
[0104] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0105] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0106] In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiments of the present disclosure are performed. According to the embodiments of the present disclosure, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0107] According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by connecting through the Internet using an Internet service provider).
[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0109] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0110] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A non-viewing human posture modeling method based on a deep learning network, characterized in that: The method comprises: Acquire target non-viewing area data, wherein the target non-viewing area data includes target three-dimensional point cloud data representing the posture of the target object behind the obstruction; Performing dimensionality reduction processing on the target three-dimensional point cloud data to obtain a target point cloud feature map; Inputting the point cloud feature map into a target posture recognition model to obtain predicted joint data, wherein the predicted joint data includes predicted joint point data and predicted joint connection data; When the predicted object posture determined by the predicted joint point data and the predicted joint connection data is a first preset posture, a target three-dimensional model is generated based on the predicted joint point data and the predicted joint connection data.
2. The method according to claim 1, characterized in that The predicted joint point data includes a plurality of sub-predicted joint points, the predicted joint connection data includes a plurality of sub-predicted joint connection data, and generating a target three-dimensional model based on the predicted joint point data and the predicted joint connection data includes: For each sub-predicted joint connection data of the non-root joint connection data, determining target root joint connection data based on the predicted joint point data and the predicted joint connection data, wherein a distance between the target root joint connection data and the predicted trunk is smaller than a distance between the sub-predicted joint connection data and the predicted trunk; Determine a root rotation angle based on the sub-predicted joint point corresponding to the target root joint connection data and the target root joint connection, wherein the root rotation angle is an angle of the target root joint connection data relative to an initial coordinate axis of an initial three-dimensional model; Taking the vector direction of the root rotation angle in the initial coordinate axis as the X-axis direction of the sub-prediction joint connection coordinate axis, and determining the rotation angle of the sub-prediction joint connection based on the predicted joint point data and the predicted joint connection data; The target three-dimensional model is obtained by adjusting the initial three-dimensional model based on all the target root joint connection data and the rotation angles of all the sub-prediction joint connections.
3. The method according to claim 1, characterized in that The step of performing dimensionality reduction processing on the target three-dimensional point cloud data to obtain a target point cloud feature map includes: Projecting the target three-dimensional point cloud data to obtain a point cloud depth map and a point cloud intensity map, wherein the point cloud depth map represents the depth information of each point in the target three-dimensional point cloud data in space, and the point cloud intensity map represents the reflection intensity information of each point in the target three-dimensional point cloud data on the target object; Performing feature extraction on the point cloud depth map to obtain point cloud depth feature data, and performing feature extraction on the point cloud intensity map to obtain point cloud intensity feature data; The point cloud feature map is obtained based on the point cloud depth map, the point cloud intensity map, the point cloud depth feature data and the point cloud intensity feature data.
4. The method according to claim 1, characterized in that: The predicted object posture is determined by the following method: Determine a target judgment joint point based on the predicted joint point data and the predicted joint connection data; In the case where the target judgment joint point satisfies a first preset posture judgment condition, determining the predicted object posture as the first preset posture; When the target judgment joint point satisfies a second preset posture judgment condition, the predicted object posture is determined as the second preset posture.
5. The method according to claim 4, characterized in that The method further comprises: When the predicted object posture is a second preset posture, rotating the predicted joint connection data to obtain rotated predicted joint connection data; The target three-dimensional model is obtained based on the predicted joint point data and the rotated predicted joint connection data.
6. The method according to claim 2, characterized in that The object prediction joint data also includes a sub-prediction joint point confidence corresponding to the sub-prediction joint point and a sub-prediction joint connection strength corresponding to the sub-prediction joint connection data. After inputting the point cloud feature data into the target posture recognition model to obtain the target object prediction joint data, the method further includes: Determine the sub-prediction joint point whose confidence level is greater than a preset confidence threshold as a target sub-prediction joint point; and The sub-prediction joint connection data whose sub-prediction joint connection strength is greater than a preset strength threshold is determined as the target sub-prediction joint connection data.
7. The method according to claim 1, characterized in that The target posture recognition model is trained by the following method: Acquire sample non-viewing area training data, the sample non-viewing area training data including sample 3D point cloud data and joint annotation information of sample objects, wherein the 3D point cloud data represents the posture of the sample object behind an occluder, the 3D point cloud data is obtained by performing non-viewing area simulation on a 3D character model corresponding to the sample object, and the 3D character model is obtained by performing posture extraction on at least one frame of an image in a video containing the sample object; Extracting features from the sample three-dimensional point cloud data to obtain a sample point cloud feature map; Inputting the sample point cloud feature map into a preset posture recognition model to obtain sample predicted joint data of the sample object; Using a preset loss function, a loss value is obtained based on the joint annotation information and the sample predicted joint data; The preset posture recognition model is adjusted based on the loss value to obtain the target posture recognition model.
8. The method according to claim 7, characterized in that The sample predicted joint data includes sample predicted joint point data, and the sample predicted joint point data includes a plurality of sub-sample predicted joint point data; the preset posture recognition model includes a plurality of convolution calculation modules, each convolution calculation module includes a convolution submodule, a pooling submodule, a convolution transposition submodule and an activation submodule, and the method further includes: For each convolution calculation module, the convolution submodule extracts features from the sample point cloud feature map based on a preset convolution weight to obtain sub-point cloud feature data, wherein the preset convolution weight is determined by a preset weight determination algorithm based on a preset number of channels and a preset convolution kernel in the convolution submodule; The pooling module downsamples the sub-point cloud features to obtain pooled sub-point cloud feature data; The convolution transposition module upsamples the pooled sub-point cloud feature data to generate candidate point cloud feature data; The activation module normalizes the candidate point cloud feature data to obtain a candidate sample prediction joint graph, wherein the sum of probabilities of each pixel point in the candidate sample prediction joint graph being determined as a sub-sample prediction joint point is 1.
9. A non-viewing human posture modeling device based on a deep learning network, characterized in that: The device comprises: An acquisition module, used to acquire target non-viewing area data, wherein the target non-viewing area data includes target three-dimensional point cloud data representing the posture of the target object behind an obstruction; A dimensionality reduction module, used for performing dimensionality reduction processing on the target three-dimensional point cloud data to obtain a target point cloud feature map; An input module, used for inputting the point cloud feature map into a target posture recognition model to obtain predicted joint data, wherein the predicted joint data includes predicted joint point data and predicted joint connection data; A generation module is used to generate a target three-dimensional model based on the predicted joint point data and the predicted joint connection data when the predicted object posture determined by the predicted joint point data and the predicted joint connection data is a first preset posture.
10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Cited By
Vehicle front terrain category identification method based on laser reflection intensity and vision fusion
CN120747903A