Information acquisition method and device, equipment and medium
Through the information prediction model, feature extraction and key point prediction are performed on the target image, and key point positions and related information of multiple body parts of the target object are obtained, which solves the problem that the key point detection model in the prior art is difficult to meet the image processing needs, and achieves more efficient and flexible image processing.
Patent Information
- Application Number
- CN202311541140.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The existing key point detection model is difficult to effectively detect key points in multiple body parts of the target object in image processing, and the detected key points are mainly based on the skeleton structure design, which is difficult to meet the needs of image processing.
Provided is a method of obtaining information, using an information prediction model to extract features of the target image, and obtain key point position information corresponding to multiple body parts of the target object, as well as length information and/or rotation angle information of the body parts.
The information prediction model obtains rich information at one time, including key point position, length information and rotation angle information of body parts, which significantly improves the flexibility and accuracy of image processing and can better meet the needs of image processing.
Smart Images

Figure CN120020878A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, apparatus, device, and medium for information acquisition. Background Art
[0002] In some image processing scenarios, it is necessary to perform beauty processing on target objects such as people included in an image. In related technologies, to perform such processing, it is usually necessary to use an existing key point detection model to detect the key points of the target object, and the key points that can be detected by the existing key point detection model are mainly points designed according to the entire skeleton structure, which is not convenient for subsequent application processing and is difficult to better meet the image processing requirements. Summary of the Invention
[0003] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for information acquisition.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for information acquisition, the method including: obtaining a target image to be processed; where the target image includes a target object; inputting the target image into a preset information prediction model; where the information prediction model includes a feature extraction network, a key point prediction network, and a body part information prediction network; extracting features of the target image through the feature extraction network to obtain image features; based on the image features, obtaining position information of key points corresponding to multiple body parts of the target object through the key point prediction network; based on the image features, obtaining body information of the target object through the body part information prediction network; the body information includes length information of a first body part and / or rotation angle information of a second body part.
[0005] In a second aspect, an embodiment of the present disclosure further provides an information acquisition apparatus, including: an image acquisition module, configured to obtain a target image to be processed; where the target image includes a target object; a model input module, configured to input the target image into a preset information prediction model; where the information prediction model includes a feature extraction network, a key point prediction network, and a body part information prediction network; a feature extraction module, configured to extract features of the target image through the feature extraction network to obtain image features; a key point position acquisition module, configured to obtain position information of key points corresponding to multiple body parts of the target object through the key point prediction network based on the image features; a body information acquisition module, configured to obtain body information of the target object through the body part information prediction network based on the image features; the body information includes length information of a first body part and / or rotation angle information of a second body part.
[0006] In a third aspect, embodiments of the present disclosure further provide an electronic device, which includes: a storage device storing a computer program thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of the information acquisition method provided by the embodiments of the present disclosure.
[0007] In a fourth aspect, embodiments of the present disclosure further provide a computer-readable storage medium storing a computer program, and the computer program is used to execute the information acquisition method provided by the embodiments of the present disclosure.
[0008] The above technical solutions provided by the embodiments of the present disclosure can directly use the feature extraction network in the information prediction model to extract features of the target image, and based on the extracted image features, use the key point prediction network in the information prediction model to obtain the position information of the key points corresponding to multiple body parts of the target object, and use the part information prediction network in the information prediction model to obtain the body information of the target object, where the body information includes the length information of the first body part and / or the rotation angle information of the second body part. The above method can directly obtain relatively rich information such as the position of the key points corresponding to the body parts and the length information and / or rotation angle information of the body parts at one time by means of the information prediction model. Such information is also more conducive to subsequent flexible application processing and can better meet the image processing requirements.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0011] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of an information acquisition method provided by an embodiment of the present disclosure;
[0013] Figure 2 It is a schematic structural diagram of an information prediction model provided by an embodiment of the present disclosure;
[0014] Figure 3Schematic diagram of an information prediction model provided by an embodiment of the present disclosure;
[0015] Figure 4 Schematic diagram of an information prediction model provided by an embodiment of the present disclosure;
[0016] Figure 5 Schematic diagram of an information prediction model provided by an embodiment of the present disclosure;
[0017] Figure 6 Schematic diagram of a prediction unit provided by an embodiment of the present disclosure;
[0018] Figure 7 Schematic diagram of an information prediction model provided by an embodiment of the present disclosure;
[0019] Figure 8 Schematic diagram of an information acquisition device provided by an embodiment of the present disclosure;
[0020] Figure 9 Schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0021] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0022] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0023] Figure 1 Schematic flowchart of an information acquisition method provided by an embodiment of the present disclosure. This method can be executed by an information acquisition device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method mainly includes the following steps S102 to step S110:
[0024] Step S102, obtain a target image to be processed; wherein, the target image includes a target object. The target object may be a person, an animal, etc., which is not limited herein. In addition, the present disclosure does not limit the acquisition manner of the target image either. For example, the target image may be an image collected by the user, or a user image (where the user is the target object) taken with the user's authorization, or an image selected by the user from an image library, etc.
[0025] Step S104: Input the target image into a preset information prediction model. The information prediction model includes a feature extraction network, a key point prediction network, and a body part information prediction network. The information prediction model is a neural network model, and the specific structures of the feature extraction network, the key point prediction network, and the body part information prediction network included in the information prediction model are not limited in the embodiments of the present disclosure.
[0026] Step S106: Extract features from the target image through the feature extraction network to obtain image features. For example, the feature extraction network can adopt a way of layer-by-layer downsampling to extract features from the target image to obtain image features of a required scale.
[0027] Step S108: Based on the image features, obtain the position information of the key points corresponding to multiple body parts of the target object through the key point prediction network.
[0028] Exemplarily, the multiple body parts include the arms and / or legs, and also include one or more of the head, neck, chest, abdomen, and waist. Compared with the related art where key points are obtained only based on the skeleton structure, in the embodiments of the present disclosure, the key points are set and obtained based on multiple body parts of the target object. The multiple body parts not only include the arms and / or legs, but can further include the head, neck, chest, abdomen, waist, etc. By predicting the key points of multiple body parts, it is more convenient for subsequent direct application and ensures the subsequent application effect. For example, when performing special effects processing such as breast augmentation, waist thinning, and swan neck on the target person in the image, if based on the existing human 2D key point protocol designed according to the skeleton structure (only 17 limb key points are set), the limb key points cannot be directly applied, and other points need to be estimated on this basis, which seriously affects the accuracy and efficiency of subsequent image processing.
[0029] Step S110: Based on the image features, obtain the body information of the target object through the body part information prediction network. The body information includes the length information of the first body part and / or the rotation angle information of the second body part.
[0030] In some implementation examples, the first body part includes the part between the shoulders and the waist of the target object, that is, the first body part can be the upper body of the target object, and the second body part includes the shoulders and / or the waist. That is, the information prediction model provided in the embodiments of the present disclosure can not only output the key points corresponding to multiple body parts, but also output the length information of the first body part and / or the rotation angle information of the second body part at the same time, which is convenient for subsequent direct application. For example, it is more convenient and flexible to process the upper body that the user is more concerned about, or to process the rotatable parts such as the shoulders and the waist.
[0031] The above method can directly obtain relatively rich information such as the key point positions corresponding to body parts, the length information of body parts, and / or the rotation angle information, etc. with the help of an information prediction model at one time. Such information is also more conducive to subsequent flexible application processing and can better meet the image processing requirements.
[0032] In some embodiments, the above feature extraction of the target image through the feature extraction network to obtain image features includes: performing multi-scale feature extraction on the target image through the feature extraction network to obtain image features of multiple scales. Exemplarily, the feature extraction network includes multiple downsampling layers, which can perform layer-by-layer downsampling on the target image to obtain multiple features with gradually decreasing scales. For example, performing 2-fold downsampling, 4-fold downsampling, 8-fold downsampling, and 16-fold downsampling on the target image to obtain the corresponding scale features respectively.
[0033] On the basis of the above, the above obtaining the position information of the key points corresponding to multiple body parts of the target object through the key point prediction network based on the image features includes: obtaining the position information of the key points corresponding to multiple body parts of the target object through the key point prediction network based on one or more features among the image features of multiple scales. In practical applications, one or more features among the features of multiple scales can be selected according to needs. It can be understood that the receptive fields of features of different scales are different, and the position information of key points can be predicted more accurately and reliably through the features of multiple scales.
[0034] In some specific implementation examples, when the key point prediction network obtains the position information of the key points corresponding to multiple body parts of the target object, it can be executed with reference to the following steps 1 to 3:
[0035] Step 1, fuse the image features of at least two target scales among multiple scales to obtain a fused feature.
[0036] Exemplarily, the target scale can include the feature with the smallest scale among the features of multiple scales, and one or more features of the intermediate scales among the features of multiple scales. For example, the features obtained by 16-fold downsampling and the features obtained by 8-fold downsampling can be selected for fusion. When specifically fusing, the features of different scales can be unified into the same scale. For example, the features obtained by 16-fold downsampling are upsampled to obtain the same size as the features obtained by 8-fold downsampling, and then fusion operations such as splicing, dot multiplication, addition, and convolution can be performed. The specific feature fusion method adopted in the embodiments of the present disclosure is not limited. The above-mentioned fused feature finally obtained fuses the feature information of different scales and carries more rich and comprehensive information.
[0037] Step 2, obtain the heat map corresponding to the key points of the target part in the target object according to the fused feature.
[0038] Exemplarily, a convolution operation can be performed on the fused feature to obtain heatmaps corresponding to the key points of multiple body parts in the target object. In practical applications, each key point can correspond to a heatmap. Assuming there are a total of 35 key points, 35 heatmaps can be obtained. By using the more comprehensive and richly informative fused feature, the heatmap corresponding to each key point can be obtained more accurately and reliably, which helps to further ensure the accuracy of key point position prediction.
[0039] Step 3, predict the position information of the key points of the target part in the target object based on the heatmap.
[0040] In practical applications, a key point detection algorithm based on the heatmap can be used to determine the position information of the key points of the target part in the target object. For specific details, reference can be made to related technologies and will not be elaborated here.
[0041] In some embodiments, based on the above image features, the body information of the target object is obtained through the part information prediction network, including the following Method 1 and / or Method 2:
[0042] Method 1, based on the features of the smallest scale among multiple scales of image features, obtain the body information of the target object through the part information prediction network. In this method, the part information prediction network is directly related to the feature extraction network.
[0043] Method 2, obtain the fused feature generated by the key point prediction network, and based on the fused feature, obtain the body information of the target object through the part information prediction network; wherein, the fused feature is the feature obtained by the key point prediction network through the first fusion process of the image features of at least two target scales among multiple scales. In this method, the part information prediction network is directly related to the key point prediction network and indirectly related to the feature extraction network.
[0044] For ease of understanding, the embodiments of the present disclosure provide Figures 2 to 4 the structural schematic diagrams of three information prediction models as shown in Figure 2 the structural schematic diagram of an information prediction model as shown, showing that the input of the feature extraction network is the target image, the feature extraction network is directly connected to the key point prediction network and the part information prediction network respectively, one or more scales of image features output by the feature extraction network are used as the input of the key point prediction network, the smallest scale of image features output by the feature extraction network is used as the output of the part information prediction network, the output of the key point prediction network is the position information of the key points of multiple body parts, and the output of the part information prediction network is the length information of the first body part and / or the rotation angle information of the second body part.
[0045] Corresponding to the aforementioned Method 2, seeFigure 3 Schematic structural diagram of an information prediction model shown, as compared with Figure 2 The difference is that the image features of multiple scales output by the feature extraction network are used as the input of the key point prediction network, and the input of the part information prediction network is the fused feature obtained from the key point prediction network. At this time, the part information prediction network is not directly associated with the key point prediction network.
[0046] Corresponding to the foregoing Method 1 and Method 2, refer to Figure 4 Schematic structural diagram of an information prediction model shown, as compared with Figure 2 and Figure 3 The difference is that the input of the part information prediction network includes both the image features of the smallest scale output by the feature extraction network and the fused feature obtained from the key point prediction network. At this time, the part information prediction network is directly associated with both the feature extraction network and the key point prediction network.
[0047] In some specific implementation examples, on the basis of Figure 4 , reference can be made to Figure 5 Schematic structural diagram of an information prediction model shown, which further shows that the specific structure of the part information prediction network includes a first prediction unit, a second prediction unit, and a result fusion unit; the first prediction unit is used to obtain a first prediction result corresponding to the body information of the target object based on the feature of the smallest scale among the image features of multiple scales; the second prediction unit is used to obtain a second prediction result corresponding to the body information of the target object based on the fused feature; the result fusion unit is used to fuse according to the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result.
[0048] In a specific implementation example where the result fusion unit fuses according to the first prediction result and the second prediction result to determine the body information of the target object, weighted fusion processing can be performed based on the first prediction result and the second prediction result to obtain a weighted fusion result; according to the weighted fusion result, the body information of the target object is determined. The embodiments of the present disclosure do not limit the respective weights of the first prediction result and the second prediction result. In some specific implementation manners, the weight corresponding to the first prediction result is not lower than the weight corresponding to the second prediction result. Through research, this weight setting manner can better ensure the reliability of the finally obtained information.
[0049] In some specific implementation examples, the network structures of the first prediction unit and the second prediction unit are the same, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein, the global processing layer is configured to perform global processing on the input information of the network structure to obtain global information; the length information prediction layer is configured to perform first decoupling processing based on the global information to predict the length information of the first body part; the rotation angle prediction layer is configured to perform second decoupling processing based on the global information to predict the rotation angle information of the second body part.
[0050] For ease of understanding, reference may be made to Figure 6 A schematic structural diagram of a prediction unit as shown, which not only shows the global processing layer, the length information prediction layer, and the rotation angle prediction layer, but also shows the specific implementation manners of each layer. Among them, the global processing layer includes a GAP (Global average pooling) network layer and two sequentially connected FC (Full Connection) network layers, and the length information prediction layer and the rotation angle prediction layer are respectively implemented by FC network layers. It should be noted that although the FC network layers are included in both the global processing layer, the length information prediction layer, and the rotation angle prediction layer, the specific parameters may be different and are not limited herein. The above method of decoupling the global information through two different branches and then respectively performing information prediction is more accurate than the method in the related art that only uses one network branch to predict all information.
[0051] For ease of understanding, the embodiments of the present disclosure further provide a Figure 7 Schematic structural diagram of the information prediction model as shown. It shows that the feature extraction network (also called the backbone network) can output 16-fold downsampled features and 8-fold downsampled features. Among them, the 16-fold downsampled features are input to the feature fusion layer together with the 8-fold downsampled features after passing through the upsampling layer. The feature fusion layer can output fused features, and the fused features are input to the heatmap processing layer. The heatmap processing layer is configured to obtain the heatmap corresponding to the key points based on the fused features and output the position information of the key points of multiple body parts based on the heatmap. Figure 7The body part information prediction network therein includes a first prediction unit, a second prediction unit, and a result fusion unit. The input of the first prediction unit is the 16-fold downsampled feature, and the output is the first prediction result corresponding to the body information of the target object. The input of the second prediction unit is the above-mentioned fusion feature, and the output is the second prediction result corresponding to the body information of the target object. The result fusion unit is used to fuse according to the first prediction result and the second prediction result, so as to obtain the length information of the first body part and / or the rotation angle information of the second body part. The above is only an exemplary description. In practical applications, it may include more or fewer networks, and the network layers inside the network can be flexibly adjusted according to needs, which will not be limited here.
[0052] Furthermore, the embodiments of the present disclosure also provide a method for obtaining the information prediction model. Exemplarily, the information prediction model is obtained according to the following steps A to C:
[0053] Step A: Obtain a sample image carrying label information; wherein, the sample image includes a target object, and the label information includes the position labels of the key points corresponding to multiple body parts in the target object and the body information label of the target object. For ease of understanding, in some specific implementation examples, step A can be executed with reference to the following steps A1 to A5:
[0054] Step A1: Obtain a sample image including a target object.
[0055] Step A2: Use a preset network model to obtain multiple body key points corresponding to the target object in the sample image. Exemplarily, the preset network model can be a 3Dmesh model, which can mark 6890 3D key points of the whole body of the target object (such as a human body).
[0056] Step A3: According to the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object, determine the position information of the key points corresponding to the multiple body parts of the target object in the sample image.
[0057] The embodiments of the present disclosure do not directly perform key point annotation on 2D sample images. Instead, a preset network model is used to first obtain multiple body key points of the three-dimensional structure corresponding to the target object in the sample image, and then through the mapping relationship (which can be represented by a regression matrix) between the multiple body key points and the key points corresponding to multiple body parts of the target object, 6890 3D key points can be regressed to the key points corresponding to multiple body parts required by the embodiments of the present disclosure (such as 35 key points). In practical applications, the mapping relationship can be obtained based on the following method: for each key point corresponding to multiple body parts of the target object, the association weights between the key point and each body key point are obtained, and then based on the association weights between the key point and each body key point, the mapping relationship between the key point and multiple body key points is obtained. On the basis of knowing multiple body key points of the three-dimensional structure corresponding to the target object through the preset network model, the key points corresponding to multiple body parts of the target object in the sample image can be efficiently and accurately determined directly through the mapping relationship, greatly saving the manual annotation cost.
[0058] Step A4, determine the body information of the target object in the sample image according to the position information of the key points corresponding to multiple body parts of the target object in the sample image. On the basis of knowing the position information of the key points corresponding to multiple body parts, the required key points can be further obtained therefrom, and the body information of the target object can be determined according to the positional relationship between the key points. Exemplarily, the following steps (1) to (3) can be referred to:
[0059] Step (1), determine at least two first target points corresponding to the first body part and at least two second target points corresponding to the second body part according to the position information of the key points corresponding to multiple body parts of the target object in the sample image.
[0060] In some specific implementation examples, the first body part includes the part between the shoulder and the waist of the target object, which can also be referred to as the upper body part of the target object; the at least two first target points include: the first target point determined based on the center point (which can be simply referred to as the first center point) between the key point corresponding to the left shoulder part and the key point corresponding to the right shoulder part, and the first target point determined based on the center point (which can be simply referred to as the second center point) between the key point corresponding to the left waist part and the key point corresponding to the right waist part. Among them, the first center point is also the shoulder midpoint, and the second center point is also the waist midpoint. The embodiments of the present disclosure can use both the first center point and the second center point as the first target points.
[0061] In some specific implementation examples, the second body part includes the shoulders, and at least two second target points include: the key points corresponding to the left shoulder part and the key points corresponding to the right shoulder part; and / or, the second body part includes the waist, and at least two second target points include: the key points corresponding to the left waist part and the key points corresponding to the right waist part.
[0062] Step (2) determines the length information of the first body part according to at least two first target points. In the case where the first body part includes the part between the shoulders and the waist of the target object, the length information of the first body part, that is, the upper body length information, can be determined according to the distance between the aforementioned first center point and the aforementioned second center point.
[0063] Step (3) determines the rotation angle information of the second body part according to at least two second target points. Exemplarily, the rotation angle information of the shoulders can be characterized by the angles between the line connecting the key points corresponding to the left shoulder part and the key points corresponding to the right shoulder part and the X, Y, and Z axes in the space coordinate system respectively. This angle can be the Euler rotation angle (x1, y1, z1). The rotation angle information of the waist can be characterized by the angles between the line connecting the key points corresponding to the left waist part and the key points corresponding to the right waist part and the X, Y, and Z axes in the space coordinate system respectively. This angle can be the Euler rotation angle (x2, y2, z2).
[0064] Step A5, based on the position information of the key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image, attach label information to the sample image.
[0065] Based on the position information of the key points corresponding to multiple body parts of the target object in the known sample image and the body information of the target object in the sample image, the label information of the sample image can be determined, and the sample image and the label information can be associated, that is, attach label information to the sample image.
[0066] Through the aforementioned steps A1 to A5, a sample image with attached label information can be obtained without manual annotation. Therefore, the cost of obtaining the sample image is relatively low, and a large number of sample images can be obtained according to the needs for model training. Moreover, compared with the label information obtained by manual annotation, the label information obtained by the above method has relatively higher accuracy, and can better avoid common problems such as manual annotation errors. Therefore, the obtained label information is more accurate and reliable, which helps to train a more reliable information prediction model in terms of both the reliability of the sample image and the quantity of the sample images.
[0067] Step B: Obtain the information prediction result output by the preset neural network model for the sample image. The information prediction result includes the position information prediction result of the key points corresponding to multiple body parts of the target object in the sample image and the body information prediction result. The structure of the neural network model is the same as that of the aforementioned information prediction model, and the image processing method is also the same. By adjusting the parameters of the neural network model, an information prediction model that can accurately output the position information of the key points corresponding to multiple body parts and the body information is finally obtained.
[0068] Step C: Train the neural network model based on the label information and the information prediction result to obtain the information prediction model based on the trained neural network model. Specifically, the parameters of the neural network model can be adjusted in the direction of reducing the difference between the label information and the information prediction result until the information prediction result of the neural network model meets the requirements, and the information prediction model is obtained.
[0069] In some specific implementation examples, Step C can be executed according to the following Steps C1 to C3:
[0070] Step C1: Determine the first loss based on the difference between the position label and the position information prediction result. Exemplarily, the first loss can be determined by using a preset first loss function based on the difference between the position label corresponding to the sample image and the position information prediction result. The first loss function in the embodiments of the present disclosure is not limited. For example, it can be the MSE (Mean-Square Error) loss function.
[0071] Step C2: Determine the second loss based on the difference between the body information label and the body information prediction result. Exemplarily, the second loss can be determined by using a preset second loss function based on the difference between the body information label corresponding to the sample image and the body information prediction result. The second loss function in the embodiments of the present disclosure is not limited. For example, it can be the L2 loss function.
[0072] Step C3: Train the neural network model based on the first loss and the second loss to obtain the information prediction model based on the trained neural network model.
[0073] In practical applications, the total loss can be determined based on the first loss and the second loss, and the network parameters in the neural network model can be adjusted based on the total loss until the training stops when the total loss converges to a preset threshold. The trained neural network model is used as the information prediction model. The above method is to train the feature extraction network, the key point prediction network, and the part information prediction network in the neural network model at the same time. In addition, if the structure of the neural network model is Figure 3 or Figure 4Regarding the structure, since the part information prediction network in the neural network model needs to rely on the information output by the key point prediction network, in this case, the feature extraction network and the key point prediction network can also be preferentially trained. Then, the parameters of the feature extraction network and the key point prediction network are fixed, and then the part information prediction network is trained. The specific training method can be flexibly selected and is not limited here.
[0074] In summary, compared with the related art that only outputs the corresponding key points based on the skeleton structure, the above method provided by the embodiments of the present disclosure can directly obtain the key point positions corresponding to multiple body parts, as well as relatively rich information such as the length information and / or rotation angle information of the body parts at one time by means of the information prediction model. Such information is also more conducive to subsequent flexible application processing and can better meet the image processing requirements. Moreover, the training data of the information prediction model provided by the embodiments of the present disclosure does not require manual annotation, which fully guarantees the reliability of the information prediction model in terms of both the quality and quantity of the training samples.
[0075] Corresponding to the foregoing information acquisition method, the embodiments of the present disclosure provide an information acquisition device. Figure 8 FIG. is a schematic structural diagram of an information acquisition device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. By executing the information acquisition method, such as Figure 8 As shown, the information acquisition device includes:
[0076] An image acquisition module 802, configured to acquire a target image to be processed; wherein, the target image includes a target object;
[0077] A model input module 804, configured to input the target image into a preset information prediction model; wherein, the information prediction model includes a feature extraction network, a key point prediction network, and a part information prediction network;
[0078] A feature extraction module 806, configured to perform feature extraction on the target image through the feature extraction network to obtain image features;
[0079] A key point position acquisition module 808, configured to obtain the position information of the key points corresponding to multiple body parts of the target object based on the image features through the key point prediction network;
[0080] A body information acquisition module 810, configured to obtain the body information of the target object based on the image features through the part information prediction network; the body information includes the length information of the first body part and / or the rotation angle information of the second body part.
[0081] The above-mentioned device can directly obtain relatively rich information such as the position of the key points corresponding to the body parts, the length information and / or the rotation angle information of the body parts at one time with the help of the information prediction model. Such information is also more conducive to subsequent flexible application and processing, and can better meet the image processing requirements.
[0082] In some embodiments, the feature extraction module 806 is specifically configured to: perform multi-scale feature extraction on the target image through the feature extraction network to obtain image features of multiple scales.
[0083] In some embodiments, the key point position acquisition module 808 is specifically configured to: based on one or more of the image features of multiple scales, obtain the position information of the key points corresponding to multiple body parts of the target object through the key point prediction network.
[0084] In some embodiments, the key point position acquisition module 808 is specifically configured to: fuse the image features of at least two target scales in the multiple scales to obtain a fused feature; obtain a heat map corresponding to the key points of the target part in the target object according to the fused feature; predict the position information of the key points of the target part in the target object based on the heat map.
[0085] In some embodiments, the body information acquisition module 810 is specifically configured to: based on the feature of the smallest scale among the image features of multiple scales, obtain the body information of the target object through the part information prediction network; and / or obtain the fused feature generated by the key point prediction network, and based on the fused feature, obtain the body information of the target object through the part information prediction network; wherein, the fused feature is the feature obtained by the key point prediction network fusing the image features of at least two target scales in the multiple scales.
[0086] In some embodiments, the part information prediction network includes a first prediction unit, a second prediction unit, and a result fusion unit; the first prediction unit is configured to obtain a first prediction result corresponding to the body information of the target object based on the feature of the smallest scale among the image features of multiple scales; the second prediction unit is configured to obtain a second prediction result corresponding to the body information of the target object based on the fused feature; the result fusion unit is configured to fuse the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result.
[0087] In some embodiments, the result fusion unit is specifically configured to: perform weighted fusion processing on the first prediction result and the second prediction result to obtain a weighted fusion result; determine the body information of the target object according to the weighted fusion result.
[0088] In some embodiments, the weight corresponding to the first prediction result is not lower than the weight corresponding to the second prediction result.
[0089] In some embodiments, the first prediction unit and the second prediction unit have the same network structure, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein, the global processing layer is configured to perform global processing on the input information of the network structure to obtain global information; the length information prediction layer is configured to perform first decoupling processing on the global information to predict the length information of the first body part; and the rotation angle prediction layer is configured to perform second decoupling processing on the global information to predict the rotation angle information of the second body part.
[0090] In some embodiments, the device further includes a model acquisition module, which is configured to obtain the information prediction model according to the following steps: obtain a sample image carrying label information; wherein, the sample image includes a target object, and the label information includes position labels of key points corresponding to multiple body parts in the target object and a body information label of the target object; obtain the information prediction result output by a preset neural network model for the sample image; the information prediction result includes position information prediction results of key points corresponding to multiple body parts of the target object in the sample image and a body information prediction result; and train the neural network model based on the label information and the information prediction result to obtain an information prediction model based on the trained neural network model.
[0091] In some embodiments, the model acquisition module is specifically configured to: obtain a sample image including a target object; use a preset network model to obtain multiple body key points of the three-dimensional structure corresponding to the target object in the sample image; determine the position information of key points corresponding to multiple body parts of the target object in the sample image according to the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object; determine the body information of the target object in the sample image according to the position information of key points corresponding to multiple body parts of the target object in the sample image; and attach label information to the sample image based on the position information of key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image.
[0092] In some embodiments, the model acquisition module is specifically configured to: determine at least two first target points corresponding to a first body part and at least two second target points corresponding to a second body part according to the position information of key points corresponding to multiple body parts of the target object in the sample image; determine the length information of the first body part according to the at least two first target points; and determine the rotation angle information of the second body part according to the at least two second target points.
[0093] In some embodiments, the first body part includes the part between the shoulder and the waist of the target object; the at least two first target points include: a first target point determined based on the center point between the key point corresponding to the left shoulder part and the key point corresponding to the right shoulder part, and a first target point determined based on the center point between the key point corresponding to the left waist part and the key point corresponding to the right waist part.
[0094] In some embodiments, the second body part includes the shoulder, and the at least two second target points include: the key point corresponding to the left shoulder part and the key point corresponding to the right shoulder part; and / or, the second body part includes the waist, and the at least two second target points include: the key point corresponding to the left waist part and the key point corresponding to the right waist part.
[0095] In some embodiments, the model acquisition module is specifically configured to: determine a first loss based on the difference between the position label and the position information prediction result; determine a second loss based on the difference between the body information label and the body information prediction result; and train the neural network model based on the first loss and the second loss to obtain an information prediction model based on the trained neural network model.
[0096] In some embodiments, the multiple body parts include the arm and / or the leg, and also include one or more of the head, neck, chest, abdomen, and waist.
[0097] The information acquisition device provided by the embodiments of the present disclosure can execute the information acquisition method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0098] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device embodiment described above can refer to the corresponding process in the method embodiment, which will not be repeated here.
[0099] Embodiments of the present disclosure provide an electronic device, which includes: a storage device storing a computer program thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure. Embodiments of the present disclosure provide an electronic device, which includes: a storage device storing a computer program thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.
[0100] Reference is made below Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of embodiments of the present disclosure.
[0101] As Figure 9 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0102] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.
[0103] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, the above functions defined in the method of the embodiment of the present disclosure are performed.
[0104] In addition to the above methods and devices, an embodiment of the present disclosure can also be a computer program product that includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the image processing method provided by the embodiment of the present disclosure. The computer program product can be written in any combination of one or more programming languages for performing the program code of the operations of the embodiment of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0105] Furthermore, an embodiment of the present disclosure can also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the information acquisition method provided by the embodiment of the present disclosure.
[0106] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] The embodiment of the present disclosure also provides a computer program product that includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the information acquisition method in the embodiment of the present disclosure is implemented.
[0108] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0109] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0110] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0111] It is understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0112] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0113] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for obtaining information, characterized in that: include: Acquire a target image to be processed; wherein the target image contains a target object; Inputting the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network; Extracting features from the target image through the feature extraction network to obtain image features; Based on the image features, obtaining position information of key points corresponding to multiple body parts of the target object through the key point prediction network; Based on the image features, body information of the target object is acquired through the body information prediction network; the body information includes length information of a first body part and / or rotation angle information of a second body part.
2. The method according to claim 1, characterized in that The step of extracting features from the target image through the feature extraction network to obtain image features includes: Multi-scale feature extraction is performed on the target image through the feature extraction network to obtain image features of multiple scales.
3. The method according to claim 2, characterized in that The step of obtaining the position information of the key points corresponding to the multiple body parts of the target object through the key point prediction network based on the image features includes: Based on one or more features of the image features at multiple scales, position information of key points corresponding to multiple body parts of the target object is obtained through the key point prediction network.
4. The method according to claim 3, characterized in that The step of obtaining the position information of key points corresponding to the plurality of body parts of the target object includes: Fusing image features of at least two target scales among the multiple scales to obtain fused features; Acquire a thermal map corresponding to a key point of a target part in the target object according to the fusion feature; The position information of the key points of the target part in the target object is predicted based on the heat map.
5. The method according to claim 2, characterized in that: The acquiring the body information of the target object through the part information prediction network based on the image features includes: Based on the smallest-scale feature among the image features of the multiple scales, obtaining the body information of the target object through the part information prediction network; and / or, Acquire the fusion feature generated by the key point prediction network, and based on the fusion feature, acquire the body information of the target object through the part information prediction network; wherein the fusion feature is a feature obtained by the key point prediction network by fusing image features of at least two target scales among the multiple scales.
6. The method according to claim 5, characterized in that The part information prediction network includes a first prediction unit, a second prediction unit and a result fusion unit; The first prediction unit is used to obtain a first prediction result corresponding to the body information of the target object based on a feature of a minimum scale among the image features of the multiple scales; The second prediction unit is used to obtain a second prediction result corresponding to the body information of the target object based on the fusion feature; The result fusion unit is used to fuse the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result.
7. The method according to claim 6, characterized in that The fusing the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result includes: Performing weighted fusion processing based on the first prediction result and the second prediction result to obtain a weighted fusion result; The body information of the target object is determined according to the weighted fusion result.
8. The method according to claim 7, characterized in that The weight corresponding to the first prediction result is not less than the weight corresponding to the second prediction result.
9. The method according to claim 7, characterized in that: The network structure of the first prediction unit and the second prediction unit is the same, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein, The global processing layer is used to perform global processing based on the input information of the network structure to obtain global information; The length information prediction layer is used to perform a first decoupling process based on the global information to predict the length information of the first body part; The rotation angle prediction layer is used to perform a second decoupling process based on the global information to predict the rotation angle information of the second body part.
10. The method according to any one of claims 1 to 9, characterized in that: The information prediction model is obtained according to the following steps: Acquire a sample image carrying label information; wherein the sample image includes a target object, and the label information includes position labels of key points corresponding to multiple body parts of the target object and a body information label of the target object; Obtaining information prediction results output by a preset neural network model for the sample image; the information prediction results include position information prediction results and body information prediction results of key points corresponding to multiple body parts of the target object in the sample image; Based on the label information and the information prediction result, the neural network model is trained to obtain an information prediction model based on the trained neural network model.
11. The method according to claim 10, characterized in that The obtaining of a sample image carrying label information includes: Acquire a sample image containing a target object; Using a preset network model to obtain a plurality of body key points of a three-dimensional structure corresponding to a target object in the sample image; Determine position information of the key points corresponding to the multiple body parts of the target object in the sample image according to the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object; Determining body information of the target object in the sample image according to position information of key points corresponding to multiple body parts of the target object in the sample image; Based on the position information of key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image, label information is attached to the sample image.
12. The method according to claim 11, characterized in that The determining the body information of the target object in the sample image according to the position information of the key points corresponding to the multiple body parts of the target object in the sample image comprises: Determine at least two first target points corresponding to a first body part and at least two second target points corresponding to a second body part according to position information of key points corresponding to multiple body parts of the target object in the sample image; Determining length information of the first body part according to the at least two first target points; The rotation angle information of the second body part is determined according to the at least two second target points.
13. The method according to claim 12, characterized in that The first body part includes the area between the shoulder and waist of the target subject; The at least two first target points include: a first target point determined based on a center point between a key point corresponding to a left shoulder and a key point corresponding to a right shoulder, and a first target point determined based on a center point between a key point corresponding to a left waist and a key point corresponding to a right waist.
14. The method according to claim 12, characterized in that The second body part includes a shoulder, and the at least two second target points include: a key point corresponding to a left shoulder part and a key point corresponding to a right shoulder part; and / or, The second body part includes a waist, and the at least two second target points include: a key point corresponding to a left waist part and a key point corresponding to a right waist part.
15. The method according to claim 10, characterized in that The step of training the neural network model based on the label information and the information prediction result to obtain an information prediction model based on the trained neural network model includes: Determining a first loss based on a difference between the location tag and the location information prediction result; determining a second loss based on a difference between the body information tag and the body information prediction result; Based on the first loss and the second loss, the neural network model is trained to obtain an information prediction model based on the trained neural network model.
16. The method according to claim 1, characterized in that The multiple body parts include arms and / or legs, and also include one or more of the head, neck, chest, abdomen and waist.
17. An information acquisition device, characterized in that: include: An image acquisition module, used to acquire a target image to be processed; wherein the target image contains a target object; A model input module, used to input the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network; A feature extraction module, used to extract features of the target image through the feature extraction network to obtain image features; A key point position acquisition module, used to acquire the position information of key points corresponding to multiple body parts of the target object through the key point prediction network based on the image features; A body information acquisition module is used to acquire the body information of the target object through the part information prediction network based on the image features; the body information includes length information of a first body part and / or rotation angle information of a second body part.
18. An electronic device, characterized in that: The electronic device comprises: a storage device having a computer program stored thereon; A processing device, used to execute the computer program in the storage device to implement the steps of the information acquisition method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the information acquisition method described in any one of claims 1 to 16.