Information acquisition method and apparatus, device and medium
Through the information prediction model, feature extraction and key point prediction of the target image is solved, and the problem of difficulty in detecting key points of multiple body parts of the target object in the prior art is solved, thereby achieving rich body information and meeting image processing needs.
Patent Information
- Application Number
- PCT/CN2024/130274
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-22
AI Technical Summary
The existing key point detection model is difficult to effectively detect key points in multiple body parts of the target object in image processing, and the key points of detection are mainly based on the skeleton structure design, which is difficult to meet the needs of image processing.
An information acquisition method is adopted to extract features of the target image through an information prediction model, and obtain key point position information and body information corresponding to multiple body parts of the target object, including length information and rotation angle information.
It realizes the acquisition of key points and related body information of multiple body parts of the target object at one time, providing richer image processing information, which can better meet image processing needs.
Smart Images

Figure CN2024130274_22052025_PF_FP_ABST
Abstract
Description
Information acquisition method, device, equipment and medium
[0001] This application claims priority to the Chinese invention patent application entitled “Information Acquisition Method, Device, Equipment and Medium” filed on November 17, 2023, with application number 202311541140.5. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of computer technology, and in particular to an information acquisition method, apparatus, device, and medium. Background Art
[0003] In some image processing scenarios, it's necessary to perform body enhancement or other processing on target objects, such as people, contained in an image. In related technologies, this type of processing typically requires using existing keypoint detection models to detect the keypoints of the target object. However, these existing keypoint detection models can only detect keypoints based on the entire skeleton structure, making them inconvenient for subsequent application processing and failing to meet image processing requirements.
[0004] Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an information acquisition method, apparatus, device and medium.
[0006] In a first aspect, an embodiment of the present disclosure provides an information acquisition method, the method comprising: acquiring a target image to be processed; wherein the target image contains a target object; inputting the target image into a preset information prediction model; wherein the information prediction model comprises a feature extraction network, a key point prediction network and a part information prediction network; performing feature extraction on the target image through the feature extraction network to obtain image features; based on the image features, acquiring position information of key points corresponding to multiple body parts of the target object through the key point prediction network; based on the image features, acquiring body information of the target object through the part information prediction network; the body information comprises length information of a first body part and / or rotation angle information of a second body part.
[0007] In the second aspect, the embodiments of the present disclosure also provide an information acquisition device, including: an image acquisition module, used to acquire a target image to be processed; wherein the target image contains a target object; a model input module, used to input the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network; a feature extraction module, used to extract features of the target image through the feature extraction network to obtain image features; a key point position acquisition module, used to acquire the position information of key points corresponding to multiple body parts of the target object through the key point prediction network based on the image features; a body information acquisition module, used to acquire body information of the target object through the part information prediction network based on the image features; the body information includes length information of a first body part and / or rotation angle information of a second body part.
[0008] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of the information acquisition method provided in an embodiment of the present disclosure.
[0009] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the information acquisition method provided by the embodiment of the present disclosure.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0012] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] FIG1 is a flow chart of an information acquisition method provided by an embodiment of the present disclosure;
[0014] FIG2 is a schematic diagram of the structure of an information prediction model provided by an embodiment of the present disclosure;
[0015] FIG3 is a schematic diagram of the structure of an information prediction model provided by an embodiment of the present disclosure;
[0016] FIG4 is a schematic diagram of the structure of an information prediction model provided by an embodiment of the present disclosure;
[0017] FIG5 is a schematic diagram of the structure of an information prediction model provided by an embodiment of the present disclosure;
[0018] FIG6 is a schematic diagram of the structure of a prediction unit provided by an embodiment of the present disclosure;
[0019] FIG7 is a schematic diagram of the structure of an information prediction model provided by an embodiment of the present disclosure;
[0020] FIG8 is a schematic structural diagram of an information acquisition device provided by an embodiment of the present disclosure;
[0021] FIG9 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0024] The above technical solution provided by the embodiment of the present disclosure can directly use the feature extraction network in the information prediction model to extract features of the target image, and based on the extracted image features, use the key point prediction network in the information prediction model to obtain the position information of the key points corresponding to multiple body parts of the target object, and use the part information prediction network in the information prediction model to obtain the body information of the target object, and the body information includes the length information of the first body part and / or the rotation angle information of the second body part. The above method can directly use the information prediction model to obtain relatively rich information such as the key point position corresponding to the body part and the length information and / or rotation angle information of the body part at one time. Such information is also more conducive to subsequent flexible application processing and can better meet the image processing needs.
[0025] Figure 1 is a flow chart of an information acquisition method provided by an embodiment of the present disclosure. The method can be executed by an information acquisition device, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. As shown in Figure 1, the method mainly includes the following steps S102 to S110:
[0026] Step S102: Acquire a target image to be processed; the target image contains a target object. The target object can be a person, an animal, or the like, without limitation. Furthermore, the disclosed embodiments do not limit the method for acquiring the target image. For example, the target image can be an image captured by the user, an image of the user captured with the user's authorization (in which case the user is the target object), or an image selected by the user from an image library.
[0027] In step S104, the target image is input into a preset information prediction model. The information prediction model includes a feature extraction network, a keypoint prediction network, and a part information prediction network. The information prediction model is a neural network model, and the disclosed embodiments do not limit the specific structures of the feature extraction network, keypoint prediction network, and part information prediction network included in the information prediction model.
[0028] Step S106: extract features from the target image using a feature extraction network to obtain image features. For example, the feature extraction network can extract features from the target image using a layer-by-layer downsampling method to obtain image features of a desired scale.
[0029] Step S108: Based on the image features, position information of key points corresponding to multiple body parts of the target object is obtained through a key point prediction network.
[0030] Exemplarily, the multiple body parts include arms and / or legs, and also include one or more of the head, neck, chest, abdomen and waist. Compared to the related art in which key points are obtained only based on the skeleton structure, the embodiment of the present disclosure sets and obtains corresponding key points based on multiple body parts of the target object. The multiple body parts include not only arms and / or legs, but may further include the head, neck, chest, abdomen and waist, etc. By predicting the key points of multiple body parts, it is more convenient to directly apply them in the future and ensure the effect of subsequent applications. For example, if the target person in the image is subjected to special effects such as breast enhancement, waist thinning, and swan neck, if it is based on the existing human 2D key point protocol designed according to the skeleton structure (only 17 limb key points are set), the limb key points cannot be directly applied, and it is necessary to estimate through other points on this basis, which seriously affects the accuracy and efficiency of subsequent image processing.
[0031] Step S110: Based on the image features, the body information of the target object is obtained through a body information prediction network; the body information includes length information of a first body part and / or rotation angle information of a second body part.
[0032] In some implementation examples, the first body part includes the area between the shoulder and waist of the target object, that is, the first body part can be the upper body of the target object, and the second body part includes the shoulder and / or waist. In other words, the information prediction model provided by the embodiments of the present disclosure can not only output key points corresponding to multiple body parts, but also simultaneously output the length information of the first body part and / or the rotation angle information of the second body part, which is convenient for subsequent direct application, such as processing the upper body that the user is more concerned about, or processing rotatable parts such as the shoulder and waist, which is more convenient and flexible.
[0033] The above method can directly use the information prediction model to obtain relatively rich information such as the key point positions corresponding to body parts and the length information and / or rotation angle information of the body parts at one time. Such information is also more conducive to subsequent flexible application processing and can better meet image processing needs.
[0034] In some embodiments, extracting features from a target image using a feature extraction network to obtain image features includes performing multi-scale feature extraction on the target image using the feature extraction network to obtain image features at multiple scales. Exemplarily, the feature extraction network includes multiple downsampling layers that can downsample the target image layer by layer to obtain multiple features at decreasing scales. For example, the target image can be downsampled by a factor of 2, 4, 8, and 16 to obtain features at corresponding scales.
[0035] Based on the foregoing, the above-mentioned method of obtaining the position information of key points corresponding to multiple body parts of the target object through a key point prediction network based on image features includes: obtaining the position information of key points corresponding to multiple body parts of the target object through a key point prediction network based on one or more features from image features at multiple scales. In actual applications, one or more features from features at multiple scales can be selected as needed. It is understandable that features at different scales have different receptive fields, and using features at multiple scales can more accurately and reliably predict the position information of key points.
[0036] In some specific implementation examples, when the key point prediction network obtains the position information of key points corresponding to multiple body parts of the target object, it can refer to the following steps 1 to 3:
[0037] Step 1: fuse the image features of at least two target scales among multiple scales to obtain fused features.
[0038] Exemplarily, the target scale may include features of the smallest scale among features of multiple scales, and one or more features of the intermediate scale among features of multiple scales. For example, features obtained by 16-fold downsampling and features obtained by 8-fold downsampling may be selected for fusion. During the specific fusion, features of different scales may be unified into the same scale, such as upsampling the features obtained by 16-fold downsampling to obtain a size consistent with that of the features obtained by 8-fold downsampling, and then fusion operations such as splicing, dot multiplication, addition, and convolution may be performed. The specific feature fusion method adopted in the embodiment of the present disclosure is not limited. The above-mentioned fused features finally obtained fuse feature information of different scales, and the information carried is richer and more comprehensive.
[0039] Step 2: Obtain the heat map corresponding to the key points of the target part in the target object based on the fusion features.
[0040] For example, a convolution operation can be performed on the fused features to generate heatmaps corresponding to key points on multiple body parts of the target object. In practice, each key point can correspond to a heatmap. Assuming there are 35 key points in total, 35 heatmaps can be generated. Utilizing fused features that carry richer and more comprehensive information allows for more accurate and reliable heatmaps to be obtained for each key point, helping to further ensure the accuracy of key point position prediction.
[0041] Step 3: Predict the location information of the key points of the target part in the target object based on the heat map.
[0042] In practical applications, a key point detection algorithm based on a heat map can be used to determine the position information of the key points of the target part of the target object. For details, please refer to relevant technologies and will not be described in detail here.
[0043] In some embodiments, the above-mentioned method of obtaining the target object's body information through a part information prediction network based on image features includes the following method 1 and / or method 2:
[0044] Method 1: Based on the smallest-scale features among the multi-scale image features, the body information of the target object is obtained through the part information prediction network. In this method, the part information prediction network is directly related to the feature extraction network.
[0045] Method 2: Receive fused features generated by the keypoint prediction network and, based on these fused features, use the part information prediction network to obtain the target object's body information. The fused features are features generated by the keypoint prediction network by first fusing image features at at least two target scales from multiple scales. In this method, the part information prediction network is directly related to the keypoint prediction network and indirectly related to the feature extraction network.
[0046] For ease of understanding, the embodiments of the present disclosure provide structural schematic diagrams of three information prediction models as shown in Figures 2 to 4. Specifically, corresponding to the aforementioned method one, refer to the structural schematic diagram of an information prediction model shown in Figure 2, which illustrates that the input of the feature extraction network is the target image, and the feature extraction network is directly connected to the key point prediction network and the part information prediction network respectively. The image features of one or more scales output by the feature extraction network serve as the input of the key point prediction network, and the image features of the minimum scale output by the feature extraction network serve as the output of the part information prediction network. The output of the key point prediction network is the position information of the key points of multiple body parts, and the output of the part information prediction network is the length information of the first body part and / or the rotation angle information of the second body part.
[0047] Corresponding to the aforementioned method 2, refer to the structural diagram of an information prediction model shown in Figure 3. The difference from Figure 2 is that the image features of multiple scales output by the feature extraction network are used as the input of the key point prediction network, and the input of the part information prediction network is the fusion features obtained from the key point prediction network. At this time, the part information prediction network and the key point prediction network are not directly related.
[0048] Corresponding to the aforementioned methods 1 and 2, refer to the structural diagram of an information prediction model shown in Figure 4. The difference from Figures 2 and 3 is that the input of the part information prediction network includes both the minimum-scale image features output by the feature extraction network and the fusion features obtained from the key point prediction network. At this time, the part information prediction network is directly associated with the feature extraction network and the key point prediction network.
[0049] In some specific implementation examples, based on Figure 4, a structural diagram of an information prediction model shown in Figure 5 can be referred to, which further illustrates that the specific structure of the part information prediction network includes a first prediction unit, a second prediction unit and a result fusion unit; the first prediction unit is used to obtain a first prediction result corresponding to the body information of the target object based on the minimum scale feature among the image features of multiple scales; the second prediction unit is used to obtain a second prediction result corresponding to the body information of the target object based on the fused feature; the result fusion unit is used to fuse the first prediction result and the second prediction result to determine the body information of the target object based on the fusion result.
[0050] In a specific implementation example of fusing the first prediction result and the second prediction result to determine the target object's physical information based on the fusion result, the result fusion unit can perform a weighted fusion process based on the first prediction result and the second prediction result to obtain a weighted fusion result; and determine the target object's physical information based on the weighted fusion result. The disclosed embodiments do not restrict the respective weights of the first prediction result and the second prediction result. In some specific implementations, the weight corresponding to the first prediction result is not less than the weight corresponding to the second prediction result. Research has shown that this weighting method can better ensure the reliability of the ultimately obtained information.
[0051] In some specific implementation examples, the network structure of the first prediction unit and the second prediction unit is the same, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein the global processing layer is used to perform global processing based on the input information of the network structure to obtain global information; the length information prediction layer is used to perform a first decoupling processing based on the global information to predict the length information of the first body part; the rotation angle prediction layer is used to perform a second decoupling processing based on the global information to predict the rotation angle information of the second body part.
[0052] For ease of understanding, you can refer to the structural diagram of a prediction unit shown in Figure 6, which not only illustrates the global processing layer, the length information prediction layer and the rotation angle prediction layer, but also illustrates the specific implementation method of each layer, wherein the global processing layer includes a GAP (Global average pooling) network layer and two sequentially connected FC (Full Connection) network layers, and the length information prediction layer and the rotation angle prediction layer are respectively implemented using FC network layers. It should be noted that although the global processing layer, the length information prediction layer and the rotation angle prediction layer all contain FC network layers, the specific parameters can be different and are not limited here. The above-mentioned method of using two different branches to decouple the global information and thus perform information prediction separately, compared with the method of using only one network branch to predict all information in the related art, the final result is more accurate.
[0053] For ease of understanding, the embodiment of the present disclosure further provides a structural schematic diagram of an information prediction model as shown in FIG7 . It illustrates that the feature extraction network (also referred to as the backbone network) can output 16x downsampled features and 8x downsampled features, wherein the 16x downsampled features are input together with the 8x downsampled features after passing through the upsampling layer into the feature fusion layer. The feature fusion layer can output fused features, which are input into the heat map processing layer. The heat map processing layer is used to obtain a heat map corresponding to the key points based on the fused features, and output the position information of the key points of multiple body parts based on the heat map. The part information prediction network in FIG7 includes a first prediction unit, a second prediction unit, and a result fusion unit. The first prediction unit takes the 16x downsampled features as input and outputs a first prediction result corresponding to the body information of the target object. The second prediction unit takes the above-mentioned fused features as input and outputs a second prediction result corresponding to the body information of the target object. The result fusion unit is used to fuse the first prediction result and the second prediction result to obtain the length information of the first body part and / or the rotation angle information of the second body part. The above description is merely an example. In actual applications, more or fewer networks may be included, and the network layers within the network may be flexibly adjusted as needed, which is not limited here.
[0054] Furthermore, the embodiments of the present disclosure also provide a method for obtaining an information prediction model. Exemplarily, the information prediction model is obtained according to the following steps A to C:
[0055] Step A: Acquire a sample image carrying label information; wherein the sample image includes a target object, and the label information includes location labels of key points corresponding to multiple body parts of the target object and a body information label of the target object. For ease of understanding, in some specific implementation examples, Step A can be performed with reference to the following Steps A1 to A5:
[0056] Step A1: Acquire a sample image containing a target object.
[0057] In step A2, a preset network model is used to obtain multiple body key points of the three-dimensional structure corresponding to the target object in the sample image. For example, the preset network model can be a 3D mesh model that can mark 6890 3D key points of the target object (such as a human body).
[0058] Step A3: determining position information of the key points corresponding to the multiple body parts of the target object in the sample image based on the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object.
[0059] The embodiment of the present disclosure does not directly perform key point annotation on the 2D sample image, but uses a preset network model to first obtain multiple body key points of the three-dimensional structure corresponding to the target object in the sample image, and then uses the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object (which can be represented by a regression matrix) to regress the 6890 3D key points to the key points corresponding to the multiple body parts required by the embodiment of the present disclosure (such as 35 key points). In actual applications, the mapping relationship can be obtained based on the following method: for each key point corresponding to the multiple body parts of the target object, the association weight of the key point with each body key point is obtained, and then the mapping relationship between the key point and the multiple body key points is obtained based on the association weight of the key point with each body key point. On the basis of the multiple body key points of the three-dimensional structure corresponding to the target object known through the preset network model, the key points corresponding to the multiple body parts of the target object in the sample image can be efficiently and accurately determined directly through the mapping relationship, which greatly saves the cost of manual annotation.
[0060] Step A4, based on the position information of the key points corresponding to the multiple body parts of the target object in the sample image, the body information of the target object in the sample image is determined. Based on the known position information of the key points corresponding to the multiple body parts, the required key points can be further obtained from the key points, and the body information of the target object can be determined based on the positional relationship between the key points. For example, the following steps (1) to (3) can be referred to:
[0061] Step (1) determines at least two first target points corresponding to a first body part and at least two second target points corresponding to a second body part based on position information of key points corresponding to multiple body parts of the target object in the sample image.
[0062] In some specific implementation examples, the first body part includes the area between the shoulder and waist of the target subject, which can also be referred to as the upper body of the target subject; the at least two first target points include: a first target point determined based on the center point between the key points corresponding to the left shoulder and the key points corresponding to the right shoulder (hereinafter referred to as the first center point), and a first target point determined based on the center point between the key points corresponding to the left waist and the key points corresponding to the right waist (hereinafter referred to as the second center point). The first center point is also the midpoint of the shoulder, and the second center point is also the midpoint of the waist. In the disclosed embodiments, both the first center point and the second center point can be used as the first target points.
[0063] In some specific implementation examples, the second body part includes a shoulder, and the at least two second target points include: a key point corresponding to the left shoulder and a key point corresponding to the right shoulder; and / or, the second body part includes a waist, and the at least two second target points include: a key point corresponding to the left waist and a key point corresponding to the right waist.
[0064] Step (2) determines the length information of the first body part based on the at least two first target points. In the case where the first body part includes the area between the shoulder and waist of the target object, the length information of the first body part, i.e., the length information of the upper body, can be determined based on the distance between the first center point and the second center point.
[0065] Step (3) determines the rotation angle information of the second body part based on the at least two second target points. For example, the rotation angle information of the shoulder can be represented by the angles between the connecting lines corresponding to the key points of the left shoulder and the key points corresponding to the right shoulder and the X, Y, and Z axes in the spatial coordinate system, respectively. The angles can be Euler rotation angles (x1, y1, z1). The rotation angle information of the waist can be represented by the angles between the connecting lines corresponding to the key points of the left waist and the key points corresponding to the right waist and the X, Y, and Z axes in the spatial coordinate system, respectively. The angles can be Euler rotation angles (x2, y2, z2).
[0066] Step A5: Based on the position information of key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image, label information is added to the sample image.
[0067] Based on the known position information of key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image, the label information of the sample image can be determined, and the sample image can be associated with the label information, that is, the sample image can be attached with label information.
[0068] Through the aforementioned steps A1 to A5, sample images with label information can be obtained without manual labeling, so the cost of obtaining sample images is low, and a large number of sample images can be obtained for model training as needed. Moreover, the label information obtained by the above method is relatively more accurate than the label information obtained by manual labeling, and can better avoid common problems such as manual labeling errors. Therefore, the obtained label information is more accurate and reliable, which helps to train a more reliable information prediction model in both the reliability of sample images and the number of sample images.
[0069] Step B: Obtain information prediction results output by a preset neural network model for the sample image; the information prediction results include position information prediction results for key points corresponding to multiple body parts of the target object in the sample image, as well as body information prediction results. The structure of the neural network model is consistent with that of the aforementioned information prediction model, and the image processing method is also consistent. By adjusting the parameters of the neural network model, an information prediction model is ultimately obtained that can accurately output the position information of key points corresponding to multiple body parts and body information.
[0070] In step C, a neural network model is trained based on the label information and the information prediction results, thereby obtaining an information prediction model based on the trained neural network model. Specifically, the parameters of the neural network model can be adjusted to reduce the discrepancy between the label information and the information prediction results until the information prediction results of the neural network model meet the requirements, thereby obtaining an information prediction model.
[0071] In some specific implementation examples, step C can be performed with reference to the following steps C1 to C3:
[0072] Step C1: Determine a first loss based on the difference between the location tag and the location information prediction result. For example, the first loss can be determined based on the difference between the location tag corresponding to the sample image and the location information prediction result using a preset first loss function. The present disclosure does not limit the first loss function; for example, it can be an MSE (Mean-Square Error) loss function.
[0073] Step C2: Determine a second loss based on the difference between the body information label and the body information prediction result. For example, the second loss can be determined based on the difference between the body information label corresponding to the sample image and the body information prediction result using a preset second loss function. The disclosed embodiments do not limit the second loss function; for example, it can be an L2 loss function.
[0074] Step C3: training the neural network model based on the first loss and the second loss to obtain an information prediction model based on the trained neural network model.
[0075] In practical applications, the total loss can be determined based on the first loss and the second loss, and the network parameters in the neural network model can be adjusted based on the total loss. Training is stopped until the total loss converges to a preset threshold, and the trained neural network model is used as the information prediction model. The above method is to train the feature extraction network, key point prediction network, and part information prediction network in the neural network model at the same time. In addition, if the structure of the neural network model is the structure of Figure 3 or Figure 4, since the part information prediction network in the neural network model needs to rely on the information output by the key point prediction network, in this case, the feature extraction network and the key point prediction network can also be trained first, and then the parameters of the feature extraction network and the key point prediction network are fixed, and then the part information prediction network is trained. The specific training method can be flexibly selected and is not limited here.
[0076] In summary, the above-mentioned method provided by the embodiment of the present disclosure, compared with the related art that only outputs corresponding key points based on the skeleton structure, can directly use the information prediction model to obtain more abundant information such as the key point positions corresponding to multiple body parts and the length information and / or rotation angle information of the body parts at one time. This type of information is also more conducive to subsequent flexible application processing and can better meet image processing requirements. Moreover, the training data of the information prediction model provided by the embodiment of the present disclosure does not require manual labeling, and the reliability of the information prediction model is fully guaranteed in terms of the two dimensions of training sample quality and training sample quantity.
[0077] Corresponding to the aforementioned information acquisition method, an embodiment of the present disclosure provides an information acquisition device. FIG8 is a schematic structural diagram of an information acquisition device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can generally be integrated into an electronic device. By executing the information acquisition method, as shown in FIG8 , the information acquisition device includes:
[0078] The image acquisition module 802 is used to acquire a target image to be processed; wherein the target image contains a target object;
[0079] The model input module 804 is used to input the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network;
[0080] A feature extraction module 806 is used to extract features from the target image through a feature extraction network to obtain image features;
[0081] A key point location acquisition module 808 is configured to acquire location information of key points corresponding to multiple body parts of a target object through a key point prediction network based on image features;
[0082] The body information acquisition module 810 is used to acquire the body information of the target object based on image features through a body information prediction network; the body information includes length information of a first body part and / or rotation angle information of a second body part.
[0083] The above-mentioned device can directly use the information prediction model to obtain relatively rich information such as the key point positions corresponding to body parts and the length information and / or rotation angle information of the body parts at one time. Such information is also more conducive to subsequent flexible application processing and can better meet image processing needs.
[0084] In some implementations, the feature extraction module 806 is specifically configured to perform multi-scale feature extraction on the target image through the feature extraction network to obtain image features at multiple scales.
[0085] In some embodiments, the key point position acquisition module 808 is specifically configured to obtain position information of key points corresponding to multiple body parts of the target object through the key point prediction network based on one or more features of the image features at multiple scales.
[0086] In some embodiments, the key point position acquisition module 808 is specifically used to: fuse image features of at least two target scales among the multiple scales to obtain fused features; obtain a thermal map corresponding to the key points of the target part in the target object based on the fused features; and predict the position information of the key points of the target part in the target object based on the thermal map.
[0087] In some embodiments, the body information acquisition module 810 is specifically used to: acquire the body information of the target object through the part information prediction network based on the minimum-scale feature among the image features of the multiple scales; and / or acquire the fusion feature generated by the key point prediction network, and acquire the body information of the target object through the part information prediction network based on the fusion feature; wherein the fusion feature is a feature obtained by the key point prediction network by fusing image features of at least two target scales among the multiple scales.
[0088] In some embodiments, the part information prediction network includes a first prediction unit, a second prediction unit and a result fusion unit; the first prediction unit is used to obtain a first prediction result corresponding to the body information of the target object based on the minimum scale feature among the multiple scale image features; the second prediction unit is used to obtain a second prediction result corresponding to the body information of the target object based on the fusion feature; the result fusion unit is used to fuse the first prediction result and the second prediction result to determine the body information of the target object based on the fusion result.
[0089] In some embodiments, the result fusion unit is specifically configured to: perform weighted fusion processing based on the first prediction result and the second prediction result to obtain a weighted fusion result; and determine the body information of the target object according to the weighted fusion result.
[0090] In some embodiments, the weight corresponding to the first prediction result is not lower than the weight corresponding to the second prediction result.
[0091] In some embodiments, the network structures of the first prediction unit and the second prediction unit are the same, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein the global processing layer is used to perform global processing based on the input information of the network structure to obtain global information; the length information prediction layer is used to perform a first decoupling processing based on the global information to predict the length information of the first body part; the rotation angle prediction layer is used to perform a second decoupling processing based on the global information to predict the rotation angle information of the second body part.
[0092] In some embodiments, the device also includes a model acquisition module for obtaining the information prediction model according to the following steps: obtaining a sample image carrying label information; wherein the sample image contains a target object, and the label information includes position labels of key points corresponding to multiple body parts of the target object and body information labels of the target object; obtaining information prediction results output by a preset neural network model for the sample image; the information prediction results include position information prediction results and body information prediction results of key points corresponding to multiple body parts of the target object in the sample image; based on the label information and the information prediction results, the neural network model is trained to obtain an information prediction model based on the trained neural network model.
[0093] In some embodiments, the model acquisition module is specifically used to: acquire a sample image containing a target object; use a preset network model to acquire multiple body key points of the three-dimensional structure corresponding to the target object in the sample image; determine the position information of the key points corresponding to the multiple body parts of the target object in the sample image based on the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object; determine the body information of the target object in the sample image based on the position information of the key points corresponding to the multiple body parts of the target object in the sample image; and attach label information to the sample image based on the position information of the key points corresponding to the multiple body parts of the target object in the sample image and the body information of the target object in the sample image.
[0094] In some embodiments, the model acquisition module is specifically used to: determine at least two first target points corresponding to the first body part and at least two second target points corresponding to the second body part based on the position information of key points corresponding to multiple body parts of the target object in the sample image; determine the length information of the first body part based on the at least two first target points; and determine the rotation angle information of the second body part based on the at least two second target points.
[0095] In some embodiments, the first body part includes the area between the shoulder and waist of the target object; the at least two first target points include: a first target point determined based on the center point between the key point corresponding to the left shoulder part and the key point corresponding to the right shoulder part, and a first target point determined based on the center point between the key point corresponding to the left waist part and the key point corresponding to the right waist part.
[0096] In some embodiments, the second body part includes a shoulder, and the at least two second target points include: a key point corresponding to a left shoulder and a key point corresponding to a right shoulder; and / or, the second body part includes a waist, and the at least two second target points include: a key point corresponding to a left waist and a key point corresponding to a right waist.
[0097] In some embodiments, the model acquisition module is specifically used to: determine a first loss based on the difference between the location tag and the location information prediction result; determine a second loss based on the difference between the body information tag and the body information prediction result; and train the neural network model based on the first loss and the second loss to obtain an information prediction model based on the trained neural network model.
[0098] In some embodiments, the plurality of body parts include arms and / or legs, and further include one or more of the head, neck, chest, abdomen, and waist.
[0099] The information acquisition device provided by the embodiments of the present disclosure can execute the information acquisition method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0100] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device embodiment can refer to the corresponding process in the method embodiment, and will not be repeated here.
[0101] An embodiment of the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method in the present disclosure. An embodiment of the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method in the present disclosure.
[0102] Reference is now made to FIG9 , which illustrates a schematic diagram of the structure of an electronic device 900 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG9 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0103] As shown in Figure 9, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0104] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG9 shows the electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0105] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0106] In addition to the above-mentioned methods and devices, the embodiments of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the image processing method provided by the embodiments of the present disclosure. The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present disclosure, the programming languages including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0107] In addition, the embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the information acquisition method provided by the embodiment of the present disclosure.
[0108] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0109] The embodiments of the present disclosure further provide a computer program product, including a computer program / instruction, which implements the information acquisition method in the embodiments of the present disclosure when executed by a processor.
[0110] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0111] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0112] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0113] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0114] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0115] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for obtaining information, comprising: Acquire a target image to be processed; wherein the target image contains a target object; Inputting the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network; Extracting features from the target image through the feature extraction network to obtain image features; Based on the image features, obtaining position information of key points corresponding to multiple body parts of the target object through the key point prediction network; Based on the image features, body information of the target object is acquired through the body information prediction network; the body information includes length information of a first body part and / or rotation angle information of a second body part.
2. The method according to claim 1, wherein: The step of extracting features from the target image through the feature extraction network to obtain image features includes: Multi-scale feature extraction is performed on the target image through the feature extraction network to obtain image features of multiple scales.
3. The method according to claim 2, wherein: The step of obtaining the position information of the key points corresponding to the multiple body parts of the target object through the key point prediction network based on the image features includes: Based on one or more features of the image features at multiple scales, position information of key points corresponding to multiple body parts of the target object is obtained through the key point prediction network.
4. The method according to claim 3, wherein: The step of obtaining the position information of key points corresponding to the plurality of body parts of the target object includes: Fusing image features of at least two target scales among the multiple scales to obtain fused features; Acquire a thermal map corresponding to a key point of a target part in the target object according to the fusion feature; The position information of the key points of the target part in the target object is predicted based on the heat map.
5. The method according to claim 2, wherein: The acquiring the body information of the target object through the part information prediction network based on the image features includes: Based on the smallest-scale feature among the image features of the multiple scales, obtaining the body information of the target object through the part information prediction network; and / or, Acquire the fusion feature generated by the key point prediction network, and based on the fusion feature, acquire the body information of the target object through the part information prediction network; wherein the fusion feature is a feature obtained by the key point prediction network by fusing image features of at least two target scales among the multiple scales.
6. The method according to claim 5, wherein: The part information prediction network includes a first prediction unit, a second prediction unit and a result fusion unit; The first prediction unit is used to obtain a first prediction result corresponding to the body information of the target object based on a feature of a minimum scale among the image features of the multiple scales; The second prediction unit is used to obtain a second prediction result corresponding to the body information of the target object based on the fusion feature; The result fusion unit is used to fuse the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result.
7. The method according to claim 6, wherein: The fusing the first prediction result and the second prediction result to determine the body information of the target object according to the fusion result includes: Performing weighted fusion processing based on the first prediction result and the second prediction result to obtain a weighted fusion result; The body information of the target object is determined according to the weighted fusion result.
8. The method according to claim 7, wherein: The weight corresponding to the first prediction result is not less than the weight corresponding to the second prediction result.
9. The method according to claim 7, wherein: The network structure of the first prediction unit and the second prediction unit is the same, and the network structure includes a global processing layer, a length information prediction layer, and a rotation angle prediction layer; wherein, The global processing layer is used to perform global processing based on the input information of the network structure to obtain global information; The length information prediction layer is used to perform a first decoupling process based on the global information to predict the length information of the first body part; The rotation angle prediction layer is used to perform a second decoupling process based on the global information to predict the rotation angle information of the second body part.
10. The method according to any one of claims 1 to 9, wherein: The information prediction model is obtained according to the following steps: Acquire a sample image carrying label information; wherein the sample image includes a target object, and the label information includes position labels of key points corresponding to multiple body parts of the target object and a body information label of the target object; Obtaining information prediction results output by a preset neural network model for the sample image; the information prediction results include position information prediction results and body information prediction results of key points corresponding to multiple body parts of the target object in the sample image; Based on the label information and the information prediction result, the neural network model is trained to obtain an information prediction model based on the trained neural network model.
11. The method according to claim 10, wherein: The obtaining of a sample image carrying label information includes: Acquire a sample image containing a target object; Using a preset network model to obtain a plurality of body key points of a three-dimensional structure corresponding to a target object in the sample image; Determine position information of the key points corresponding to the multiple body parts of the target object in the sample image according to the multiple body key points and the mapping relationship between the multiple body key points and the key points corresponding to the multiple body parts of the target object; Determining body information of the target object in the sample image according to position information of key points corresponding to multiple body parts of the target object in the sample image; Based on the position information of key points corresponding to multiple body parts of the target object in the sample image and the body information of the target object in the sample image, label information is attached to the sample image.
12. The method according to claim 11, wherein: The determining the body information of the target object in the sample image according to the position information of the key points corresponding to the multiple body parts of the target object in the sample image comprises: Determine at least two first target points corresponding to a first body part and at least two second target points corresponding to a second body part according to position information of key points corresponding to multiple body parts of the target object in the sample image; Determining length information of the first body part according to the at least two first target points; The rotation angle information of the second body part is determined according to the at least two second target points.
13. The method according to claim 12, wherein: The first body part includes the area between the shoulder and waist of the target subject; The at least two first target points include: a first target point determined based on a center point between a key point corresponding to a left shoulder and a key point corresponding to a right shoulder, and a first target point determined based on a center point between a key point corresponding to a left waist and a key point corresponding to a right waist.
14. The method according to claim 12, wherein: The second body part includes a shoulder, and the at least two second target points include: a key point corresponding to a left shoulder part and a key point corresponding to a right shoulder part; and / or, The second body part includes a waist, and the at least two second target points include: a key point corresponding to a left waist part and a key point corresponding to a right waist part.
15. The method according to claim 10, wherein: The step of training the neural network model based on the label information and the information prediction result to obtain an information prediction model based on the trained neural network model includes: Determining a first loss based on a difference between the location tag and the location information prediction result; determining a second loss based on a difference between the body information tag and the body information prediction result; Based on the first loss and the second loss, the neural network model is trained to obtain an information prediction model based on the trained neural network model.
16. The method according to claim 1, wherein: The multiple body parts include arms and / or legs, and also include one or more of the head, neck, chest, abdomen and waist.
17. An information acquisition device, comprising: An image acquisition module, used to acquire a target image to be processed; wherein the target image contains a target object; A model input module, used to input the target image into a preset information prediction model; wherein the information prediction model includes a feature extraction network, a key point prediction network and a part information prediction network; A feature extraction module, used to extract features of the target image through the feature extraction network to obtain image features; A key point position acquisition module, used to acquire the position information of key points corresponding to multiple body parts of the target object through the key point prediction network based on the image features; A body information acquisition module is used to acquire the body information of the target object through the part information prediction network based on the image features; the body information includes length information of a first body part and / or rotation angle information of a second body part.
18. An electronic device, comprising: a storage device having a computer program stored thereon; A processing device, used to execute the computer program in the storage device to implement the steps of the information acquisition method according to any one of claims 1 to 16.
19. A computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program is used to execute the information acquisition method described in any one of claims 1 to 16.
Citation Information
Patent Citations
Target object key point detection method and apparatus, and deep learning neural network
CN108229343A
Method for predicting attitude of three-dimensional model and electronic equipment
CN115482588A
Image processing method and device
CN116612495A
A method and system for body part measurement for skin treatment
WO2023094377A1