A human attribute prediction method and device, electronic equipment and storage medium

By acquiring and training human body key points and combining multiple human body attribute prediction models, the problems of cumbersome and inefficient human body attribute prediction in existing technologies are solved, and fast and simplified multi-human body attribute prediction and 3D contour display are realized.

CN114120351BActive Publication Date: 2026-04-24SO-YOUNG INT INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SO-YOUNG INT INC
Filing Date
2020-08-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing methods for predicting human attributes are too cumbersome, require the collection of large amounts of sample data, and have low prediction efficiency, making it impossible to effectively utilize non-contact measurement technology.

Method used

By acquiring a preset number of human key points of the target object, training a detection model, and combining a first human attribute prediction model and a second human attribute prediction model (such as based on pupil distance and a preset neural network model), the coordinates of the human key points are output to predict human attributes. The prediction results of multiple models are integrated to improve accuracy and efficiency.

Benefits of technology

It enables the rapid and simplified prediction of multiple human attributes in a single image, improving detection efficiency and accuracy. It can display the 3D contour of the target object and simultaneously identify and detect key points of the human contour in multi-target scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120351B_ABST
    Figure CN114120351B_ABST
Patent Text Reader

Abstract

The application discloses a human attribute prediction method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a preset number of human key points of a target object; inputting the preset number of human key points into a detection model for training, and outputting a preset number of human key point coordinates; based on the preset number of human key point coordinates, predicting each human attribute of the target object according to at least one human attribute prediction model to obtain a prediction result. By using the embodiment of the application, the size of the sample data obtained in different application scenarios can be effectively combined, and a human attribute prediction model matched with the current application scenario is preferably selected, so that the prediction efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for predicting human attributes. Background Technology

[0002] Traditional methods for obtaining anthropometric information primarily involve testing personnel using tools such as measuring tapes to measure clients. While this method boasts high accuracy, its drawbacks are also significant: it is time-consuming, cumbersome, and requires in-person contact, necessitating the user's presence on-site.

[0003] In recent years, non-contact measurement has gradually entered the public eye. For example, the iPin laser ruler uses internal devices to calculate the time it takes for the laser to travel from emission to reflection, thus determining the distance between the object being measured and the phone. Using this distance as a scale, a mobile app can calculate all the measurement values ​​of objects on the same plane in a photograph. However, this method requires the participation of an additional laser device, limiting its versatility. Furthermore, this method can only measure length and width; it is ineffective for measurements of circumferences such as chest and waist.

[0004] Existing methods for predicting human attributes are often too cumbersome, requiring the collection of large amounts of sample data. Furthermore, the process of making predictions based on sample data is also often too cumbersome. Therefore, the existing prediction methods are not very efficient.

[0005] Therefore, how to improve the prediction efficiency of existing methods for predicting various human attributes of users is a technical problem to be solved. Summary of the Invention

[0006] Therefore, it is necessary to provide a human attribute prediction method, device, electronic device, and storage medium to address the problem of low prediction efficiency of existing human indicator prediction methods.

[0007] In a first aspect, embodiments of this application provide a method for predicting human attributes, the method comprising:

[0008] Obtain a preset number of human body key points of the target object;

[0009] The preset number of human key points are input into the detection model for training, and the preset number of human key point coordinates are output.

[0010] Based on the preset number of human body key point coordinates, and according to at least one human body attribute prediction model, various human body attributes of the target object are predicted to obtain prediction results, which are used to predict various human body attributes of the target object.

[0011] In one embodiment, the human attribute prediction model includes a first human attribute prediction model and a second human attribute prediction model. The first human attribute prediction model is a prediction model that predicts human attributes based on the interpupillary distance of the target object, and the second human attribute prediction model is a preset neural network prediction model.

[0012] Based on the preset number of human body key point coordinates, and according to at least one human body attribute prediction model, the prediction results for various human body attributes of the target object include:

[0013] Based on the preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a first prediction result; or...

[0014] Based on the preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a second prediction result; or...

[0015] Based on the preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model respectively, and the corresponding prediction results are obtained. The two prediction results are then fused to obtain the corresponding prediction result.

[0016] In one implementation, the step of predicting various human attributes of the target object based on the preset number of human key point coordinates and according to the first human attribute prediction model to obtain a first prediction result includes:

[0017] Select any one of the first human attributes from the first category as the current first human attribute;

[0018] Obtain the calculation formula used to calculate the attributes of the current first human body;

[0019] Based on the preset number of human body key point coordinates, and taking the pupil distance as a reference, the current first human body attribute is predicted according to the calculation formula to obtain the first prediction result; wherein, the first prediction result includes: a prediction result obtained based on the pupil distance and used to characterize any one of the first human body attributes in the first type of human body attributes.

[0020] In one implementation, the step of predicting various human attributes of the target object based on the preset number of human key point coordinates and according to the second human attribute prediction model to obtain a second prediction result includes:

[0021] Obtain big data that can determine any one of the second-class human attributes;

[0022] By analyzing the big data, preset conditions can be obtained to determine any one of the second human attributes;

[0023] Select any one of the second type of human attributes as the current second human attribute;

[0024] Obtain preset sub-conditions that match the current second human body attribute from the preset conditions;

[0025] Based on the preset number of human body key point coordinates, the current second human body attribute is predicted according to the preset sub-conditions to obtain the second prediction result.

[0026] In one implementation, the step of predicting various human attributes of the target object based on the preset number of human key point coordinates, according to the first human attribute prediction model and the second human attribute prediction model respectively, to obtain the corresponding prediction results includes:

[0027] Based on the preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object to obtain a first prediction result.

[0028] Based on the preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object to obtain a second prediction result.

[0029] The first prediction result and the second prediction result are fused to obtain and output the fused prediction result.

[0030] In one implementation, the step of predicting various human attributes of the target object based on the preset number of human key point coordinates and according to the second human attribute prediction model to obtain a second prediction result further includes:

[0031] Read the coordinates of a preset number of key human body points of the target object;

[0032] The coordinates of a preset number of human body key points of the target object are normalized and preprocessed to obtain training data;

[0033] The training data is input into the preset neural network prediction model to obtain and output the second prediction result.

[0034] Secondly, embodiments of this application provide a human attribute prediction device, the device comprising:

[0035] The acquisition unit is used to acquire a preset number of human body key points of the target object;

[0036] The training unit is used to input the preset number of human key points acquired by the acquisition unit into the detection model for training, and output the coordinates of the preset number of human key points.

[0037] The prediction unit is used to predict various human attributes of the target object based on the preset number of human key point coordinates output by the training unit and according to at least one human attribute prediction model, and to obtain prediction results. The prediction results are used to predict various human attributes of the target object.

[0038] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method steps described above.

[0039] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method steps described above.

[0040] The technical solutions provided in this application embodiment may include the following beneficial effects:

[0041] In this embodiment of the application, a preset number of human body key points of the target object are obtained; the preset number of human body key points are input into the detection model for training, and the preset number of human body key point coordinates are output; based on the preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to at least one human body attribute prediction model, and the prediction results are obtained. The prediction results are used to predict the various human body attributes of the target object. By employing the embodiments of this application, the size of the sample data obtained under different application scenarios can be effectively combined to select the human attribute prediction model that best matches the current application scenario, thus effectively improving prediction efficiency. Furthermore, compared to existing detection methods that can only detect individual skeletal key points (i.e., the detection result only displays the two-dimensional contour of the target object), the embodiments of this application, because the output result includes the coordinates of a preset number of human contour key points, include not only the coordinates of all contour points covering the entire body contour of the target object but also the coordinates of all contour points covering the facial features of the target object. Thus, the detection result can display the three-dimensional contour of the target object, providing a sense of depth. Moreover, the detection method provided by this invention, when an image includes multiple objects, and all of these objects are target objects, can simultaneously identify multiple target objects and synchronously detect the coordinates of a preset number of human contour key points for each of the multiple target objects, thereby improving detection efficiency.

[0042] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0043] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0044] Figure 1 This is an application scenario diagram of a human attribute prediction method provided in an embodiment of this disclosure;

[0045] Figure 2 This is a flowchart illustrating a human attribute prediction method provided in an embodiment of this disclosure;

[0046] Figure 3 This is a schematic diagram of the human body contour points of a target object in a specific application scenario provided in this embodiment of the disclosure;

[0047] Figure 4 This is a flowchart illustrating another human attribute prediction method provided in this embodiment of the disclosure;

[0048] Figure 5 This is a flowchart illustrating another human attribute prediction method provided in this embodiment of the disclosure;

[0049] Figure 6 A schematic diagram is shown of an index used to evaluate any one of the first human attributes in the first category;

[0050] Figure 7 Another schematic diagram is shown of the indicators used to evaluate any one of the first human attributes in the first category;

[0051] Figure 8 This is a schematic diagram of the neural network structure of the preset neural network prediction model in the embodiments of this disclosure;

[0052] Figure 9 This is a schematic diagram of the structure of a human attribute prediction device provided in an embodiment of this disclosure;

[0053] Figure 10 A schematic diagram of an electronic device connection structure according to an embodiment of the present disclosure is shown. Detailed Implementation

[0054] The following description and accompanying drawings fully illustrate specific embodiments of the invention to enable those skilled in the art to practice them.

[0055] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0056] The optional embodiments of this disclosure are described in detail below with reference to the accompanying drawings.

[0057] like Figure 1 The diagram illustrates an application scenario of this disclosure. In this scenario, multiple users operate a client installed on a mobile phone or other terminal device, and the client communicates with a backend server via a network. A specific application scenario involves a prediction process based on a preset number of human body keypoint coordinates to predict various human body attributes. However, this is not limited to this single application scenario; any scenario applicable to this implementation is included. For ease of explanation, this embodiment uses the application scenario of predicting various human body attributes of a target object in a predicted image as an example. Figure 1 As shown, the images obtained, which at least include the target object, often come from sources such as... Figure 1 One of the multiple clients shown.

[0058] like Figure 2 As shown in the figure, this disclosure provides a method for predicting human attributes, applied on the server side, and specifically includes the following method steps:

[0059] S202: Obtain a preset number of human body key points of the target object.

[0060] In practical applications, to obtain multiple images containing different objects, one can start from, for example... Figure 1 The diagram shows multiple clients acquiring images of multiple objects. Furthermore, to make the prediction results obtained by the prediction method more accurate, images are often selected that fully cover at least one object, and these full-body images clearly show the facial features of the object.

[0061] In this step, there is no specific limit to the preset number of key points on the target object's body. The preset number is sufficient to cover all key areas corresponding to the target object's overall outline and to fully cover the target object's facial features. In practical applications, the preset number can be adjusted according to the needs of different application scenarios; therefore, no specific limit is placed on the preset number here.

[0062] like Figure 3 The diagram shown is a schematic representation of the human body contour points of a target object in a specific application scenario provided by an embodiment of this disclosure. For ease of explanation and illustration, Figure 3Instead of using real photos, images of models were used; this is merely for illustrative purposes. Figure 3 As shown, the target object's human body contour points are marked with a total of 77 points. Specifically, the location represented by each point and the numerical value of each point can be found in [the table / reference]. Figure 3 As is clearly seen in the text, I will not elaborate further here. Through methods such as... Figure 3 The annotation of the 77 contour key points shown can achieve precise positioning of the facial features and human body contour of the target object. The facial features of the target object include the eyes, nose and mouth.

[0063] S204: Input a preset number of human key points into the detection model for training, and output the coordinates of the preset number of human key points.

[0064] In practical applications, at least one target object's preset number of human contour key points are input into a preset detection model for training, and the detection results are output. The detection results include the coordinates of multiple key points used to determine the boundary of the detection box, as well as the coordinates of the preset number of human contour key points. Compared to existing detection methods that can only detect individual human skeletal key points (i.e., the detection results only display the two-dimensional contour of the target object), the present invention, by including the coordinates of a preset number of human contour key points in the output results, provides a three-dimensional representation of the target object. These coordinates include not only the coordinates of all contour points that fully cover the entire body contour of the target object, but also the coordinates of all contour points that cover the facial features. Furthermore, the detection method provided by this invention, when an image contains multiple objects, and all of these objects are target objects, can simultaneously identify multiple target objects and synchronously detect the coordinates of the preset number of human contour key points for each of the multiple target objects, thereby improving detection efficiency. In one possible implementation, inputting at least one target object's preset number of human contour key points into a preset detection model for training includes the following steps:

[0065] Obtain the dataset, which includes a preset number of first human contour key points. The first human contour key points are human contour key points annotated by an annotation tool. The annotation tool can be Labelme, a common image annotation tool that allows users to create customized annotation tasks or perform image annotation.

[0066] The dataset is stored in a file with a predefined format, which can be a JSON (JavaScript Object Notation) file. JSON is a lightweight data-interchange format that uses a text format completely independent of programming languages ​​to store and represent data. In practical applications, its concise and clear hierarchical structure makes JSON an ideal data exchange language. JSON is easy for humans to read and write, and also easy for machines to parse and generate, effectively improving network transmission efficiency.

[0067] Images and files in a preset format are used as training samples and input into a preset detection model for training. The prediction results are output, including a preset number of second human body contour key points, which are the predicted human body contour key points.

[0068] Based on a preset number of first human contour key points and a preset number of second human contour key points, the loss function of the preset detection model is calculated, wherein the neural network corresponding to the preset detection model is converged through the loss function.

[0069] The training samples are iteratively trained to obtain a preset detection model, and a weight file including the preset detection model is output.

[0070] It should be noted that in practical applications, before inputting at least a preset number of human contour key points of the target object into the preset detection model for training, a batch of training images should be prepared. Each image should include the target object, and it should be ensured that the entire body of the target object in the image is not obscured by external occlusions, that is, the full body contour of the target object can be clearly identified.

[0071] In this step, the preset detection model is constructed based on a preset detection algorithm. The preset detection algorithm in the detection method of this embodiment is described below:

[0072] The default detection algorithm uses a detection framework like CenterNet. By replacing the basic network structure with EfficientNet-B0, a good balance between accuracy and speed can be achieved. The overall network structure is composed of EfficientNet-B0 plus deconvolution modules.

[0073] The EfficientNet-B0 part consists of a normal convolutional layer plus seven MBconv layers. The structure of the MBconv layers is as follows: Conv represents convolution, Batchnorm represents normalization, DepthwiseConv represents channel-separating convolution, Swish represents the activation function, drop_connect represents a module that randomly deactivates neuron connections, which can increase the model's generalization ability, Average pooling represents average pooling, and sigmoid represents the activation function.

[0074] In the EfficientNet-B0 part, the input image size is 512*512*3 (height, width, channel). After 5 downsampling operations (reducing width and height), the final output feature map size of this module is 16*16*320 (512 / 25 = 16). Then, it goes through 3 deconvolution modules, each of which upsamples the input feature map once (increasing width and height). The final output feature map size after 3 deconvolutions is 128*128*256. Based on this, a 1*1 convolution is performed to obtain the final 6 network output results.

[0075] In one possible implementation, inputting a predetermined number of human contour key points of at least one target object into a predetermined detection model for training, and outputting the detection results includes the following steps:

[0076] At least one target object with a preset number of human contour key points is input into a preset detection model for training, and the network output results in multiple dimensions are output. The network output results in multiple dimensions include a first network output result that represents the offset value of the preset number of human contour key points based on the center point, a second network output result that represents the feature map of the preset number of human contour key points, and a third network output result that represents the offset value of multiple human contour key points in the preset number.

[0077] Based on the output results of the first network, the second network, and the third network, determine and output the coordinates of a preset number of human body contour key points.

[0078] In a specific application scenario, the output results of the above three networks can be as follows:

[0079] Hps (corresponding to the output of the first network): The key points of the object are offset based on the center point, and the output has k*2 channels, where k represents the number of key points. Assuming that in a certain application scenario, the human body contour key points of the target object are 77 human body contour key points, and 2 represents the x and y offsets, the final size is 77*2=154.

[0080] hm_hp (corresponding to the output of the second network): Feature map of key points of the object, outputting k channels, one channel per point, with a final size of 128*128*77.

[0081] hp_offset (corresponding to the output of the third network): the offset of K key points, all of which use the same offset, with two indicators, x and y, outputting 2 channels, and the final size is 128*128*2.

[0082] In one possible implementation, the detection result also includes the coordinates of multiple key points used to determine the boundary of the detection box. The process of inputting a preset number of human contour key points of at least one target object into a preset detection model for training, and outputting the detection result, further includes the following steps:

[0083] At least one target object with a preset number of human contour key points is input into a preset detection model for training, and the network output results in multiple dimensions are output. The network output results in multiple dimensions also include the fourth network output result corresponding to the feature map used to characterize the target object category, the fifth network output result used to characterize the size of the target object, and the sixth network output result corresponding to the offset value of the center point of the target object.

[0084] Based on the outputs of the fourth, fifth, and sixth networks, the coordinates of multiple key points used to determine the boundaries of the detection box are determined and output.

[0085] In a specific application scenario, the output results of the above three networks can be as follows:

[0086] Hm (corresponding to the output of the fourth network): the feature map of the object, one channel (thickness) for one category. Since the detection method of this embodiment only detects the human body category, the size of this layer is 128*128*1.

[0087] Wh (corresponding to the output of the fifth network): The size of the object, that is, the width and height. Therefore, there are 2 channels, and the final size is 128*128*2.

[0088] Reg (corresponding to the output of the sixth network): The offset of the object's center point, including two offsets, x and y. Therefore, it has two channels and the final size is 128*128*2.

[0089] Use the following formula to convert the keypoint coordinates, as detailed below:

[0090]

[0091] Here, (x, y) represents the coordinates of the center point of the target object, i.e., the hm branch. Jxyj represents the offset of the (x, y) coordinates of the j-th point, i.e., the hps branch. However, the error in predicting keypoint coordinates using this direct regression-based method is relatively large.

[0092] Lj = hm_hp + hp_offset

[0093] Direct regression can determine the approximate location of each contour point. Then, the point Lj(hm_hp) with the highest confidence score (0-1 score) greater than 0.1 near this location is selected as the true contour point. Finally, the coordinates of the contour point are obtained by adding an offset (hp_offset) to this point. It should be noted that if there are two or more points with a confidence score greater than 0.1, the point closest to lj is selected, as shown in the following formula:

[0094]

[0095] It's important to note that `lj` represents the coordinates of keypoints obtained using the regression-based method, i.e., the HPS branch. `Lj` represents the coordinates of keypoints obtained using the feature map-based method. There are a total of 77 keypoints. For each of the 77 regression points `lj`, the distance between them and the corresponding feature map branch `Lj` is calculated. The point in `Lj` with the closest distance is taken as the final predicted keypoint. For example, if the keypoint for the eye position is predicted based on regression, then in the feature map of the channel containing the eye, the point with the closest confidence level higher than 0.1 is found and used as the final keypoint for the eye position.

[0096] In this way, the above detection model can accurately detect entry and exit points. Figure 3 The coordinates of 77 key points of the target object's human silhouette are shown, so that the various human attributes of the target object can be predicted based on these 77 key points.

[0097] S206: Based on a preset number of human body key point coordinates, predict various human body attributes of the target object according to at least one human body attribute prediction model, and obtain prediction results. The prediction results are used to predict various human body attributes of the target object.

[0098] In this step, the human attribute prediction model that matches the current application scenario can be selected based on the size of the sample data to be predicted by the human attribute prediction model in different application scenarios.

[0099] Specifically, when the amount of sample data in the current application scenario is large, for example, when a large amount of sample data can be obtained based on existing big data statistical technology, the first human attribute prediction model can be used to predict various human attributes of the target object. The first human attribute prediction model is a prediction model for predicting human attributes based on the pupil distance of the target object.

[0100] In the case of a small amount of sample data in the current application scenario, a second human attribute prediction model can be used to predict various human attributes of the target object. The second human attribute prediction model is a preset neural network prediction model.

[0101] Furthermore, given the large volume of sample data in the current application scenario and the high accuracy requirements for the prediction results, it is advisable to simultaneously use a first human attribute prediction model to predict various human attributes of the target object to obtain a first prediction result, and a second human attribute prediction model to predict various human attributes of the target object to obtain a second prediction result. The first and second prediction results are then fused to obtain and output the fused prediction result.

[0102] In practical applications, it is only necessary to use, as follows Figure 1 Any image acquisition device on any of the clients shown can acquire at least one image of the target object. The requirements for the image of the target object are: a frontal image of the target object, which fully covers the entire outer contour of the target object and clearly shows the facial features of the target object.

[0103] Because only a single clear, frontal image of the target object is required, the prediction process for various human attributes of the target object using a predictive model is simplified, improving prediction efficiency. Furthermore, since only one frontal image of the target object needs to be acquired, the image acquisition process is simplified, ultimately improving the user experience.

[0104] Among them, any of the above-mentioned human body key point coordinates includes not only the horizontal coordinate but also the corresponding vertical coordinate.

[0105] In one possible implementation, the human attribute prediction model includes a first human attribute prediction model and a second human attribute prediction model. The first human attribute prediction model is a prediction model that predicts human attributes based on the interpupillary distance of the target object. The second human attribute prediction model is a preset neural network prediction model. Based on a preset number of human key point coordinates, and according to at least one human attribute prediction model, the various human attributes of the target object are predicted to obtain the prediction results, which includes the following steps:

[0106] Based on a preset number of human body key point coordinates, various human body attributes of the target object are predicted according to the first human body attribute prediction model to obtain a first prediction result. The first human body attribute prediction model is suitable for application scenarios with a large amount of sample data for prediction; or...

[0107] Based on a preset number of human body key point coordinates, various human body attributes of the target object are predicted according to the second human body attribute prediction model, resulting in a second prediction result. The second human body attribute prediction model is suitable for application scenarios where the amount of sample data for prediction is relatively small; or...

[0108] Based on a preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model respectively, and the corresponding prediction results are obtained. The two prediction results are then fused together. At the same time, the first human body attribute prediction model and the second human body attribute prediction model are used to predict the various human body attributes of the target object respectively, and the final prediction result is fused and output.

[0109] In one possible implementation, based on a preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model, respectively, to obtain the corresponding prediction results, including the following steps:

[0110] Based on a preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, and the first prediction result is obtained.

[0111] Based on a preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, and a second prediction result is obtained.

[0112] The first and second prediction results are fused to obtain and output the fused prediction result.

[0113] In this step, human attributes include a first human attribute characterized by any first type of human attribute and a second human attribute characterized by any second type of human attribute. The first human attribute prediction model is a prediction model for predicting human attributes based on the pupil distance of the target object, and the second human attribute prediction model is a preset neural network prediction model.

[0114] It should be noted that the first type of human attributes refers to the human attributes of the first sample data used in the prediction by the first human attribute prediction model, such as height, chest circumference, lower chest circumference, waist circumference, belly circumference, hip circumference, hip width, hip circumference, shoulder length, arm length, arm circumference, thigh circumference, calf circumference, leg length, head length, length from shoulder to the base of the thumb, vertical length from shoulder to crotch line, and shoulder width.

[0115] The second type of human attribute refers to the human attributes possessed by the second sample data used in the prediction using the second human attribute prediction model. This second type of human attribute refers to the coordinates of various key human points of the input target object, for example... Figure 3 The coordinates of the key points of the 77 target objects.

[0116] In this step, the second human attribute prediction model is a preset neural network prediction model. For example... Figure 8 The diagram shown is a schematic diagram of the neural network structure of the preset neural network prediction model in an embodiment of this disclosure.

[0117] In the prediction method provided in this embodiment, based on the preset neural network prediction model, predictions for 18 specific numerical values ​​(cm) and 11 specific categories (from fine to coarse) of indicators are simultaneously achieved.

[0118] Based on this pre-defined neural network prediction model, the input consists of normalized values ​​corresponding to the coordinates of various key human body points of the target object. The normalization method is as follows.

[0119] Assume the coordinates of the key points on the human body are,

[0120] P = {p1, p2, p3, ..., p77}

[0121] Therefore, all of the target object, such as Figure 3 Among the key points shown, the smallest coordinate, the upper left corner point Pmin, and the largest coordinate, the lower right corner point Pmax, can be determined.

[0122] Pmin = (min(Px), min(Py));

[0123] Pmax = (max(Px), max(Py));

[0124] The coordinates of point P after normalization are Pnorm.

[0125] Pnorm=(P-Pmin) / (Pmax-Pmin).

[0126] like Figure 3 As shown, there are a total of 77 key points on the target object's human body, specifically as follows: Figure 3As shown, further details are omitted. Each keypoint has two coordinates, x and y, and the final network input is 77 * 2 = 154. It then passes through five fully connected layers, each outputting two branches. The first branch has a dimension of 18 and is primarily responsible for predicting 18 specific numerical values ​​(cm). The second branch outputs a dimension of 11 * 5, where 11 represents the 11 specific categories of indicators (from fine to coarse), and 5 represents that each indicator has five categories. Thus, ultimately, a single neural network can predict all human attribute indicators.

[0127] It's important to note that the 11 specific categories (from finest to coarser) of indicators predicted here are directly the final results. The 18 specific numerical values ​​(cm) are normalized outputs, with the output range also between 0 and 1. Multiplying this output by the user's height yields the final results for each indicator. The user's height is still obtained using the interpupillary distance method described above.

[0128] The training process of this neural network:

[0129] The training data was obtained from actual data collected from 1000 models, including images, such as... Figure 3 The 18 cm indicators and 11 classification indicators for the target object shown have been obtained. Using a similar normalization approach as above, normalizing these real measurements will yield the training labels, i.e., the ground truth values.

[0130] The key point coordinates of multiple human figures in the target object are extracted using conventional methods. The x and y coordinates are then normalized using the same normalization method described above, resulting in the training data.

[0131] By inputting the training data into the neural network above, the output result can be obtained. The loss function is calculated by comparing the output result with the ground truth value and then fed back to continuously optimize the neural network. After repeated iterations, the weight file of the trained neural network model can be obtained. Based on this weight file, the various human attributes of the target object can be predicted.

[0132] In this step, the first prediction result includes: a prediction result based on pupil distance prediction, used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on preset conditions, used to characterize any one of the second human attributes in the second category of human attributes.

[0133] In this step, the second prediction result includes a prediction result based on a preset neural network prediction model, which is used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on a preset neural network prediction model, which is used to characterize any one of the second human attributes in the second category of human attributes.

[0134] It should be noted that the fusion method used to combine the first and second prediction results is a conventional data fusion method, which will not be elaborated here.

[0135] In this step, the first prediction result and the second prediction result are fused to obtain and output the fused prediction result, which includes the following steps:

[0136] Based on the first prediction result, the first weight value pre-configured based on the first prediction result, the second prediction result, and the second weight value pre-configured based on the second prediction result, the first prediction result and the second prediction result are fused together to obtain and output the fused prediction result.

[0137] In practical applications, different first and second weight values ​​can be configured according to the needs of different application scenarios. A common application scenario is to configure the first and second weight values ​​to be equal or nearly equal, for example, the first weight value accounts for 50% and the second weight value also accounts for 50%. This can effectively balance the first and second prediction results, making the fused prediction output more consistent with the real application scenario.

[0138] In this step, there are no specific restrictions on the values ​​of the first and second weight values. In a certain application scenario, in order to balance the first prediction result obtained by the first human attribute prediction model and the second prediction result obtained by the second human attribute prediction surface model, the first weight value is often configured to be equal to the second weight value. In different application scenarios, the first and second weight values ​​can be adjusted, which will not be elaborated here.

[0139] like Figure 4 As shown in the figure, this disclosure provides another method for predicting human attributes, applied to the server side, specifically including the following method steps:

[0140] S402: Obtain key points of the human body.

[0141] For methods for obtaining key points of the human body, please refer to the description of the same or similar parts in the foregoing method embodiments, which will not be repeated here.

[0142] S404: Input a preset number of human key points into the detection model for training, and output the coordinates of the preset number of human key points.

[0143] Based on the method of outputting a preset number of human body key point coordinates of the target object through continuous training, please refer to the description of the same or similar parts in the foregoing method embodiments, which will not be repeated here.

[0144] S406: Based on a preset number of human body key point coordinates, predict various human body attributes of the target object according to the first human body attribute prediction model to obtain the first prediction result.

[0145] In this step, the first human attribute prediction model is a prediction model for predicting human attributes based on the pupil distance of the target object. The first prediction result includes: a prediction result based on pupil distance prediction, which is used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on preset conditions, which is used to characterize any one of the second human attributes in the second category of human attributes.

[0146] In one possible implementation, predicting various human attributes of the target object based on a preset number of human key point coordinates and according to a first human attribute prediction model includes the following steps:

[0147] Select any one of the first human attributes from the first category as the current first human attribute;

[0148] Obtain the calculation formula used to calculate the attributes of the current first human body;

[0149] Based on a preset number of human body key point coordinates, and taking the pupil distance as a benchmark, the current first human body attribute is predicted according to the calculation formula. The first prediction result includes: the prediction result of any one of the first human body attributes in the first category of human body attributes, which is obtained based on the pupil distance prediction.

[0150] In one possible implementation, predicting various human attributes of the target object based on a preset number of human key point coordinates and according to the second human attribute prediction model further includes the following steps:

[0151] Obtain big data that can determine any one of the second-class human attributes;

[0152] By analyzing big data, preset conditions can be obtained to determine any one of the second human attributes;

[0153] Select any one of the second type of human attributes as the current second human attribute;

[0154] Retrieve preset sub-conditions that match the current second human body attributes from preset conditions;

[0155] Based on a preset number of human body key point coordinates, the current second human body attribute is predicted according to preset sub-conditions. The second prediction result also includes: the prediction result of any second human body attribute in the second category of human body attributes, which is obtained according to the preset conditions.

[0156] S408: Based on a preset number of human body key point coordinates, predict various human body attributes of the target object according to the second human body attribute prediction model to obtain the second prediction result.

[0157] In this step, the second human attribute prediction model is a preset neural network prediction model, and the second prediction result includes the prediction result of any first human attribute in the first category of human attributes, which is predicted based on the preset neural network prediction model, and the prediction result of any second human attribute in the second category of human attributes, which is predicted based on the preset neural network prediction model.

[0158] In one possible implementation, predicting various human attributes of the target object based on a preset number of human key point coordinates and according to the second human attribute prediction model further includes the following steps:

[0159] Read the coordinates of a preset number of human body key points of the target object;

[0160] The coordinates of a predetermined number of human body key points of the target object are normalized and preprocessed to obtain training data;

[0161] The training data is input into a preset neural network prediction model to obtain and output a second prediction result; wherein, the second prediction result includes a prediction result based on the preset neural network prediction model and used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on the preset neural network prediction model and used to characterize any one of the second human attributes in the second category of human attributes.

[0162] S410: Fuse the first prediction result and the second prediction result to obtain and output the fused prediction result.

[0163] In this step, the first prediction result and the second prediction result are fused to obtain and output the fused prediction result, which includes the following steps:

[0164] Based on the first prediction result, the first weight value pre-configured based on the first prediction result, the second prediction result, and the second weight value pre-configured based on the second prediction result, the first prediction result and the second prediction result are fused together to obtain and output the fused prediction result.

[0165] In one possible implementation, after obtaining and outputting the fused prediction result, the method further includes the following steps:

[0166] The fused prediction results are pushed to the target object's display device; this allows users to view the fused prediction results on the display device at any time, improving the user experience.

[0167] like Figure 5 The diagram illustrates another human attribute prediction method provided in this disclosure, applied to the server side, and specifically includes the following method steps:

[0168] S502: Obtain a preset number of human body key points of the target object.

[0169] For methods for obtaining key points of the human body, please refer to the description of the same or similar parts in the foregoing method embodiments, which will not be repeated here.

[0170] S504: Input a preset number of human key points into the detection model for training, and output the coordinates of the preset number of human key points.

[0171] Based on the method of outputting a preset number of human body key point coordinates of the target object through continuous training, please refer to the description of the same or similar parts in the foregoing method embodiments, which will not be repeated here.

[0172] S506: Read the first type of human body attribute and the second type of human body attribute from the human body attributes of the target object.

[0173] In this step, the first type of human body attributes of the target object read includes at least one of the following: height, chest circumference, lower chest circumference, waist circumference, belly circumference, hip circumference, hip width, hip circumference, shoulder length, arm length, arm circumference, thigh circumference, calf circumference, leg length, head length, length from shoulder to the base of the thumb, vertical length from shoulder to crotch line, and shoulder width.

[0174] The specific indicators used to evaluate any one of the first-category human attributes mentioned above are as follows:

[0175] These 18 indicators are divided into length-based indicators and circumference-based indicators.

[0176] The length indicators include: height, hip width, shoulder length, arm length, leg length, head length, length from shoulder to the base of the thumb, vertical length from shoulder to crotch line, and shoulder width.

[0177] The measurements include: bust, underbust (for women), waist, belly, hips, pelvis, arms, thighs, and calves.

[0178] First, for estimating length metrics, interpupillary distance (IPD) is used as the standard scale in the image. The average IPD for men is between 60 and 73 millimeters, and for women, it is between 53 and 68 millimeters. Therefore, we take 66.5 for men and 60.5 for women as the IPD values. Then, based on the following proportional formula, the nine dimensions of length can be calculated.

[0179] For women, the following formula is used for calculation:

[0180]

[0181] For example, the actual height is calculated as follows: Height in the image x 60.5 / Interpupillary distance in the image. For males, the following formula is used for calculation:

[0182]

[0183] The relationship between the length index calculation and the corresponding keypoint index is shown in Table 1 below, as detailed below:

[0184]

[0185]

[0186] Table 1

[0187] For the estimation of circumference indicators, elliptical models are used to fit the bust, underbust (for women), waist, belly, hip, and groin circumference, while circular models are used to fit the arm, thigh, and calf circumference.

[0188] The formula for the perimeter of the elliptical model is as follows:

[0189] L = 2πb + 4(ab)

[0190] Where a and b represent the lengths of the major and minor semi-axis, respectively.

[0191] Since the prediction algorithm in the prediction method provided in this embodiment uses frontal user images, it can only measure the length of indicators from the front, such as the major axis (d) of the chest circumference. For the minor axis (d), the relationship between the minor axis length and the major axis length (d) is obtained by statistically analyzing 1000 real user data points. Here, the number of real user data points is not specifically limited to 1000; it can be adjusted according to the needs of different application scenarios. Generally, the number of user data points is positively correlated with the accuracy of the prediction result.

[0192] The detailed relationship between the major semi-axis a, the minor semi-axis b, and the major axis d is shown in Table 2 below.

[0193]

[0194] Table 2

[0195] The formula for the circumference of a circular model is as follows:

[0196] L=πd

[0197] Where d represents the diameter length.

[0198] Table 3 below shows the corresponding keypoint index formulas used for hip circumference, thigh circumference, and calf circumference, as detailed below:

[0199] index Corresponding key point index formula arm circumference (L(p16, p24) + L(p60, p68)) / 2 Thigh circumference (L(p31, p42) + L(p42, p53)) / 2 Calf circumference (L(p34, p39) + L(p45, p50)) / 2

[0200] Table 3

[0201] Where, midpoint(p1, p2) represents the coordinates of the midpoint between points p1 and p2, i.e., midpoint

[0202] L(p1, p2) represents the L2 distance between points p1 and p2, which is the Euclidean distance and also the axis length d.

[0203]

[0204] Y(p1, p2) represents the calculation of the perpendicular distance between points p1 and p2.

[0205] Y(p1, p2) = |(p2.y - p1.y)|.

[0206] In this step, the second type of human body attributes of the target object read includes at least one of the following: chest circumference attribute determined by upper chest and lower chest circumference, multiple levels determined by the chest circumference attribute, and the numerical range corresponding to the chest circumference attribute under each level; waist circumference attribute, multiple levels determined by the waist circumference attribute, and the numerical range corresponding to the waist circumference attribute under each level; belly circumference attribute, multiple levels determined by the belly circumference attribute, and the numerical range corresponding to the belly circumference attribute under each level; hip circumference attribute, multiple levels determined by the hip circumference attribute, and the numerical range corresponding to the hip circumference attribute under each level; groin circumference attribute, multiple levels determined by the groin circumference attribute, and the numerical range corresponding to the groin circumference attribute under each level. The numerical range corresponding to the attributes; shoulder length attribute, multiple levels determined by shoulder length, and the numerical range corresponding to shoulder length under each level; arm length attribute, multiple levels determined by arm length, and the numerical range corresponding to arm length under each level; arm circumference attribute, multiple levels determined by arm circumference, and the numerical range corresponding to arm circumference under each level; thigh circumference attribute, multiple levels determined by thigh circumference, and the numerical range corresponding to thigh circumference under each level; calf circumference attribute, multiple levels determined by calf circumference, and the numerical range corresponding to calf circumference under each level; leg length attribute, multiple levels determined by leg length, and the numerical range corresponding to leg length under each level.

[0207] Figure 6 The indicators used to evaluate any one of the first human attributes in the first category are shown.

[0208] Figure 7 Another schematic diagram is shown, illustrating an index used to evaluate any one of the first human attributes in the first category.

[0209] For details of each indicator, please refer to [link / reference]. Figure 6 and Figure 7 As shown, it will not be elaborated further here.

[0210] The 11 specific categories of indicators for the human body are determined using the rules defined above. These rules are summarized from data from 1000 real models.

[0211] The calculation is performed using the following formula.

[0212] C 胸围 = f(upper bust circumference - lower bust circumference);

[0213] C 腰围 = f(waist circumference / hip circumference);

[0214] C 肚腩围 = f(belly circumference / hip circumference);

[0215] C 臀围 = f(waist circumference / hip circumference);

[0216] C 跨围 = f(waist circumference / hip circumference);

[0217] C 肩宽 = f(shoulder length / hip circumference);

[0218] C 臂围 = f(arm circumference / thigh circumference);

[0219] C 大腿围 = f(thigh circumference / leg length);

[0220] C 小腿围 = f(calf circumference / leg length);

[0221] C 腿长 = f(leg length / height);

[0222] C 臂长 = f(length from shoulder to the base of the thumb / vertical length from shoulder to the crotch line);

[0223] Where C represents the category corresponding to each indicator, and f represents the function that transforms the parameters in f into specific categories according to the rules in the table above.

[0224] S508: Based on a preset number of human body key point coordinates, predict various human body attributes of the target object according to the first human body attribute prediction model to obtain the first prediction result.

[0225] In this step, the first human attribute prediction model is a prediction model that predicts human attributes based on the interpupillary distance of the target object.

[0226] The first prediction result includes: a prediction result based on pupil distance prediction, used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on preset conditions, used to characterize any one of the second human attributes in the second category of human attributes.

[0227] For a detailed description of this step, see [link / reference]. Figures 1 to 4 The same or related descriptions will not be repeated here.

[0228] S510: Based on a preset number of human body key point coordinates, predict various human body attributes of the target object according to the second human body attribute prediction model to obtain the second prediction result.

[0229] In this step, the second human attribute prediction model is a preset neural network prediction model.

[0230] The second prediction result includes a prediction result based on a preset neural network prediction model, used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on a preset neural network prediction model, used to characterize any one of the second human attributes in the second category of human attributes.

[0231] For a detailed description of this step, see [link / reference]. Figures 1 to 4 The same or related descriptions will not be repeated here.

[0232] S512: Fuse the first prediction result and the second prediction result to obtain and output the fused prediction result.

[0233] In this step, the first prediction result includes: a prediction result based on pupil distance prediction, used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on preset conditions, used to characterize any one of the second human attributes in the second category of human attributes.

[0234] The second prediction result includes a prediction result based on a preset neural network prediction model, used to characterize any one of the first human attributes in the first category of human attributes, and a prediction result based on a preset neural network prediction model, used to characterize any one of the second human attributes in the second category of human attributes.

[0235] For a detailed description of this step, see [link / reference]. Figures 1 to 4 The same or related descriptions will not be repeated here.

[0236] In this embodiment of the disclosure, a preset number of human key points of the target object are obtained; the preset number of human key points are input into the detection model for training, and the preset number of human key point coordinates are output; based on the preset number of human key point coordinates, the various human attributes of the target object are predicted according to at least one human attribute prediction model, and the prediction results are used to predict the various human attributes of the target object. By employing the embodiments of this application, the size of the sample data obtained under different application scenarios can be effectively combined to select the human attribute prediction model that matches the current application scenario, thus effectively improving prediction efficiency. Furthermore, compared to existing detection methods that can only detect individual human skeletal key points (i.e., the detection result only displays the two-dimensional contour of the target object), the embodiments of this application, because the output result includes the coordinates of a preset number of human contour key points, include not only the coordinates of all contour points that can fully cover the entire contour of the target object, but also the coordinates of all contour points that can cover the facial features of the target object. Thus, the detection result can display the three-dimensional contour of the target object, providing a sense of depth. Moreover, the detection method provided by this invention, when an image includes multiple objects, and all of these objects are target objects, can simultaneously identify multiple target objects and synchronously detect the coordinates of a preset number of human contour key points for each of the multiple target objects, thereby improving detection efficiency.

[0237] The following are embodiments of the human attribute prediction device of this disclosure, which can be used to execute the human attribute prediction method of this disclosure. For details not disclosed in the embodiments of the human attribute prediction device of this disclosure, please refer to the embodiments of the human attribute prediction method of this disclosure.

[0238] Please see Figure 9 This diagram illustrates the structure of a human attribute prediction device provided in an exemplary embodiment of the present invention. This human attribute prediction device can be implemented as all or part of a terminal through software, hardware, or a combination of both. The human attribute prediction device includes an acquisition unit 902, a training unit 904, and a prediction unit 906.

[0239] Specifically, the acquisition unit 902 is used to acquire a preset number of human body key points of the target object;

[0240] The training unit 904 is used to input a preset number of human key points acquired by the acquisition unit 902 into the detection model for training, and output the coordinates of the preset number of human key points.

[0241] The prediction unit 906 is used to predict various human attributes of the target object based on a preset number of human key point coordinates output by the training unit 904 and according to at least one human attribute prediction model, and obtain prediction results. The prediction results are used to predict various human attributes of the target object.

[0242] Optionally, the human attribute prediction model includes a first human attribute prediction model and a second human attribute prediction model. The first human attribute prediction model is a prediction model for predicting human attributes based on the interpupillary distance of the target object, and the second human attribute prediction model is a preset neural network prediction model. The prediction unit 906 is used for:

[0243] Based on a preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a first prediction result; or...

[0244] Based on a preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a second prediction result; or...

[0245] Based on a preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model respectively, and the corresponding prediction results are obtained. The two prediction results are then fused to obtain the corresponding prediction result.

[0246] Optionally, prediction unit 906 is specifically used for:

[0247] Select any one of the first human attributes from the first category as the current first human attribute;

[0248] Obtain the calculation formula used to calculate the attributes of the current first human body;

[0249] Based on a preset number of human body key point coordinates, and taking the pupil distance as a benchmark, the current first human body attribute is predicted according to the calculation formula to obtain the first prediction result; wherein, the first prediction result includes: the prediction result of any one of the first human body attributes in the first category of human body attributes obtained based on the pupil distance prediction.

[0250] Optionally, the prediction unit 906 is also specifically used for:

[0251] Obtain big data that can determine any one of the second-class human attributes;

[0252] By analyzing big data, preset conditions can be obtained to determine any one of the second human attributes;

[0253] Select any one of the second type of human attributes as the current second human attribute;

[0254] Obtain preset sub-conditions that match the current second human body attributes from preset conditions;

[0255] Based on a preset number of human body key point coordinates, the current second human body attributes are predicted according to preset sub-conditions to obtain the second prediction result.

[0256] Optionally, the prediction unit 906 is also specifically used for:

[0257] Based on a preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, and the first prediction result is obtained.

[0258] Based on a preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, and a second prediction result is obtained.

[0259] The first and second prediction results are fused to obtain and output the fused prediction result.

[0260] Optionally, the prediction unit 906 is also specifically used for:

[0261] Read the coordinates of a preset number of human body key points of the target object;

[0262] The coordinates of a predetermined number of human body key points of the target object are normalized and preprocessed to obtain training data;

[0263] The training data is input into a preset neural network prediction model to obtain and output a second prediction result.

[0264] Optionally, the prediction unit 906 is also specifically used for:

[0265] Based on the first prediction result, the first weight value pre-configured based on the first prediction result, the second prediction result, and the second weight value pre-configured based on the second prediction result, the first prediction result and the second prediction result are fused together to obtain and output the fused prediction result.

[0266] It should be noted that the human attribute prediction device provided in the above embodiments is only illustrated by the division of the above functional units when executing the human attribute prediction method. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. In addition, the human attribute prediction device and the human attribute prediction method embodiments provided in the above embodiments belong to the same concept, and the implementation process can be found in the human attribute prediction method embodiments, which will not be repeated here.

[0267] In this embodiment of the disclosure, the acquisition unit is used to acquire a preset number of human key points of the target object; the training unit is used to input the preset number of human key points acquired by the acquisition unit into the detection model for training, and output the preset number of human key point coordinates; the prediction unit is used to predict various human attributes of the target object based on the preset number of human key point coordinates output by the training unit and according to at least one human attribute prediction model, and obtain prediction results, which are used to predict various human attributes of the target object. By employing the embodiments of this application, the size of the sample data obtained under different application scenarios can be effectively combined to select the human attribute prediction model that matches the current application scenario, thus effectively improving prediction efficiency. Furthermore, compared to existing detection methods that can only detect individual human skeletal key points (i.e., the detection result only displays the two-dimensional contour of the target object), the embodiments of this application, because the output result includes the coordinates of a preset number of human contour key points, include not only the coordinates of all contour points that can fully cover the entire contour of the target object, but also the coordinates of all contour points that can cover the facial features of the target object. Thus, the detection result can display the three-dimensional contour of the target object, providing a sense of depth. Moreover, the detection method provided by this invention, when an image includes multiple objects, and all of these objects are target objects, can simultaneously identify multiple target objects and synchronously detect the coordinates of a preset number of human contour key points for each of the multiple target objects, thereby improving detection efficiency.

[0268] like Figure 10 As shown, this embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the method steps described above.

[0269] This disclosure provides a storage medium storing computer-readable instructions, on which a computer program is stored, and the program is executed by a processor to implement the method steps described above.

[0270] The following is for reference. Figure 10 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0271] like Figure 10 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0272] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0273] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0274] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0275] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0276] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0277] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0278] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not, in some cases, intended to limit the specific unit.

Claims

1. A method for predicting human attributes, characterized in that, The method includes: Obtain a preset number of key human body points of the target object, wherein the preset number can completely cover the key parts corresponding to the outer contour of the entire body of the target object, and fully cover the facial features of the target object; The preset number of human key points are input into the detection model for training, and the preset number of human key point coordinates are output. Based on the preset number of human body key point coordinates, and according to at least one human body attribute prediction model, the various human body attributes of the target object are predicted to obtain prediction results, which are used to predict the various human body attributes of the target object. The human attribute prediction model includes a first human attribute prediction model and a second human attribute prediction model. The first human attribute prediction model is a prediction model that predicts human attributes based on the interpupillary distance of the target object, and the second human attribute prediction model is a preset neural network prediction model. Based on the preset number of human body key point coordinates, and according to at least one human body attribute prediction model, the prediction results for various human body attributes of the target object include: Based on the preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a first prediction result; or... Based on the preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a second prediction result; or... Based on the preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model respectively, and the corresponding prediction results are obtained. The two prediction results are then fused to obtain the corresponding prediction result. The indicators used to evaluate any one of the first human body attributes are divided into length-based indicators and girth-based indicators. For the estimation of the length indicator, the interpupillary distance is used as the standard scale in the image, which includes multiple objects, and all of the multiple objects are target objects. The method further includes: Based on the size of the sample data to be predicted by the human attribute prediction model in different application scenarios, select the human attribute prediction model that matches the current application scenario. The step of inputting the preset number of human key points into the detection model for training and outputting the coordinates of the preset number of human key points includes: At least one target object with a preset number of human contour key points is input into a preset detection model for training, and the network output results in multiple dimensions are output. The network output results in multiple dimensions include a first network output result that represents the offset value of the preset number of human contour key points based on the center point, a second network output result that represents the feature map of the preset number of human contour key points, and a third network output result that represents the offset value of multiple human contour key points in the preset number. Based on the output results of the first network, the second network, and the third network, the coordinates of a preset number of key points of the human body contour are determined and output.

2. The method according to claim 1, characterized in that, The first prediction result obtained by predicting various human attributes of the target object based on the preset number of human key point coordinates and according to the first human attribute prediction model, includes: Select any one of the first human attributes from the first category as the current first human attribute; Obtain the calculation formula used to calculate the attributes of the current first human body; Based on the preset number of human body key point coordinates, and taking the pupil distance as a reference, the current first human body attribute is predicted according to the calculation formula to obtain the first prediction result; wherein, the first prediction result includes: a prediction result obtained based on the pupil distance and used to characterize any one of the first human body attributes in the first type of human body attributes.

3. The method according to claim 1, characterized in that, The second prediction result, obtained by predicting various human attributes of the target object based on the preset number of human key point coordinates and according to the second human attribute prediction model, includes: Obtain big data that can determine any one of the second-class human attributes; By analyzing the big data, preset conditions can be obtained to determine any one of the second human attributes; Select any one of the second type of human attributes as the current second human attribute; Obtain preset sub-conditions that match the current second human body attribute from the preset conditions; Based on the preset number of human body key point coordinates, the current second human body attribute is predicted according to the preset sub-conditions to obtain the second prediction result.

4. The method according to claim 1, characterized in that, The process of predicting various human attributes of the target object based on the preset number of human key point coordinates and the second human attribute prediction model to obtain the second prediction result further includes: Read the coordinates of a preset number of key human body points of the target object; The coordinates of a preset number of human body key points of the target object are normalized and preprocessed to obtain training data; The training data is input into the preset neural network prediction model to obtain and output the second prediction result.

5. The method according to claim 1, characterized in that, The process of fusing the first prediction result and the second prediction result to obtain and output the fused prediction result includes: Based on the first prediction result, the first weight value preset based on the first prediction result, the second prediction result, and the second weight value preset based on the second prediction result, the first prediction result and the second prediction result are fused together to obtain and output the fused prediction result.

6. A human attribute prediction device, characterized in that, The device includes: The acquisition unit is used to acquire a preset number of human body key points of the target object, wherein the preset number can completely cover the key parts corresponding to the outer contour of the entire body of the target object, and fully cover the facial features of the target object; The training unit is used to input the preset number of human key points acquired by the acquisition unit into the detection model for training, and output the coordinates of the preset number of human key points. The prediction unit is used to predict various human attributes of the target object based on the preset number of human key point coordinates output by the training unit and according to at least one human attribute prediction model, and to obtain prediction results. The prediction results are used to predict various human attributes of the target object. The human attribute prediction model includes a first human attribute prediction model and a second human attribute prediction model. The first human attribute prediction model is a prediction model that predicts human attributes based on the interpupillary distance of the target object, and the second human attribute prediction model is a pre-defined neural network prediction model. The prediction unit is used for: Based on a preset number of human body key point coordinates, the first human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a first prediction result; or... Based on a preset number of human body key point coordinates, the second human body attribute prediction model is used to predict various human body attributes of the target object, resulting in a second prediction result; or... Based on a preset number of human body key point coordinates, the various human body attributes of the target object are predicted according to the first human body attribute prediction model and the second human body attribute prediction model respectively, and the corresponding prediction results are obtained. The two prediction results are then fused to obtain the corresponding prediction result. The indicators used to evaluate any one of the first human body attributes are divided into length-based indicators and girth-based indicators. For the estimation of the length indicator, the interpupillary distance is used as the standard scale in the image, which includes multiple objects, and all of the multiple objects are target objects. The device is also used for: Based on the size of the sample data to be predicted by the human attribute prediction model in different application scenarios, select the human attribute prediction model that matches the current application scenario. The training unit is further used for: At least one target object with a preset number of human contour key points is input into a preset detection model for training, and the network output results in multiple dimensions are output. The network output results in multiple dimensions include a first network output result that represents the offset value of the preset number of human contour key points based on the center point, a second network output result that represents the feature map of the preset number of human contour key points, and a third network output result that represents the offset value of multiple human contour key points in the preset number. Based on the output results of the first network, the second network, and the third network, the coordinates of a preset number of key points of the human body contour are determined and output.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Human physiological-parameter detecting method and device, storage medium and system

    CN108511069A

  • User sign determination method and device, equipment and storage medium

    CN109409348A

  • Face attribute recognition method and device, computer equipment and storage medium

    CN111507285A