A method, device, electronic device and storage medium for detecting human body contour points

By training preset detection models and labeling tools, the coordinates of key points of human contours are output, and the problem of two-dimensional contours in existing detection methods is solved, the generation of detailed human contours and the identification of multi-object objects is realized, and the detection efficiency and user experience are improved.

CN114120352BActive Publication Date: 2025-07-11SO-YOUNG INT INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010874022.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-26
Publication Date
2025-07-11
Estimated Expiration
2040-08-26

AI Technical Summary

Technical Problem

The existing human contour detection methods can only detect key points of the user's bones, and cannot accurately obtain the contour point coordinates of various parts of the user's human body, resulting in the detection result being a two-dimensional contour, which has a low user experience.

Method used

By obtaining the image of the target object, using the preset detection model training and labeling tool to output the coordinates of the preset number of human contour key points, including the whole body and facial features contour points, to generate detailed human contours.

Benefits of technology

Accurate detection of human contours is realized, detailed human attribute information is generated, detection efficiency and user experience are improved, and contour points of multiple target objects can be identified at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120352B_ABST
    Figure CN114120352B_ABST
Patent Text Reader

Abstract

The present invention discloses a human contour point detection method, device, electronic device, and storage medium. The method includes: inputting a preset number of human contour key points of at least one target object into a preset detection model for training, and outputting a detection result to generate the human contour of the target object according to the detection result. Compared with the existing detection methods that can only detect each human bone key point, by adopting the embodiment of the present application, the human contour of the target object can be generated from the detection result, and the human attribute information can be generated from the labeled human contour of the target object; in addition, when there are multiple objects in the picture and all the multiple objects are target objects, multiple target objects can be recognized simultaneously, and the coordinates of the preset number of human contour key points of each of the multiple target objects can be detected synchronously, improving the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, electronic device and storage medium for detecting human contour points. Background Art

[0002] With the development of artificial intelligence technology, artificial intelligence technology is widely used. For example, when an image of a user is captured by an image acquisition device, the key points of the user's human skeleton can be given according to the user's image. However, the detection results obtained based on this detection method are often not accurate enough. The detection results only include the coordinates of multiple key points associated with the user's skeleton, and cannot give the coordinates of more and more detailed human contour points distributed in various parts of the user's body.

[0003] In a certain application scenario, the user remotely waits to customize high-end clothing or accessories. Therefore, it is necessary to accurately obtain multiple human contour points of the current user and the respective coordinates corresponding to the multiple human contour points. In this way, the number of contour points in the detection results obtained by the existing detection methods is too small, and the obtained contour points are mostly the skeleton contour points of the user. In this way, the detection results can only show the two-dimensional contour of the user, and the user experience is low. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, electronic device and storage medium for detecting human contour points for the problem that the detection results obtained based on the existing detection methods can only show the two-dimensional contour of the user and the user experience is low.

[0005] In a first aspect, an embodiment of the present application provides a method for detecting human contour points, and the method includes:

[0006] Obtain a picture including at least a target object;

[0007] Determine at least one target object from at least one object in the picture;

[0008] Input the human contour key points of the at least one target object in a preset number into a preset detection model, and output a detection result to generate the human contour of the target object according to the detection result. The detection result includes the coordinates of multiple key points for determining the boundary of the detection frame and the coordinates of the human contour key points in the preset number.

[0009] In an implementation manner, the inputting the human contour key points of the at least one target object in a preset number into the preset detection model for training includes:

[0010] Obtain a data set, where the data set includes the preset number of first human body contour key points, and the first human body contour key points are human body contour key points marked based on a marking tool;

[0011] Store the data set in a file in a preset format.

[0012] In one implementation, the step of inputting the preset number of human body contour key points of at least one target object into a preset detection model for training further includes:

[0013] Use the picture and the file in the preset format as training samples, input them into the preset detection model for training, and output a prediction result, where the prediction result includes the preset number of second human body contour key points, and the second human body contour key points are predicted human body contour key points;

[0014] Calculate the loss function of the preset detection model according to the preset number of first human body contour key points and the preset number of second human body contour key points;

[0015] Iteratively train the training samples to obtain the preset detection model, and output a weight file including the preset detection model.

[0016] In one implementation, the step of inputting the preset number of human body contour key points of at least one target object into a preset detection model for training and outputting a detection result includes:

[0017] Input the preset number of human body contour key points of at least one target object into the preset detection model for training, and output network output results in multiple dimensions;

[0018] The network output results in multiple dimensions include a first network output result for characterizing the offset value of the preset number of human body contour key points based on the center point, a second network output result for characterizing the feature map of the preset number of human body contour key points, and a third network output result for characterizing the offset value of multiple human body contour key points among the preset number;

[0019] Determine and output the coordinates of the preset number of human body contour key points according to the first network output result, the second network output result, and the third network output result.

[0020] In one implementation, after determining and outputting the coordinates of the preset number of human body contour key points, the method further includes:

[0021] Read the first coordinate for characterizing the first key point of the nose of the target object and the second coordinate for characterizing the second key point of the crotch line of the target object;

[0022] Set an included angle, where the included angle is the angle between the line connecting the first key point determined by the first coordinate and the second key point determined by the second coordinate and the vertical perpendicular line;

[0023] If the included angle is greater than a preset included angle threshold, then determine that the current body posture of the target object is inclined; or,

[0024] Read the coordinates of any one of the preset number of body contour key points;

[0025] If the area corresponding to the coordinates of any one of the key points is in a non-picture area, then determine that the target object is in a state where the body is blocked.

[0026] In one implementation, each coordinate for characterizing each body contour key point of the target object further includes a third coordinate and a fourth coordinate. The third coordinate is the coordinate of the third key point for characterizing the first shoulder of the target object, and the fourth coordinate is the coordinate of the fourth key point for characterizing the second shoulder of the target object. After determining and outputting the coordinates of the preset number of body contour key points, the method further includes:

[0027] Read the ordinate in the first coordinate, the ordinate in the third coordinate, and the ordinate in the fourth coordinate;

[0028] According to the ordinate in the third coordinate and the ordinate in the fourth coordinate, calculate the height of the shoulders of the target object to obtain the average value of the shoulder height of the target object;

[0029] According to the ordinate in the first coordinate and the average value of the shoulder height of the target object, judge the current body posture of the target object. If the height value corresponding to the ordinate in the first coordinate is greater than the average value of the shoulder height of the target object, judge that the current body posture of the target object is bending down.

[0030] In one implementation, the determining at least one target object from the picture includes:

[0031] Obtain multiple objects in the picture;

[0032] Obtain the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object;

[0033] Calculate the scores of each of the multiple objects respectively according to the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object;

[0034] Determine at least one object with the highest score among the multiple objects as the target object.

[0035] In one implementation, the detection result further includes the coordinates of multiple key points for determining the boundary of the detection box. Inputting the preset number of human contour key points of at least one target object into the preset detection model for training, and the output of the detection result further includes:

[0036] Input the preset number of human contour key points of at least one target object into the preset detection model for training, and output network output results of multiple dimensions. The network output results of multiple dimensions further include a fourth network output result corresponding to a feature map for characterizing the category of the target object, a fifth network output result for characterizing the size of the target object, and a sixth network output result corresponding to the offset value of the center point of the target object;

[0037] Determine and output the coordinates of multiple key points for determining the boundary of the detection box according to the fourth network output result, the fifth network output result, and the sixth network output result.

[0038] In a second aspect, an embodiment of the present application provides a human contour point detection device, and the device includes:

[0039] An acquisition module for acquiring a picture including at least a target object;

[0040] A determination module for determining at least one target object from at least one object in the picture acquired by the acquisition module;

[0041] A processing module for inputting the preset number of human contour key points of at least one target object determined by the determination module into the preset detection model, outputting a detection result, and generating a human contour of the target object according to the detection result. The detection result includes the coordinates of multiple key points for determining the boundary of the detection box and the coordinates of the preset number of human contour key points.

[0042] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the method steps as described above.

[0043] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method steps as described above.

[0044] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0045] In the embodiment of the present application, a picture including a target object is obtained; at least one target object is determined from the picture; a preset number of human body contour key points of at least one target object are input into a preset detection model for training, and a detection result is output to generate a human body contour of the target object according to the detection result, and the detection result includes the coordinates of a preset number of human body contour key points. Compared with the existing detection methods that can only detect each human bone key point, that is, the detection result can only show the two-dimensional contour of the target object. However, in the embodiment of the present application, since the output result includes the coordinates of a preset number of human body contour key points, the coordinates of the preset number of human body contour key points not only include the coordinates of each contour point that can fully cover the whole body contour of the target object, but also include the coordinates of each contour point that can cover the facial features contour of the target object. In this way, a human body contour of the target object can be generated from the detection result, and human body attribute information can be generated from the marked human body contour of the target object. In addition, the detection method provided by the present invention can simultaneously identify multiple target objects and synchronously detect the coordinates of a preset number of human body contour key points of each of the multiple target objects when there are multiple objects in the picture and all the multiple objects are target objects, thereby improving the detection efficiency. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings

[0046] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0047] Figure 1 is an application scenario diagram of a human body contour point detection method provided by an embodiment of the present disclosure;

[0048] Figure 2 is a flowchart of a human body contour point detection method provided by an embodiment of the present disclosure;

[0049] Figure 3 is a diagram of human body contour points of a target object in a specific application scenario provided by an embodiment of the present disclosure;

[0050] Figure 4 is a diagram of determining the coordinates of a detection frame provided by an embodiment of the present disclosure;

[0051] Figure 5 It is a schematic flowchart of another human body contour point detection method provided by an embodiment of the present disclosure;

[0052] Figure 6 It is a schematic flowchart of another human body contour point detection method provided by an embodiment of the present disclosure;

[0053] Figure 7 It is a schematic flowchart of another human body contour point detection method provided by an embodiment of the present disclosure;

[0054] Figure 8 It is a schematic flowchart of another human body contour point detection method provided by an embodiment of the present disclosure;

[0055] Figure 9 It is a schematic diagram of a detection frame and each human body contour key point marked in a specific application scenario provided by an embodiment of the present disclosure;

[0056] Figure 10 It is a schematic structural diagram of a human body contour point detection device provided by an embodiment of the present disclosure;

[0057] Figure 11 It shows a schematic diagram of the connection structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0058] The following description and drawings fully illustrate the specific embodiments of the present invention so that those skilled in the art can practice them.

[0059] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0060] The optional embodiments of the present disclosure will be described in detail below with reference to the drawings.

[0061] As Figure 1 shown, it is an application scenario diagram of an embodiment of the present disclosure. In this application scenario, multiple users operate a client installed on the terminal device such as a mobile phone, and the client communicates with the background server through the network for data. A particular application scenario is the detection process of the human body contour points of at least one target object, but it is not limited to this unique application scenario. It can be understood that any scenario that can be applied to this implementation scheme is included. For the convenience of description, this embodiment takes the application scenario of detecting the human body contour points of a target object in a picture as an example for description. As Figure 1 shown, the obtained picture including at least the target object often comes from one of the multiple clients as Figure 1 shown.

[0062] As Figure 2 shown, an embodiment of the present disclosure provides a method for detecting human contour points, which is applied to the server side and specifically includes the following method steps:

[0063] S202: Obtain a picture including a target object.

[0064] In practical applications, in order to obtain multiple pictures including different objects, pictures of multiple objects can be obtained from multiple clients as Figure 1 shown. In addition, in order to make the detection results obtained by the detection method more accurate, the pictures selected are often full-body pictures that can completely cover at least one object, and the full-body pictures are pictures that can clearly show the facial contour of the object.

[0065] S204: Determine at least one target object from the picture.

[0066] In practical applications, there are often multiple objects in a picture. At this time, it is necessary to determine which object is the target object or which ones are the target objects from multiple objects.

[0067] Specifically, determining at least one target object from at least one object in the picture includes the following steps:

[0068] Obtain multiple objects in the picture;

[0069] Obtain the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object;

[0070] According to the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object, calculate the score of each object in multiple objects respectively;

[0071] Determine at least one object with the highest score among multiple objects as the target object.

[0072] In a certain application scenario, the rule for determining at least one target object is pre-configured as: the target object is one of the target objects that occupy the most pixels in the picture, and the prediction score of the target object is also relatively high; in this way, using the above rule to calculate the score of each object in the picture, the one with the highest score is the determined target object.

[0073] The specific calculation formula is:

[0074] score = 0.5 * (person width * person height ) / picwidth *pic height +0.5*score person 。

[0075] Through the above calculation formula, at least one target object can be accurately determined from the picture.

[0076] When the scores of two or more objects in the picture are relatively close, it is determined that two or more objects in the picture are all target objects. Therefore, the detection method provided by the embodiments of the present disclosure can detect the human body contour key points of multiple objects in the picture at the same time and output the coordinates of the human body contour points of multiple different target objects; in this way, the detection efficiency is greatly improved.

[0077] S206: Input the human body contour key points of at least one target object with a preset number into a preset detection model for training, and output a detection result to generate the human body contour of the target object according to the detection result. The detection result includes the coordinates for the human body contour key points with a preset number. Compared with the existing detection methods that can only detect each human body bone key point, that is, the detection result can only show the two-dimensional contour of the target object. However, in the embodiments of the present application, since the output result includes the coordinates of the human body contour key points with a preset number, the coordinates of the human body contour key points with a preset number not only include the coordinates of each contour point that can completely cover the whole body contour of the target object, but also include the coordinates of each contour point that can cover the facial features contour of the target object. In this way, the human body contour of the target object can be generated from the detection result, and the human body attribute information can be generated from the labeled human body contour of the target object.

[0078] It should be noted that in this step, the human body attribute information can be the information used to represent the following human body attributes: height information, chest circumference information, lower chest circumference information, waist circumference information, belly circumference information, hip circumference information, hip width information, hip circumference information, shoulder length information, arm length information, arm circumference information, thigh circumference information, calf circumference information, leg length information, head length information, length information from the shoulder to the tiger's mouth, vertical length information from the shoulder to the crotch line, and shoulder width information.

[0079] In addition, the detection method provided by the present invention can identify multiple target objects at the same time and synchronously detect the coordinates of the human body contour key points with a preset number of each of the multiple target objects when there are multiple objects in the picture and all the multiple objects are target objects, thereby improving the detection efficiency.

[0080] In this step, after determining at least one target object, during the process of detecting the human contour points of at least one target object, assuming that the number of target objects to be detected currently is one, for this target object, there is no specific limit on the preset number of human contour key points. As long as this preset number can: fully cover the key parts corresponding to the whole body outline of this target object, and comprehensively cover the facial features of this target object. In practical applications, the preset number can be adjusted according to the needs of different application scenarios, and no specific limit is imposed on the preset number here.

[0081] In a possible implementation manner, inputting the human contour key points with a preset number of at least one target object into a preset detection model for training includes the following steps:

[0082] Obtain a data set, where the data set includes the first human contour key points with a preset number. The first human contour key points are the human contour key points marked based on a marking tool. Among them, the marking tool can be the labelme marking tool. Labelme is a common image marking tool. Through this marking tool, users can create customized marking tasks or perform image marking;

[0083] Store the data set in a file with a preset format. Among them, the preset format can be a file in json (JavaScript Object Notation, JS object notation) format. Json is a lightweight data exchange format. It uses a text format that is completely independent of programming languages to store and represent data. In practical applications, the simple and clear hierarchical structure makes json an ideal data exchange language. Json is easy for people to read and write, and is also easy for machines to parse and generate, and effectively improves the network transmission efficiency.

[0084] In a possible implementation manner, inputting the human contour key points with a preset number of at least one target object into a preset detection model for training further includes the following steps:

[0085] Use the picture and the file with the preset format as training samples, input them into the preset detection model for training, and output a prediction result. The prediction result includes the second human contour key points with a preset number. The second human contour key points are the predicted human contour key points;

[0086] Calculate the loss function of the preset detection model according to the first human contour key points with a preset number and the second human contour key points with a preset number. Among them, through this loss function, the neural network corresponding to this preset detection model is made to converge;

[0087] Iteratively train the training samples to obtain a preset detection model, and output a weight file including the preset detection model.

[0088] It should be noted that in practical applications, before inputting the preset number of human body contour key points of at least one target object into the preset detection model for training, a batch of training pictures are prepared. Each picture includes the target object, and it is ensured that the entire body of the target object in the picture is not blocked by external obstacles, that is, the whole body contour of the target object can be clearly recognized.

[0089] As Figure 3 shown, it is a schematic diagram of the human body contour points of the target object in the specific application scenario provided by the embodiments of the present disclosure. For the convenience of explanation and illustration, Figure 3 instead of using pictures of real people, pictures of models are used here, just for example. As Figure 3 shown, the total number of human body contour points of the target object marked is 77 points. Specifically, the part represented by each point and the number represented by each point can be clearly seen in Figure 3 and will not be elaborated here. Through the marking of the 77 contour key points as Figure 3 shown, the precise positioning of the facial features and the human body contour line of the target object can be achieved. Among them, the facial features of the target object include the eyes, nose, and mouth of the target object.

[0090] In this step, the preset detection model is constructed based on a preset detection algorithm. Here, the preset detection algorithm in the detection method of the embodiments of the present disclosure is described as follows:

[0091] The preset detection algorithm uses a detection framework such as CenterNet. On the basis of the original CenterNet, the basic network structure is replaced with EfficientNet-B0, which can achieve a good balance between accuracy and speed. The overall network structure is composed of EfficientNet-B0 plus a deconvolution module. The EfficientNet-B0 part consists of a normal conv plus 7 MBconvs. The structure of the MBconv is as follows, where Conv represents convolution, Batchnorm represents normalization operation, DepthwiseConv represents depthwise separable convolution, Swish represents an activation function, drop_connect represents a module that randomly deactivates neuron connections, which can increase the generalization ability of the model, Average pooling represents average pooling, and sigmoid represents an activation function.

[0092] In the EfficientNet-B0 part, the size of the input image is 512*512*3 (height, width, channel). After 5 downsampling (processing to reduce width and height) operations, the size of the output feature map of this module is finally 16*16*320 (512 / 2^5 = 16). Then, through 3 transposed convolution modules, each transposed convolution upsamples (processing to increase width and height) the input feature map once. Finally, the size of the output feature map after 3 transposed convolutions is 128*128*256. On this basis, through a 1*1 convolution processing, the final 6 network output results are obtained.

[0093] In a possible implementation, input the preset number of human contour key points of at least one target object into a preset detection model for training, and the steps for outputting the detection result include:

[0094] Input the preset number of human contour key points of at least one target object into a preset detection model for training, and output network output results in multiple dimensions. The network output results in multiple dimensions include a first network output result used to represent the offset value of the preset number of human contour key points based on the center point, a second network output result used to represent the feature map of the preset number of human contour key points, and a third network output result used to represent the offset value of multiple human contour key points among the preset number.

[0095] According to the first network output result, the second network output result, and the third network output result, determine and output the coordinates of the preset number of human contour key points.

[0096] In a specific application scenario, the above three network output results can be specifically:

[0097] Hps (corresponding to the first network output result): The offset of the key points of the object based on the center point, outputting k*2 channels, where k represents the number of key points. Assume that in a certain application scenario, the human contour key points of the target object are 77 human contour key points, and 2 represents the two offsets of x and y. The final size is 77*2 = 154.

[0098] hm_hp (corresponding to the second network output result): The feature map of the object key points, outputting k channels, one channel for one point. The final size is 128*128*77.

[0099] hp_offset (corresponding to the third network output result): The offset of K key points, and all these key points use the same offset amount, with two indicators of x and y, outputting 2 channels. The final size is 128*128*2.

[0100] In a possible implementation, the detection result further includes the coordinates of multiple key points for determining the boundary of the detection box. Inputting the preset number of human contour key points of at least one target object into a preset detection model for training, and the steps for outputting the detection result further include the following:

[0101] Input the preset number of human contour key points of at least one target object into a preset detection model for training, and output network output results in multiple dimensions. The network output results in multiple dimensions further include a fourth network output result corresponding to a feature map for characterizing the category of the target object, a fifth network output result for characterizing the size of the target object, and a sixth network output result corresponding to the offset value of the center point of the target object;

[0102] Determine and output the coordinates of multiple key points for determining the boundary of the detection box according to the fourth network output result, the fifth network output result, and the sixth network output result.

[0103] In a specific application scenario, the above three network output results can be specifically:

[0104] Hm (corresponding to the fourth network output result): The feature map of the object, one channel (thickness) for one category. Since the detection method in the embodiments of the present disclosure only detects one category of human body, the size of this layer is 128 * 128 * 1.

[0105] Wh (corresponding to the fifth network output result): The size of the object, that is, width and height. Therefore, it is 2 channels, and the final size is 128 * 128 * 2.

[0106] Reg (corresponding to the sixth network output result): The offset of the center point of the object, including two offset amounts of x and y. Therefore, it is 2 channels, and the final size is 128 * 128 * 2.

[0107] In a specific application scenario, after obtaining the above six network output results, use the following formula to convert them into the final coordinates. As Figure 4 shown, it is a schematic diagram for determining the coordinates of the detection box provided by the embodiments of the present disclosure;

[0108] (x i + δx i - w i / 2, y i + δy i - h i / 2,

[0109] x i + δx i + w i / 2, y i + δyi +h i / 2)

[0110] Among them, (xi, yi) represents the predicted center point x and y coordinates, that is, the hm branch.

[0111] (δxi, δyi) represents the offset of the predicted center point coordinates x and y, that is, the reg branch.

[0112] (wi, hi) represents the size width and height of the object box.

[0113] By adding the offset of the center point coordinates in the x and y directions to the center point coordinates, the accurate center point coordinates are obtained. Then, by subtracting half of the width and height from the center point coordinates respectively, the upper left coordinate point (x1, y1) of the human target box is obtained, and by adding half of the width and height to the center point coordinates respectively, the lower right coordinate point (x2, y2) of the human target box is obtained.

[0114] It is transformed into the final key point coordinates using the following formula, as described below:

[0115]

[0116] Among them, (x, y) represents the coordinates of the center point of the target object, that is, the hm branch. J xyj represents the offset of the (x, y) coordinates of the j-th point, that is, the hps branch. However, the error of the key point coordinates predicted by such a direct regression-based method is relatively large.

[0117] L j = hm_hp + hp_offset

[0118] According to direct regression, the approximate position of each contour point can be determined, and then the point Lj(hm_hp) with the largest confidence (0-1 score) greater than 0.1 near this position is found and used as the true contour point; then, the coordinates of the final contour point are obtained by adding the offset (hp_offset) to this point. It should be noted here that if there are more than two points greater than 0.1, the point closest to lj is taken. The formula is as follows:

[0119]

[0120] It should be noted that lj represents the coordinates of the key points obtained by the regression-based method, that is, the hps branch. Lj represents the coordinates of the key points obtained by the feature map-based method. There are a total of 77 key points here. For each of the 77 regression points lj, the distance is calculated from the corresponding feature map branch Lj, and the point in Lj with the closest distance is used as the final predicted output key point. For example, based on the regression, the key point position of the eye position is predicted, and then in the feature map of the channel where the eye is located, the point with the closest distance and a confidence level higher than 0.1 is found as the final key point of the eye position.

[0121] As Figure 5 shown, it is another human body contour point detection method provided by an embodiment of the present disclosure, which is applied to the server side and specifically includes the following method steps:

[0122] S502: Obtain a picture including the target object.

[0123] In practical applications, in order to obtain multiple pictures including different objects, pictures of multiple objects can be obtained from multiple clients as Figure 1 shown. In addition, in order to make the detection results obtained by the detection method more accurate, the pictures selected are usually full-body pictures that can fully cover at least one object, and the full-body pictures are pictures that can clearly show the facial features contours of the object.

[0124] S504: Determine at least one target object from the picture.

[0125] It should be noted that for the specific description of the step of determining at least one target object from the picture, please refer to the description of the same or similar parts above, and will not be repeated here.

[0126] S506: Input the human body contour key points of at least one target object with a preset number into a preset detection model for training, and output the detection result, so as to generate the human body contour of the target object according to the detection result. The detection result includes the coordinates of the human body contour key points with a preset number.

[0127] It should be noted that for the specific description of the step of inputting the human body contour key points of at least one target object with a preset number into a preset detection model for training and outputting the detection result, please refer to the description of the same or similar parts above, and will not be repeated here.

[0128] S508: Read the first coordinate of the first key point for representing the nose of the target object and the second coordinate of the second key point for representing the crotch line of the target object.

[0129] This step is completed after determining and outputting the coordinates of a preset number of key points of the human body contour. Among them, the first coordinate includes the abscissa and ordinate of the first key point, and the second coordinate also includes the abscissa and ordinate of the second key point.

[0130] S510: Set an included angle, which is the included angle between the line connecting the first key point determined by the first coordinate and the second key point determined by the second coordinate and the vertical perpendicular line. By reading the value of this included angle, it can be judged whether the current body posture of the target object is inclined and the degree of inclination.

[0131] In a schematic diagram such as Figure 3 The first key point is the key point numbered 3, whose coordinates correspond to the first coordinate, and this number 3 is used to indicate the nose of the target object; the second key point is the key point numbered 42, whose coordinates correspond to the second coordinate, and this number 42 is used to indicate the crotch line of the target object. Here is just an example, and other numbers can also be used to indicate the nose of the target object or other numbers to indicate the crotch line of the target object.

[0132] S512: If the included angle is greater than the preset included angle threshold, then it is determined that the current body posture of the target object is inclined.

[0133] In this step, the preset included angle threshold can adjust its value according to the requirements of different application scenarios.

[0134] The formula for calculating the included angle θ can be described as follows:

[0135] θ = arccos(abs(y42 - y3) / sqrt((y42 - y3) 2 +(x42 - x3) 2 ))

[0136] The numbers in the above formula are explained as follows: In a schematic diagram such as Figure 3 The first key point is the key point numbered 3, whose coordinates correspond to the first coordinate, and this number 3 is used to indicate the nose of the target object; the second key point is the key point numbered 42, whose coordinates correspond to the second coordinate, and this number 42 is used to indicate the crotch line of the target object.

[0137] In different application scenarios, different included angle thresholds can be assumed, and no specific limit is imposed on the included angle threshold here. If in a certain specific application scenario, the included angle threshold is configured as 10 degrees, then when the included angle corresponding to θ calculated according to the above formula is greater than this included angle threshold of 10 degrees, it is determined that the current body posture of the target object is inclined.

[0138] Further, if θ calculated by the above formula is much greater than the included angle threshold of 10 degrees, it is determined that the current inclination amplitude of the target object is relatively large. On the contrary, if θ calculated by the above formula is only slightly greater than the included angle threshold of 10 degrees, for example, 11 degrees, it is determined that the current inclination amplitude of the target object is very small.

[0139] As Figure 6 shown, another method for detecting human body contour points provided by an embodiment of the present disclosure is applied to a server side, and specifically includes the following method steps:

[0140] S602: Obtain a picture including a target object.

[0141] In practical applications, in order to obtain multiple pictures including different objects, pictures of multiple objects can be obtained from multiple clients as Figure 1 shown. In addition, in order to make the detection results obtained by the detection method more accurate, the pictures selected are usually full-body pictures that can fully cover at least one object, and the full-body pictures can clearly show the facial contour of the object.

[0142] S604: Determine at least one target object from the picture.

[0143] It should be noted that for the specific description of the step of determining at least one target object from the picture, please refer to the description of the same or similar parts above, and details are not repeated here.

[0144] S606: Input all the human body contour key points of at least one target object with a preset quantity into a preset detection model for training, and output a detection result, so as to generate a human body contour of the target object according to the detection result. The detection result includes the coordinates of the human body contour key points with the preset quantity.

[0145] In this step, after determining at least one target object, during the process of detecting the human body contour points of at least one target object, assuming that the number of currently detected target objects is one, for this target object, the preset quantity of human body contour key points is not specifically limited. As long as the preset quantity can cover all the key parts corresponding to the whole body outline of the target object and fully cover the facial features of the target object. In practical applications, the preset quantity can be adjusted according to the needs of different application scenarios, and the preset quantity is not specifically limited here.

[0146] Through this step, the coordinates of the human body contour key points with the preset quantity can be determined and output. For the specific description, please refer to the same or similar parts in the above method embodiments, and details are not repeated here.

[0147] It should be noted that the coordinates of each human body contour key point used to represent the target object also include a third coordinate and a fourth coordinate. The third coordinate is the coordinate of the third key point used to represent the first shoulder of the target object, and the fourth coordinate is the coordinate of the fourth key point used to represent the second shoulder of the target object.

[0148] S608: Read the ordinate in the first coordinate, the ordinate in the third coordinate, and the ordinate in the fourth coordinate.

[0149] In a schematic diagram such as Figure 3 the first key point is the key point numbered 3, and its coordinate corresponds to the first coordinate. The number 3 is used to indicate the nose of the target object; the third key point is the key point numbered 13, and its coordinate corresponds to the third coordinate. The number 13 is used to indicate the first shoulder of the target object, that is, one side of the shoulder; the fourth key point is the key point numbered 71, and its coordinate corresponds to the fourth coordinate. The number 71 is used to indicate the second shoulder of the target object, that is, the opposite shoulder corresponding to the first shoulder. This is just an example, and other numbers can also be used to indicate the nose of the target object, or other numbers can be used to indicate the first shoulder of the target object, or other numbers can be used to indicate the second shoulder of the target object.

[0150] S610: Calculate the height of the target object's shoulder based on the ordinate in the third coordinate and the ordinate in the fourth coordinate, and obtain the average value of the target object's shoulder height.

[0151] In a specific application scenario, the calculation formula for calculating the average value of the target object's shoulder height is as follows:

[0152] yshoulder = (y13 + y71) / 2.0

[0153] The numbers in the above formula are explained as follows: In a schematic diagram such as Figure 3 the third key point is the key point numbered 13, and its coordinate corresponds to the third coordinate. The number 13 is used to indicate the first shoulder of the target object, that is, one side of the shoulder; the fourth key point is the key point numbered 71, and its coordinate corresponds to the fourth coordinate. The number 71 is used to indicate the second shoulder of the target object, that is, the opposite shoulder corresponding to the first shoulder.

[0154] S612: Determine the current body posture of the target object based on the ordinate in the first coordinate and the average value of the target object's shoulder height. When the height value corresponding to the ordinate in the first coordinate is greater than the average value of the target object's shoulder height, it is determined that the current body posture of the target object is bending over. In a specific application scenario, the following formula is used to determine whether the current body posture of the target object is bending over.

[0155]

[0156] Wherein, ynose is the ordinate in the first coordinate, and yshoulder is the average value of the shoulder height of the target object.

[0157] As Figure 7 shown, another method for detecting human body contour points provided by an embodiment of the present disclosure is applied to the server side, and specifically includes the following method steps:

[0158] S702: Obtain a picture including the target object.

[0159] In practical applications, in order to obtain multiple pictures including different objects, pictures of multiple objects can be obtained from multiple clients as Figure 1 shown. In addition, in order to make the detection results obtained by the detection method more accurate, the pictures selected are usually full-body pictures that can completely cover at least one object, and the full-body pictures are pictures that can clearly show the facial features of the object.

[0160] S704: Determine at least one target object from the picture.

[0161] It should be noted that for the specific description of the step of determining at least one target object from the picture, please refer to the description of the same or similar parts above, and will not be repeated here.

[0162] S706: Input the human body contour key points of the preset number of at least one target object into a preset detection model for training, and output the detection result, so as to generate the human body contour of the target object according to the detection result. The detection result includes the coordinates of the human body contour key points of the preset number.

[0163] In this step, after determining at least one target object, during the process of detecting the human body contour points of at least one target object, assuming that the number of currently to-be-detected target objects is one, for this target object, there is no specific limit on the preset number of human body contour key points. As long as the preset number can cover all the key parts corresponding to the whole body outline of the target object and fully cover the facial features of the target object. In practical applications, the preset number can be adjusted according to the needs of different application scenarios, and no specific limit is made on the preset number here.

[0164] Through this step, the coordinates of the human body contour key points of the preset number can be determined and output. For the specific description, please refer to the same or similar parts in the above method embodiments, and will not be repeated here.

[0165] S708: Read the coordinates of any one of the human body contour key points of the preset number.

[0166] In a specific application scenario, through this step, the coordinates such asFigure 3 The coordinates of the 77 key points shown, including the abscissa and the corresponding ordinate of each key point.

[0167] S710: When the area corresponding to the coordinates of any key point is a non-picture area, it is determined that the target object is in a state where the body is blocked; wherein, the target object may be in a state where the body is partially blocked, or the target object may also be in a state where the body is completely blocked. When the target object is in a state where the body is blocked, the target object is reminded to move until it moves to a position where the entire contour of the target object can be detected.

[0168] In a specific application scenario, for example, Figure 3 In the application scenario of the 77 key points shown, when the area corresponding to the coordinates of any key point is a non-picture area, it can be determined that the target object is in a state where the body is blocked.

[0169] Specifically, among the above 77 key points, if the area corresponding to any key point is a non-picture area, it is determined that the target object is partially blocked, and the number of key points in the non-picture area is positively correlated with the degree of occlusion of the target object. That is: the more the number of key points in the non-picture area, the higher the degree of occlusion of the target object. When all the above 77 key points are in the non-picture area, it is determined that the target object is completely blocked. For the sake of convenience of description, here, for example, Figure 3 The 77 key points shown are described. In the detection method provided by the embodiments of the present disclosure, the number of key points is not specifically limited. For specific descriptions, refer to the same or related descriptions above, and details are not repeated here.

[0170] Such as Figure 8 As shown, another human body contour point detection method provided by the embodiments of the present disclosure is applied to the server side, and specifically includes the following method steps:

[0171] S802: Obtain a picture including the target object.

[0172] In practical applications, in order to obtain multiple pictures including different objects, pictures of multiple objects can be obtained from multiple clients as shown, for example, Figure 1 In addition, in order to make the detection results obtained by the detection method more accurate, the pictures selected are usually full-body pictures that can completely cover at least one object, and the full-body pictures can clearly show the facial contour of the object.

[0173] S804: Determine at least one target object from the picture.

[0174] It should be noted that for the specific description of the steps of determining at least one target object from the picture, please refer to the description of the same or similar parts above, and will not be elaborated here.

[0175] S806: Input the human contour key points of the preset quantity of at least one target object into a preset detection model for training, and output a detection result to generate the human contour of the target object according to the detection result. The detection result includes the coordinates of the human contour key points of the preset quantity.

[0176] In this step, after determining at least one target object, during the process of detecting the human contour points of at least one target object, assuming that the number of currently to-be-detected target objects is one, for this target object, there is no specific limitation on the preset quantity of the human contour key points. As long as the preset quantity can cover all the key parts corresponding to the entire outer contour of the target object and fully cover the facial features of the target object. In practical applications, the preset quantity can be adjusted according to the needs of different application scenarios, and no specific limitation is made on the preset quantity here.

[0177] In the detection method provided by the embodiments of the present disclosure, in addition to the coordinates of the human contour key points of the above-mentioned preset quantity, the detection result also includes the coordinates of multiple key points for determining the boundary of the detection frame.

[0178] The detection method provided by the embodiments of the present disclosure, the finally output detection result not only includes the coordinates of the human contour key points of the preset quantity, but also includes the coordinates of multiple key points for determining the boundary of the detection frame. Through the coordinates of the multiple key points for determining the boundary of the detection frame, the detection frame can be accurately generated and marked; in this way, when a certain body part of the target object is outside the boundary of the detection frame, the target object is timely reminded to move the body until all body parts of the target object are within the boundary of the detection frame.

[0179] S808: Obtain the coordinates of multiple key points for determining the boundary of the detection frame in the detection result, and automatically generate a detection frame according to the coordinates of the multiple key points for determining the boundary of the detection frame.

[0180] In practical applications, based on the finally output detection result not only including the coordinates of the human contour key points of the preset quantity, but also including the coordinates of multiple key points for determining the boundary of the detection frame. In this way, not only can the human contour key points of the target object be marked according to the coordinates of the human contour key points of the preset quantity, and finally the clear human contour of the target object can be presented; but also the boundary of the detection frame can be marked according to the coordinates of the multiple key points for determining the boundary of the detection frame. Figure 9As shown, it is a schematic diagram of the detection frame and each key point of the human body contour marked in a specific application scenario; by determining the coordinates of multiple key points on the boundary of the detection frame, the detection frame can be clearly marked; for the specific algorithm for determining the detection frame, please refer to the same or related parts above and will not be elaborated here. By marking the detection frame, it is possible to accurately determine whether the human body contour of the target object exceeds the detection frame, and when any key point of the human body contour of the target object exceeds the detection frame, the target object will be reminded to make adjustments until any key point of the human body contour does not exceed the detection frame.

[0181] In addition, when there are multiple target objects to be detected in the picture, different target objects can be clearly distinguished by their respective corresponding detection frames, and it is also convenient to remind each target object whether it exceeds its corresponding detection frame.

[0182] Similarly, after marking the key points of the human body contour of the target object, it is possible to clearly see the positions of each key point of the human body contour from the schematic diagram as Figure 9 shown. As shown in the schematic diagram Figure 9 shown, it includes a preset number of key points of the human body contour, where the preset number is 77 key points of the human body contour. Figure 9 The coordinates of the 77 key points of the human body contour are shown. The coordinates of these 77 key points of the human body contour not only include the coordinates of each contour point that can fully cover the whole body contour of the target object, but also include the coordinates of each contour point that can cover the facial features contour of the target object. Since the obtained contour points are sufficient in number and fully cover, it can meet the requirements of various application scenarios for the number and accuracy of user contour points, and improve the user experience.

[0183] In an embodiment of the present disclosure, an image including a target object is obtained; at least one target object is determined from the image; a preset number of human body contour key points of at least one target object are all input into a preset detection model for training, and a detection result is output to generate a human body contour of the target object according to the detection result. The detection result includes the coordinates of a preset number of human body contour key points. Compared with the existing detection methods that can only detect each human body bone key point, that is, the detection result can only show the two-dimensional contour of the target object. However, in the embodiment of the present application, since the output result includes the coordinates of a preset number of human body contour key points, the coordinates of the preset number of human body contour key points not only include the coordinates of each contour point that can completely cover the whole body contour of the target object, but also include the coordinates of each contour point that can cover the facial features contour of the target object. In this way, a human body contour of the target object can be generated from the detection result, and human body attribute information can be generated from the labeled human body contour of the target object. In addition, the detection method provided by the present invention can simultaneously identify multiple target objects and synchronously detect the coordinates of a preset number of human body contour key points of each of the multiple target objects when the image includes multiple objects and all the multiple objects are target objects, thereby improving the detection efficiency.

[0184] The following is an embodiment of the human body contour point detection device according to the embodiment of the present disclosure, which can be used to execute the embodiment of the human body contour point detection method according to the embodiment of the present disclosure. For the details not disclosed in the embodiment of the human body contour point detection device according to the embodiment of the present disclosure, please refer to the embodiment of the human body contour point detection method according to the embodiment of the present disclosure.

[0185] Please refer to Figure 10 , which shows a schematic structural diagram of a human body contour point detection device provided by an exemplary embodiment of the present invention. The human body contour point detection device can be implemented as all or part of a terminal through software, hardware, or a combination of both. The human body contour point detection device includes an acquisition unit 1002, a determination unit 1004, and a processing unit 1006.

[0186] Specifically, the acquisition unit 1002 is configured to acquire an image including a target object;

[0187] The determination unit 1004 is configured to determine at least one target object from the image acquired by the acquisition module;

[0188] The processing unit 1006 is configured to input a preset number of human body contour key points of at least one target object determined by the determination module into a preset detection model for training, and output a detection result to generate a human body contour of the target object according to the detection result. The detection result includes the coordinates of a preset number of human body contour key points.

[0189] Optionally, the processing unit 1006 is specifically configured to:

[0190] Obtain a data set, where the data set includes a preset number of first human body contour key points, and the first human body contour key points are human body contour key points marked based on a marking tool;

[0191] Store the data set in a file in a preset format;

[0192] Use the picture and the file in the preset format as training samples, input them into a preset detection model for training, and output a prediction result, where the prediction result includes a preset number of second human body contour key points, and the second human body contour key points are predicted human body contour key points;

[0193] Calculate the loss function of the preset detection model according to the preset number of first human body contour key points and the preset number of second human body contour key points;

[0194] Iteratively train the training samples to obtain a preset detection model, and output a weight file including the preset detection model.

[0195] Optionally, the processing unit 1006 is specifically further configured to:

[0196] Input the preset number of human body contour key points of at least one target object into the preset detection model for training, and output network output results in multiple dimensions, where the network output results in multiple dimensions include a first network output result for characterizing the offset value of the preset number of human body contour key points based on the center point, a second network output result for characterizing the feature map of the preset number of human body contour key points, and a third network output result for characterizing the offset value of multiple human body contour key points among the preset number;

[0197] Determine and output the coordinates of the preset number of human body contour key points according to the first network output result, the second network output result, and the third network output result.

[0198] Optionally, the processing unit 1006 is specifically further configured to:

[0199] After determining and outputting the coordinates of the preset number of human body contour key points, read the first coordinate of the first key point for characterizing the nose of the target object and the second coordinate of the second key point for characterizing the crotch line of the target object;

[0200] Set an included angle, where the included angle is the included angle between the line connecting the first key point determined by the first coordinate and the second key point determined by the second coordinate and the vertical perpendicular line;

[0201] In the case where the included angle is greater than a preset included angle threshold, it is determined that the current body posture of the target object is inclined; or,

[0202] Read the coordinates of any one of the preset number of key points of the human body contour;

[0203] When the area corresponding to the coordinates of any one of the key points is in a non-picture area, it is determined that the target object is in a state where the body is blocked.

[0204] Optionally, the processing unit 1006 is further specifically configured to:

[0205] After determining and outputting the coordinates of the preset number of key points of the human body contour, read the ordinate in the first coordinate, the ordinate in the third coordinate, and the ordinate in the fourth coordinate, where the third coordinate is the coordinate of the third key point used to represent the first shoulder of the target object, and the fourth coordinate is the coordinate of the fourth key point used to represent the second shoulder of the target object;

[0206] According to the ordinate in the third coordinate and the ordinate in the fourth coordinate, calculate the height of the target object's shoulders to obtain the average value of the target object's shoulder height;

[0207] According to the ordinate in the first coordinate and the average value of the target object's shoulder height, determine the current body posture of the target object. When the height value corresponding to the ordinate in the first coordinate is greater than the average value of the target object's shoulder height, it is determined that the current body posture of the target object is bending over.

[0208] Optionally, the determining unit 1004 is specifically configured to:

[0209] Obtain multiple objects in the picture;

[0210] Obtain the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object;

[0211] According to the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object, calculate the score of each object in the multiple objects respectively;

[0212] Determine at least one object with the highest score among the multiple objects as the target object.

[0213] Optionally, the detection result further includes the coordinates of multiple key points for determining the boundary of the detection frame, and the processing unit 1006 is further specifically configured to:

[0214] Input the preset number of human contour key points of at least one target object into a preset detection model for training, and output network output results in multiple dimensions. The network output results in multiple dimensions also include a fourth network output result corresponding to a feature map for characterizing the category of the target object, a fifth network output result for characterizing the size of the target object, and a sixth network output result corresponding to an offset value for characterizing the center point of the target object.

[0215] Determine and output the coordinates of multiple key points for determining the boundary of the detection frame according to the fourth network output result, the fifth network output result, and the sixth network output result.

[0216] For the specific description of the above six network data results, please refer to the description of the same or similar parts in the above method embodiments, which will not be repeated here.

[0217] It should be noted that when the human contour point detection device provided in the above embodiment executes the human contour point detection method, only the above division of each functional unit is used as an example. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above. In addition, the human contour point detection device provided in the above embodiment and the human contour point detection method embodiment belong to the same concept. The implementation process is detailed in the human contour point detection method embodiment, which will not be repeated here.

[0218] In the embodiment of the present disclosure, the acquisition unit acquires a picture including a target object; the determination unit determines at least one target object from the picture; the processing unit inputs the preset number of human contour key points of at least one target object into a preset detection model for training and outputs a detection result to generate a human contour of the target object according to the detection result. The detection result output by the processing unit includes the coordinates of the preset number of human contour key points. Compared with the existing detection devices that can only detect each human bone key point, that is, the detection result can only show the two-dimensional contour of the target object. In the embodiment of the present application, since the output result includes the coordinates of the preset number of human contour key points, the coordinates of the preset number of human contour key points not only include the coordinates of each contour point that can fully cover the whole body contour of the target object, but also include the coordinates of each contour point that can cover the facial features contour of the target object. In this way, a human contour of the target object can be generated from the detection result, and human attribute information can be generated from the labeled human contour of the target object. In addition, the detection method provided by the present invention can simultaneously identify multiple target objects and synchronously detect the coordinates of the preset number of human contour key points of each of the multiple target objects when there are multiple objects in the picture and all the multiple objects are target objects, thereby improving the detection efficiency.

[0219] Such asFigure 11 As shown, this embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the method steps as described above.

[0220] This embodiment of the present disclosure provides a storage medium storing computer-readable instructions, on which a computer program is stored, and the program is executed by a processor to implement the method steps as described above.

[0221] Next, refer to Figure 11 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0222] As Figure 11 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1102 or the program loaded from the storage device 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device are also stored. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.

[0223] Generally, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device shown has various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or provided alternatively.

[0224] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0225] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0226] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.

[0227] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0228] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0229] The units described in the embodiments of this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.

Claims

1. A method for detecting human body contour points, characterized in that, The method includes: Obtaining a picture including a target object; Determining at least one of the target objects from the picture; Inputting the human body contour key points of at least one target object in a preset number into a preset detection model for training, and outputting a detection result to generate the human body contour of the target object according to the detection result, where the detection result includes the coordinates of the human body contour key points in the preset number; The step of inputting the human body contour key points of at least one target object in a preset number into a preset detection model for training and outputting a detection result includes: Inputting the human body contour key points of at least one target object in a preset number into the preset detection model for training, and outputting network output results in multiple dimensions; The network output results in multiple dimensions include a first network output result for characterizing the offset value of the human body contour key points in the preset number based on the center point, a second network output result for characterizing the feature map of the human body contour key points in the preset number, and a third network output result for characterizing the offset value of multiple human body contour key points in the preset number; Determining and outputting the coordinates of the human body contour key points in the preset number according to the first network output result, the second network output result, and the third network output result.

2. The method according to claim 1, wherein The step of inputting the human body contour key points of at least one target object in a preset number into a preset detection model for training includes: Obtaining a data set, where the data set includes the human body contour key points in the preset number of the first type, and the human body contour key points in the first type are the human body contour key points marked by a marking tool; Storing the data set in a file in a preset format.

3. The method according to claim 2, wherein The step of inputting the human body contour key points of at least one target object in a preset number into a preset detection model for training further includes: Taking the picture and the file in the preset format as training samples, inputting them into the preset detection model for training, and outputting a prediction result, where the prediction result includes the human body contour key points in the preset number of the second type, and the human body contour key points in the second type are the predicted human body contour key points; Calculating the loss function of the preset detection model according to the human body contour key points in the preset number of the first type and the human body contour key points in the preset number of the second type; Iteratively training the training samples to obtain the preset detection model, and outputting a weight file including the preset detection model.

4. The method according to claim 1, characterized in that After determining and outputting the coordinates of the human body contour key points in the preset number, the method further includes: Reading the first coordinate of the first key point for characterizing the nose of the target object and the second coordinate of the second key point for characterizing the crotch line of the target object; Setting an included angle, where the included angle is the included angle between the connecting line obtained by connecting the first key point determined by the first coordinate and the second key point determined by the second coordinate and the vertical perpendicular line; In the case where the included angle is greater than a preset included angle threshold, determining that the current body posture of the target object is inclined; or, Reading the coordinates of any one of the human body contour key points in the preset number; When the area corresponding to the coordinates of any of the key points is in a non-picture area, it is determined that the target object is in a state where the body is occluded.

5. The method according to claim 4, characterized in that The respective coordinates of each of the human contour key points for characterizing the target object further include a third coordinate and a fourth coordinate. The third coordinate is the coordinate of the third key point for characterizing the first shoulder of the target object, and the fourth coordinate is the coordinate of the fourth key point for characterizing the second shoulder of the target object. After determining and outputting the coordinates of the preset number of human contour key points, the method further includes: Reading the ordinate in the first coordinate, the ordinate in the third coordinate, and the ordinate in the fourth coordinate; Calculating the height of the shoulders of the target object based on the ordinate in the third coordinate and the ordinate in the fourth coordinate to obtain the average value of the shoulder height of the target object; Judging the current body posture of the target object based on the ordinate in the first coordinate and the average value of the shoulder height of the target object. When the height value corresponding to the ordinate in the first coordinate is greater than the average value of the shoulder height of the target object, it is judged that the current body posture of the target object is bending over.

6. The method according to claim 1, characterized in that The determining at least one target object from the picture includes: Obtaining a plurality of objects in the picture; Obtaining the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object; Calculating the score of each object among the plurality of objects respectively according to the pixel ratio of each object in the picture, the first weight value corresponding to the pixel ratio of each object, the prediction score of each object, and the second weight value corresponding to the prediction score of each object; Determining at least one object with the highest score among the plurality of objects as the target object.

7. The method according to claim 1, characterized in that, The detection result further includes the coordinates of a plurality of key points for determining the boundary of the detection frame. The inputting the preset number of human contour key points of at least one target object into a preset detection model for training and outputting the detection result further includes: Inputting the preset number of human contour key points of at least one target object into the preset detection model for training, and outputting a network output result in multiple dimensions. The network output result in multiple dimensions further includes a fourth network output result corresponding to a feature map for characterizing the category of the target object, a fifth network output result for characterizing the size of the target object, and a sixth network output result corresponding to an offset value for characterizing the center point of the target object; Determining and outputting the coordinates of a plurality of key points for determining the boundary of the detection frame according to the fourth network output result, the fifth network output result, and the sixth network output result.

8. A human body contour point detection device, characterized in that The device includes: An obtaining unit for obtaining a picture including a target object; A determining unit for determining at least one target object from the picture obtained by the obtaining unit; A processing unit, configured to input at least one target object's preset number of human contour key points determined by the determining unit into a preset detection model for training, and output a detection result, so as to generate a human contour of the target object according to the detection result, where the detection result includes coordinates of the preset number of human contour key points; Specifically, the processing unit is further configured to: Input at least one target object's preset number of human contour key points into the preset detection model for training, and output a network output result in multiple dimensions; The network output result in multiple dimensions includes a first network output result for characterizing an offset value of the preset number of human contour key points based on a center point, a second network output result for characterizing a feature map of the preset number of human contour key points, and a third network output result for characterizing an offset value of multiple human contour key points among the preset number; Determine and output coordinates of the preset number of human contour key points according to the first network output result, the second network output result, and the third network output result.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method and mobile terminal

    CN105303523A