Human body attribute recognition method, device and equipment, medium and vehicle

By adjusting and combining basic confidence and auxiliary confidence, the problem of distortion of human trunk length affecting the accuracy of human body attribute recognition in complex environments is solved, and the accuracy of recognition is improved.

CN119964196APending Publication Date: 2025-05-09BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311482631.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In complex environments, distortion of the length of the human body's trunk affects the accuracy of human body's attribute recognition, resulting in low error detection and recognition accuracy.

Method used

By obtaining the basic confidence and auxiliary confidence in the image to be detected, the basic target confidence is adjusted and the step coefficient is assigned, and the human body attribute is determined by comparing the first confidence and the second confidence.

Benefits of technology

The error detection rate caused by distortion of the human torso length is reduced, and the accuracy of human attribute recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964196A_ABST
    Figure CN119964196A_ABST
Patent Text Reader

Abstract

The invention relates to a human body attribute recognition method, device and equipment, a medium and a vehicle, and the method comprises the steps: obtaining a to-be-detected image; obtaining a basic confidence coefficient and an auxiliary confidence coefficient corresponding to the target human body in the to-be-detected image; determining a basic target confidence coefficient meeting the target condition, and adjusting the basic target confidence coefficient; adding the adjusted basic target confidence and the corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, and adding the basic confidence except the basic target confidence and the corresponding auxiliary confidence to obtain a second confidence corresponding to the target human body; and comparing the first confidence coefficient with the second confidence coefficient, and determining the attribute information of the target human body in the to-be-detected image, so that human body attribute false detection caused by distortion of the human body trunk length in some special states can be avoided by adjusting the basic target confidence coefficient, and the accuracy of human body attribute recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image recognition technology, and in particular to a method, device, equipment, medium and vehicle for human attribute recognition. Background Art

[0002] Human attributes are the main basis for identifying the human body. The identification of human attributes can be used to track and locate the human body. When identifying human attributes, the image features and key point information of the human body are obtained mainly based on a preset machine learning model, and then the first attribute information of the human body is determined based on the image features, and the second attribute information of the human body is determined based on the key point information of the human body, and then the final attribute result of the human body is determined through the first attribute information and the second attribute information.

[0003] However, when the human body is in a complex environment, such as inside the cabin of a vehicle, due to the complexity of the cabin environment and the diversity of human behavior and posture, the human torso length will be greatly distorted under some special conditions. For example, when a child stands on a seat and leans forward, or when a child is in the middle aisle between two rows of seats with his body very close to the camera, the imaging principle of the camera will cause the child's torso length to be extremely magnified, thereby affecting the accuracy of the second attribute information obtained based on the key points of the human body, causing misdetection of human attributes, and leading to the problem of low accuracy in human attribute recognition. Summary of the invention

[0004] In order to solve the above technical problems, the present disclosure provides a human attribute recognition method, device, equipment, medium and vehicle.

[0005] A first aspect of an embodiment of the present disclosure provides a method for identifying human attributes, the method comprising:

[0006] Acquire the image to be detected;

[0007] Obtaining a basic confidence and an auxiliary confidence corresponding to a target human body in an image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence;

[0008] Determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of a basic adult confidence and a basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence;

[0009] The adjusted basic target confidence and its corresponding auxiliary confidence are added to obtain a first confidence corresponding to the target human body, and the basic confidences other than the basic target confidence and their corresponding auxiliary confidences are added to obtain a second confidence corresponding to the target human body;

[0010] The first confidence level is compared with the second confidence level, and the attribute information of the target human body in the image to be detected is determined according to the comparison result.

[0011] A second aspect of an embodiment of the present disclosure provides a human attribute recognition device, the device comprising:

[0012] An image acquisition module, used for acquiring an image to be detected;

[0013] A confidence acquisition module is used to acquire a basic confidence and an auxiliary confidence corresponding to a target human body in an image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, and the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence;

[0014] A confidence adjustment module, used to determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of a basic adult confidence and a basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence;

[0015] A confidence calculation module, used for adding the adjusted basic target confidence and the corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, and adding the basic confidence other than the basic target confidence and the corresponding auxiliary confidence to obtain a second confidence corresponding to the target human body;

[0016] The attribute determination module is used to compare the first confidence level with the second confidence level, and determine the attribute information of the target human body in the image to be detected according to the comparison result.

[0017] A third aspect of an embodiment of the present disclosure provides an electronic device, the device comprising:

[0018] Memory;

[0019] Processor; and

[0020] A computer program, wherein the computer program is stored in a memory and is configured to be executed by a processor to implement the human attribute recognition method of the first aspect as described above.

[0021] A fourth aspect of the embodiments of the present disclosure provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the human attribute recognition method as described in the first aspect above is implemented.

[0022] A fifth aspect of an embodiment of the present disclosure provides a vehicle, comprising the electronic device of the third aspect described above.

[0023] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0024] The human attribute recognition method, device, equipment, medium and vehicle provided by the embodiments of the present disclosure can obtain an image to be detected, obtain a basic confidence and an auxiliary confidence corresponding to a target human in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, and the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence, determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence, The adjusted basic target confidence is added to its corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, the basic confidences other than the basic target confidence are added to their corresponding auxiliary confidences to obtain a second confidence corresponding to the target human body, the first confidence is compared with the second confidence, and the attribute information of the target human body in the image to be detected is determined according to the size of the comparison result. Thus, the basic target confidence that meets the target condition can be adjusted, and a step coefficient is introduced into the basic target confidence to reduce the influence of the auxiliary confidence obtained based on the key point information of the human body due to the distortion of the human body torso length under some special conditions on the confidence of the target human body through the adjustment of the basic target confidence, thereby avoiding the misdetection of human attributes caused by the distortion of the human body torso length under some special conditions, and improving the accuracy of human attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0027] Figure 1 is a flow chart of a human attribute recognition method provided by an embodiment of the present disclosure;

[0028] Figure 2 is a flow chart of a method for determining human attributes provided by an embodiment of the present disclosure;

[0029] Figure 3 is a structural schematic diagram of a human attribute recognition device provided by an embodiment of the present disclosure;

[0030] Figure 4 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0033] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0034] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0036] Figure 1 It is a flow chart of a human attribute recognition method provided by an embodiment of the present disclosure. The method can be executed by a human attribute recognition device. The human attribute recognition device can be implemented in software and / or hardware. The human attribute recognition device can be configured in an electronic device, such as a server or a terminal or a server cluster, wherein the terminal can specifically include a mobile phone, a computer or a tablet computer, a vehicle-mounted terminal, etc., which is not limited here.

[0037] like Figure 1 As shown, the human attribute recognition method provided by the embodiment of the present disclosure includes the following steps.

[0038] S110: Acquire an image to be detected.

[0039] In the embodiment of the present disclosure, the electronic device can acquire the image to be detected based on the image acquisition device.

[0040] The image to be detected may be an image of the interior of a vehicle cabin, wherein the image to be detected may include at least one target human body, for example, the target human body may be a passenger in the vehicle cabin.

[0041] The image acquisition device may be at least one camera installed inside the vehicle cabin.

[0042] In some embodiments of the present disclosure, the electronic device may control the image acquisition device to acquire the image to be detected in real time, and upload the acquired image to be detected to the electronic device based on the image acquisition device, so that the electronic device obtains the image to be detected.

[0043] In some other embodiments of the present disclosure, after receiving the human attribute recognition instruction, the electronic device may obtain the image to be detected corresponding to the human attribute recognition instruction by loading from a local memory based on the human attribute recognition instruction.

[0044] S120. Obtain a basic confidence and an auxiliary confidence corresponding to a target human body in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence.

[0045] In the disclosed embodiment, the basic adult confidence and the basic child confidence are confidences obtained based on the image features of the target human body, and the auxiliary adult confidence and the auxiliary child confidence are confidences obtained based on the key point information of the target human body.

[0046] In the disclosed embodiment, the basic adult confidence and the basic child confidence are based on a preset machine learning model to perform feature extraction and feature recognition and classification on the image to be detected, and obtain the probability that the target human body belongs to an adult and the probability that the target human body belongs to a child.

[0047] In an embodiment of the present disclosure, the preset machine learning model may be a model obtained by training using sample images containing a target human body, and is used for detecting, classifying, and identifying attributes of a target human body, such as a human body detection and classification network.

[0048] Specifically, after acquiring the image to be detected, the electronic device inputs the image to be detected into a preset machine learning model, and the preset machine learning model performs detection classification and attribute recognition on the image to be detected to obtain a detection frame of the target human body corresponding to the image to be detected and a basic adult confidence and a basic child confidence corresponding to the target human body.

[0049] Optionally, the preset machine learning model can be a convolutional neural network (Region-based Convolutional Neural Networks, RCNN), a YOLO model, etc.

[0050] In the disclosed embodiment, after obtaining the detection frame of the target human body corresponding to the image to be detected, the electronic device inputs the area image of the target human body corresponding to the detection frame of the target human body into the human key point detection network, and the human key point detection network recognizes the area image to obtain the key point information corresponding to the target human body, and then determines the auxiliary adult confidence and auxiliary child confidence corresponding to the target human body based on the human key point information.

[0051] Optionally, the human key point detection network may be a pre-trained network model for identifying and predicting key points of a target human body in an image to be detected, such as a human key point detection network (High-Resolution Net, HRNet) and the like.

[0052] S130. Determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of a basic adult confidence and a basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence.

[0053] In the embodiment of the present disclosure, the preset confidence threshold includes a preset adult confidence threshold and a preset child confidence threshold. The preset adult confidence threshold is a preset confidence threshold for determining whether the target human body is an adult; the preset child confidence threshold is a preset confidence threshold for determining whether the target human body is a child.

[0054] The preset adult confidence threshold is obtained by collecting multiple images of adults in various positions and postures, calculating the adult confidence for each image, arranging the multiple adult confidences corresponding to the multiple images in ascending order, and determining the preset adult confidence threshold as the target value corresponding to the adult confidence higher than a certain target value in the arrangement result and the proportion of images with adult confidence higher than a certain target value greater than the preset proportion.

[0055] The preset child confidence threshold is obtained by collecting multiple images of children in various positions and postures, calculating the child confidence for each image, arranging the multiple child confidences corresponding to the multiple images in order from small to large, and determining the preset child confidence threshold as the target value corresponding to the child confidence higher than a certain target value in the arrangement result and the proportion of images with child confidence higher than a certain target value greater than the preset proportion.

[0056] Exemplarily, the number of images of adults in various positions and postures collected is 100. When the number of images whose adult confidence values ​​corresponding to the 100 images are greater than 80 accounts for more than 90% of the total number of images, the adult confidence threshold is determined to be 80. It should be noted that this is just an example for ease of understanding and does not limit this solution.

[0057] When the basic confidence is the basic adult confidence, the target condition is to make the basic adult confidence greater than or equal to the preset adult confidence threshold; when the basic confidence is the basic child confidence, the target condition is to make the basic child confidence greater than or equal to the preset child confidence threshold.

[0058] In the disclosed embodiment, the ladder coefficient can be understood as giving a certain weight to the basic target confidence, so as to avoid misdetection of human attributes due to distortion of the human torso length in some special cases.

[0059] Specifically, after obtaining the basic confidence, the electronic device compares the basic confidence with a preset confidence threshold, determines the basic target confidence that meets the target condition, and adjusts the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment.

[0060] S140. Add the adjusted basic target confidence and its corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, and add the basic confidences other than the basic target confidence and their corresponding auxiliary confidences to obtain a second confidence corresponding to the target human body.

[0061] In some embodiments of the present disclosure, when the basic target confidence is the basic adult confidence, the adjusted basic adult confidence and the auxiliary adult confidence are added to obtain a first confidence, wherein the first confidence can be characterized as the confidence that the target person is an adult, and the basic child confidence and the auxiliary child confidence are added to obtain a second confidence, wherein the second confidence can be characterized as the confidence that the target person is a child.

[0062] In other embodiments of the present disclosure, when the basic target confidence is the basic child confidence, the adjusted basic child confidence and the auxiliary child confidence are added to obtain a first confidence. At this time, the first confidence can be characterized as the confidence that the target person is a child. The basic adult confidence and the auxiliary adult confidence are added to obtain a second confidence. At this time, the second confidence can be characterized as the confidence that the target person is an adult.

[0063] S150: Compare the first confidence level with the second confidence level, and determine attribute information of the target human body in the image to be detected according to the comparison result.

[0064] An example is given in which the first confidence level is represented by the confidence level that the target person is an adult, and the second confidence level is represented by the confidence level that the target person is a child.

[0065] In some embodiments of the present disclosure, when the first confidence is greater than the second confidence, the attribute information of the target person is determined to be an adult; when the adult confidence is less than the child confidence, the attribute information of the target person is determined to be a child; when the adult confidence is equal to the child confidence, the attribute information of the target person is not output.

[0066] In other embodiments of the present disclosure, when the first confidence is greater than or equal to the second confidence, the attribute information of the target person is determined to be an adult, and when the first confidence is less than the second confidence, the attribute information of the target person is determined to be a child.

[0067] In some other embodiments of the present disclosure, when the first confidence is greater than the second confidence, the attribute information of the target person is determined to be an adult, and when the first confidence is less than or equal to the second confidence, the attribute information of the target person is determined to be a child.

[0068] Similarly, when the first confidence level is represented as the confidence level that the target person is a child and the second confidence level is represented as the confidence level that the target person is an adult, the specific implementation method of the first confidence level being represented as the confidence level that the target person is an adult and the second confidence level being represented as the confidence level that the target person is a child is similar to the above-mentioned first confidence level being represented as the confidence level that the target person is an adult and the second confidence level being represented as the confidence level that the target person is a child is not repeated here.

[0069] In the disclosed embodiment, it is possible to obtain an image to be detected, and obtain a basic confidence and an auxiliary confidence corresponding to a target human body in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, and the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence, determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence, and adjusting the adjusted basic target confidence. The first confidence corresponding to the target human body is obtained by adding the corresponding auxiliary confidences, the second confidence corresponding to the target human body is obtained by adding the basic confidences other than the basic target confidences and the corresponding auxiliary confidences, the first confidence is compared with the second confidence, and the attribute information of the target human body in the image to be detected is determined according to the size of the comparison result. Thus, the basic target confidence that meets the target conditions can be adjusted, and a step coefficient is introduced into the basic target confidence to reduce the influence of the auxiliary confidence obtained based on the key point information of the human body due to the distortion of the human body torso length under some special conditions on the confidence of the target human body through the adjustment of the basic target confidence, thereby avoiding the misdetection of human attributes caused by the distortion of the human body torso length under some special conditions, and improving the accuracy of human attribute recognition.

[0070] On the basis of the above-mentioned embodiments of the present invention, basic confidence and auxiliary confidence corresponding to the target human body in the image to be detected are obtained, including obtaining the basic adult confidence and basic child confidence corresponding to the target human body in the image to be detected, and obtaining the auxiliary adult confidence and auxiliary child confidence corresponding to the target human body.

[0071] Among them, obtaining the basic adult confidence and basic child confidence respectively corresponding to the target human body in the image to be detected can specifically include: inputting the image to be detected into a human body detection and classification network, and the human body detection and classification network detects and classifies the image to be detected, and obtains a first detection result corresponding to the target human body in the image to be detected, the first detection result including a detection frame corresponding to the target human body and a basic adult confidence and a basic child confidence corresponding to the target human body.

[0072] Obtaining the auxiliary adult confidence and auxiliary child confidence corresponding to the target human body respectively may specifically include: determining the area image corresponding to the target human body based on the detection frame corresponding to the target human body, inputting the area image into the human key point detection network, and using the human key point detection network to perform human key point recognition on the area image to obtain a second detection result, wherein the second detection result includes key point information corresponding to the target human body; inputting the second detection result into a preset adult key point confidence mapping function to obtain the auxiliary adult confidence corresponding to the target human body, and at the same time inputting the second detection result into a preset child key point confidence mapping function to obtain the auxiliary child confidence corresponding to the target human body.

[0073] In the embodiment of the present disclosure, the key point confidence mapping function is a function that maps human key point detection information to human attribute confidence, including an adult key point confidence mapping function and a child key point confidence mapping function.

[0074] The preset adult key point confidence mapping function and the preset child key point confidence mapping function are respectively generated based on the test data, using mathematical statistics to calculate the area position of the human body detection frame, the human body trunk length, the human body trunk length kernel density, etc., and are used to calculate the auxiliary adult confidence and auxiliary child confidence mapping functions respectively.

[0075] In some embodiments of the present disclosure, a human body detection classification network may include a first backbone network, a first shoulder network, and a first head network. Based on the first backbone network, the first shoulder network, and the first head network, backbone features, shoulder features, and head features are extracted from the image to be detected, respectively, to obtain a feature map corresponding to the image to be detected. The feature map after feature extraction is traversed in the form of candidate frames to obtain a detection frame corresponding to each target human body in the image to be detected, and attribute recognition is performed on the detection frame corresponding to each target human body to obtain a first detection result, wherein the first detection result includes a detection frame corresponding to each target human body in the image to be detected and basic attributes of the human body ( Including basic adult confidence and basic child confidence), and then input the detection frame corresponding to the target human body into the human key point detection network, wherein the human key point detection network may include a second backbone network, a second shoulder network and a second head network, and based on the second backbone network, the second shoulder network and the second head network, the backbone features, shoulder features and head features of the target human body in the detection frame are extracted respectively to obtain a feature map corresponding to the target human body in the detection frame, and then the key points of the target human body are identified, and the identification results are matched with the preset key points to obtain a second detection result, wherein the second detection result includes the human key point information corresponding to the target human body.

[0076] Specifically, after the electronic device acquires the image to be detected, it first preprocesses the image to be detected, wherein the preprocessing may include image cropping, image normalization processing, image filtering processing, etc., and inputs the preprocessed image to be detected into the human body detection classification network, and the human body detection classification network performs feature extraction on the preprocessed image to be detected to obtain a feature map corresponding to the image to be detected, and then the human body detection classification network performs human body recognition and human body attribute recognition on the feature map to obtain a first detection result corresponding to the target human body in the image to be detected, wherein the first detection result includes a detection frame of the target human body and basic human body attributes corresponding to the target human body (including basic adult confidence and basic child confidence), and the target human body corresponding to The regional image corresponding to the detection frame, that is, the regional image of the target human body, is segmented, and the regional image of the target human body is input into the human key point detection network. The human key point detection network performs human key point recognition on the regional image, and the recognition result is matched with the preset key points to obtain a second detection result, wherein the second detection result includes key point information corresponding to the target human body, and further, the second detection result, that is, the key point information corresponding to the target human body, is input into a preset adult key point confidence mapping function to obtain an auxiliary adult confidence corresponding to the target human body, and at the same time, the second detection result, that is, the key point information corresponding to the target human body, is input into a preset child key point confidence mapping function to obtain an auxiliary child confidence corresponding to the target human body.

[0077] In the embodiments of the present disclosure, the basic adult confidence and basic child confidence corresponding to the target human body in the image to be detected can be obtained based on the human body detection classification network, and the key point information corresponding to the target human body can be obtained based on the human body key point detection network, and the key point information is respectively input into the preset adult key point confidence mapping function and the preset child key point confidence mapping function to obtain the auxiliary adult confidence and auxiliary child confidence corresponding to the target human body, thereby improving the accuracy of the obtained basic adult confidence and basic child confidence as well as the auxiliary adult confidence and auxiliary child confidence.

[0078] Further, in some embodiments of the present disclosure, determining the basic target confidence that meets the target conditions may specifically include: comparing the basic adult confidence with a preset adult confidence threshold to obtain a first comparison result, and when the first comparison result shows that the basic adult confidence is greater than or equal to the preset adult confidence threshold, determining the basic adult confidence to be the basic target confidence.

[0079] When the basic target confidence is the basic adult confidence, adjusting the basic target confidence may specifically include: inputting a preset adult confidence threshold and the basic adult confidence into a first exponential function corresponding to a preset first step function to obtain a first step coefficient, wherein the first step coefficient obtained based on the first exponential function is greater than 1; and multiplying the first step coefficient by the basic adult confidence to obtain an adjusted basic adult confidence.

[0080] When the basic adult confidence level based on the first comparison result is less than a preset adult confidence level threshold, the first step coefficient is determined to be 1.

[0081] In the embodiment of the present disclosure, when the first comparison result is that the basic adult confidence is greater than or equal to the preset adult confidence threshold, the expression of the preset first step function is specifically as follows:

[0082]

[0083] Among them, T1 represents the first step coefficient, X represents the basic adult confidence, and Ta represents the preset adult confidence threshold. Represents the first exponential function.

[0084] Furthermore, the expression of the adult confidence corresponding to the target human body is as follows:

[0085] A=X*T1+a

[0086] Among them, A represents the adult confidence corresponding to the target person, X represents the basic adult confidence, a represents the auxiliary adult confidence, and X*T1 represents the adjusted basic adult confidence.

[0087] When the first comparison result shows that the basic adult confidence is less than the preset adult confidence threshold, the first step coefficient is determined to be 1, that is, the basic adult confidence is not adjusted. At this time, the expression of the adult confidence corresponding to the target human body is as follows:

[0088] A=X+a

[0089] In some other embodiments of the present disclosure, determining the basic target confidence that meets the target conditions may specifically include: comparing the basic child confidence with a preset child confidence threshold to obtain a second comparison result, and when the second comparison result shows that the basic child confidence is greater than or equal to the preset child confidence threshold, determining the basic child confidence to be the basic target confidence.

[0090] When the basic target confidence is the basic child confidence, adjusting the basic target confidence may specifically include: inputting a preset child confidence threshold and the basic child confidence into a second exponential function corresponding to a preset second step function to obtain a second step coefficient, wherein the second step coefficient obtained based on the second exponential function is greater than 1; and multiplying the second step coefficient by the basic child confidence to obtain an adjusted basic child confidence.

[0091] When the second comparison result shows that the basic child confidence is less than the preset child confidence threshold, the second step coefficient is determined to be 1.

[0092] In the embodiment of the present disclosure, when the second comparison result is that the basic child confidence is greater than or equal to the preset child confidence threshold, the expression of the preset second step function is specifically as follows:

[0093]

[0094] Among them, T2 represents the second step coefficient, Y represents the basic child confidence, and Tc represents the preset child confidence threshold. Represents the second exponential function.

[0095] Furthermore, the expression of the child confidence corresponding to the target human body is as follows:

[0096] C=Y*T2+b

[0097] Among them, C represents the child confidence corresponding to the target human body, Y represents the basic child confidence, b represents the auxiliary child confidence, and Y*T2 represents the adjusted basic child confidence.

[0098] When the second comparison result shows that the basic child confidence is less than the preset child confidence threshold, the second step coefficient is determined to be 1, that is, the basic child confidence is not adjusted. At this time, the expression of the child confidence corresponding to the target human body is as follows:

[0099] C=Y+b

[0100] Among them, the preset first step function and the preset second step function are to calculate the basic adult confidence and basic child confidence of the image data collected when people with different body features, different genders, and different attributes (adults and children) make different sitting postures in different seats of the vehicle, determine the confidence threshold based on the basic adult confidence and basic child confidence and the real attributes corresponding to the image data, and then determine the coordinates of the confidence threshold and the accuracy of the human attribute condition, and the function obtained by curve fitting in MATLAB software based on the coordinates of the confidence threshold and the accuracy of the human attribute condition, wherein the accuracy of the human attribute condition is the proportion of the attributes of adults in the number of images that meet the preset conditions. For the preset first step function, the preset conditions are that the basic adult confidence is greater than the basic child confidence, and the basic adult confidence is greater than the basic adult confidence threshold, and for the preset second step function, the preset conditions are that the basic child confidence is greater than the basic adult confidence, and the basic child confidence is greater than the basic child confidence threshold.

[0101] In some examples, when a child is very close to an image acquisition device such as a camera, the auxiliary adult confidence obtained by identifying human key points in the image captured by the camera and based on the human key point information will be very large, while the auxiliary child confidence will be very small. Therefore, the obtained adult confidence will be very large, and the child will be misdetected as an adult. However, when the basic adult confidence is greater than or equal to the preset adult confidence threshold, a basic adult confidence weight, i.e., a first-step coefficient, is assigned, or when the basic child confidence is greater than or equal to the preset child confidence threshold, a basic child confidence weight, i.e., a second-step coefficient, is assigned. This can balance the impact of the auxiliary adult confidence or the auxiliary child confidence on the adult confidence or child confidence corresponding to the target human body, thereby reducing the false detection rate of human attribute results.

[0102] In the embodiment of the present disclosure, the first step coefficient of the basic adult confidence can be determined based on the comparison result of the basic adult confidence and the preset adult confidence threshold, and the second step coefficient of the basic child confidence can be determined based on the comparison result of the basic child confidence and the preset child confidence threshold. The first step coefficient and the second step coefficient are used to avoid misdetection of human attributes caused by distortion of the human trunk length under some special conditions, thereby improving the accuracy of human attribute recognition.

[0103] Figure 2 It is a flow chart of a method for determining human attributes provided by an embodiment of the present disclosure. In the embodiment of the present disclosure, based on the attribute information of the target human body in the image to be detected, the attributes of the target human body in the image sequence to be detected are voted to obtain the human attribute results of the target human body, and steps S210-S230 can be specifically executed.

[0104] S210: Acquire a sequence of images to be detected.

[0105] S220: Determine the number of first images in which the target human body is an adult and the number of second images in which the target human body is a child in the image sequence to be detected.

[0106] S230. Compare the number of first images with the number of second images. When the number of first images is greater than or equal to the number of second images, determine that the human attribute result of the target person is an adult. When the number of second images is greater than the number of first images, determine that the human attribute result of the target person is a child.

[0107] Specifically, the electronic device is capable of acquiring a sequence of images to be detected, wherein the sequence of images to be detected may include multiple images to be detected, acquiring attribute information of a target human body in each image to be detected, and after obtaining the attribute information of the target human body in each image to be detected, determining the number of first images in which the target human body belongs to an adult and the number of second images in which the target human body belongs to a child in the sequence of images to be detected, and determining a human attribute result of the target human body based on the number of first images in which the target human body belongs to an adult and the number of second images in which the target human body belongs to a child in the sequence of images to be detected; when the number of first images is greater than or equal to the number of second images, determining the human attribute result of the target human body is an adult, and when the number of second images is greater than the number of first images, determining the human attribute result of the target human body is a child.

[0108] In the disclosed embodiment, the attributes of the target human body in the image sequence to be detected can be voted to obtain the human body attribute results of the target human body, thereby improving the accuracy of obtaining the human body attribute results of the target human body.

[0109] Figure 3 : is a schematic diagram of the structure of a human attribute recognition device provided in an embodiment of the present disclosure. The human attribute recognition device in the embodiment of the present disclosure can be set in an electronic device, and the electronic device can be a server or a terminal or a server cluster, wherein the terminal can specifically include a mobile phone, a computer or a tablet computer, and a vehicle-mounted terminal, which is not limited here.

[0110] like Figure 3 As shown, the human attribute recognition device 300 may include an image acquisition module 310 , a confidence acquisition module 320 , a confidence adjustment module 330 , a confidence calculation module 340 and an attribute determination module 350 .

[0111] The image acquisition module 310 can be used to acquire an image to be detected.

[0112] The confidence acquisition module 320 can be used to obtain the basic confidence and auxiliary confidence corresponding to the target human body in the image to be detected, wherein the basic confidence includes the basic adult confidence and the basic child confidence, the auxiliary confidence includes the auxiliary adult confidence and the auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence.

[0113] The confidence adjustment module 330 can be used to determine the basic target confidence that meets the target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence.

[0114] The confidence calculation module 340 can be used to add the adjusted basic target confidence and its corresponding auxiliary confidence to obtain a first confidence corresponding to the target person, and add the basic confidence other than the basic target confidence and its corresponding auxiliary confidence to obtain a second confidence corresponding to the target person.

[0115] The attribute determination module 350 may be used to compare the first confidence level with the second confidence level, and determine the attribute information of the target human body in the image to be detected according to the comparison result.

[0116] In the disclosed embodiment, it is possible to obtain an image to be detected, and obtain a basic confidence and an auxiliary confidence corresponding to a target human body in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, and the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence, determine a basic target confidence that meets a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence, and adjusting the adjusted basic target confidence. The first confidence corresponding to the target human body is obtained by adding the corresponding auxiliary confidences, the second confidence corresponding to the target human body is obtained by adding the basic confidences other than the basic target confidences and the corresponding auxiliary confidences, the first confidence is compared with the second confidence, and the attribute information of the target human body in the image to be detected is determined according to the size of the comparison result. Thus, the basic target confidence that meets the target conditions can be adjusted, and a step coefficient is introduced into the basic target confidence to reduce the influence of the auxiliary confidence obtained based on the key point information of the human body due to the distortion of the human body torso length under some special conditions on the confidence of the target human body through the adjustment of the basic target confidence, thereby avoiding the misdetection of human attributes caused by the distortion of the human body torso length under some special conditions, and improving the accuracy of human attribute recognition.

[0117] In some embodiments of the present disclosure, the confidence acquisition module 320 can be specifically used to input the image to be detected into a human body detection and classification network, and the human body detection and classification network detects and classifies the image to be detected to obtain a first detection result corresponding to the target human body in the image to be detected, the first detection result including a detection frame corresponding to the target human body and a basic adult confidence and a basic child confidence corresponding to the target human body; based on the detection frame corresponding to the target human body, determine the regional image corresponding to the target human body, input the regional image into the human body key point detection network, and the human body key point detection network performs human body key point recognition on the regional image to obtain a second detection result, the second detection result including key point information corresponding to the target human body; input the second detection result into a preset adult key point confidence mapping function to obtain an auxiliary adult confidence corresponding to the target human body, and at the same time input the second detection result into a preset child key point confidence mapping function to obtain an auxiliary child confidence corresponding to the target human body.

[0118] In some embodiments of the present disclosure, the confidence adjustment module 330 can be specifically used to compare the basic adult confidence with a preset adult confidence threshold to obtain a first comparison result; when the first comparison result shows that the basic adult confidence is greater than or equal to the preset adult confidence threshold, the basic adult confidence is determined to be the basic target confidence; when the basic target confidence is the basic adult confidence, the preset adult confidence threshold and the basic adult confidence are input into a first exponential function corresponding to a preset first step function to obtain a first step coefficient, wherein the first step coefficient obtained based on the first exponential function is greater than 1; the first step coefficient is multiplied by the basic adult confidence to obtain an adjusted basic adult confidence.

[0119] In some embodiments of the present disclosure, the confidence adjustment module 330 can be specifically used to compare the basic child confidence with a preset child confidence threshold to obtain a second comparison result. When the second comparison result shows that the basic child confidence is greater than or equal to the preset child confidence threshold, the basic child confidence is determined to be the basic target confidence; when the basic target confidence is the basic child confidence, the preset child confidence threshold and the basic child confidence are input into a second exponential function corresponding to a preset second step function to obtain a second step coefficient, wherein the second step coefficient obtained based on the second exponential function is greater than 1; the second step coefficient is multiplied by the basic child confidence to obtain an adjusted basic child confidence.

[0120] In some embodiments of the present disclosure, the human attribute recognition device 300 may include an attribute voting module.

[0121] The attribute voting module can be used to obtain a sequence of images to be detected; based on the attribute information of the target human body in each image to be detected in the sequence of images to be detected, the attributes of the target human body in the sequence of images to be detected are voted to obtain the human body attribute results of the target human body.

[0122] In some embodiments of the present disclosure, the attribute voting module can be specifically used to determine the number of first images in which the target human body belongs to an adult and the number of second images in which the target human body belongs to a child in the image sequence to be detected; compare the number of first images with the number of second images, and when the number of first images is greater than or equal to the number of second images, determine that the human attribute result of the target human body is an adult, and when the number of second images is greater than the number of first images, determine that the human attribute result of the target human body is a child.

[0123] It should be noted that Figure 3 The human attribute recognition device 300 shown can execute each step in the above method embodiment and realize each process and effect in the above method embodiment, which will not be described in detail here.

[0124] Figure 4A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown.

[0125] In the disclosed embodiment, Figure 4 The electronic device shown may be a server or a terminal or a server cluster, wherein the terminal may specifically include a mobile phone, a computer or a tablet computer, a vehicle-mounted terminal, or any device that can be used for the human attribute recognition method, etc., and is not limited here.

[0126] like Figure 4 As shown, the electronic device may include a processor 410 and a memory 420 storing computer program instructions.

[0127] Specifically, the processor 410 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0128] The memory 420 may include a large capacity memory for information or instructions. By way of example and not limitation, the memory 420 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 420 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 420 may be inside or outside the integrated gateway device. In a particular embodiment, the memory 420 is a non-volatile solid-state memory. In a particular embodiment, the memory 420 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (Electrically Erasable Programmable ROM, EEPROM), an electrically rewritable ROM (EAROM) or a flash memory, or a combination of two or more of these.

[0129] The processor 410 reads and executes the computer program instructions stored in the memory 420 to perform the steps of the human attribute recognition method provided in the embodiment of the present disclosure.

[0130] In one example, the electronic device may further include a transceiver 430 and a bus 440. Figure 4 As shown, the processor 410, the memory 420 and the transceiver 430 are connected via a bus 440 and communicate with each other.

[0131] The bus 440 includes hardware, software, or both. For example, but not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 440 may include one or more buses.

[0132] The embodiments of the present disclosure further provide a computer-readable storage medium, which may store a computer program. When the computer program is executed by a processor, the processor implements the human attribute recognition method provided by the embodiments of the present disclosure.

[0133] The above-mentioned storage medium may, for example, include a memory 420 of computer program instructions, and the above-mentioned instructions may be executed by the processor 410 of the electronic device to complete the human attribute recognition method provided in the embodiment of the present disclosure. Optionally, the storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (Random Access Memory, RAM), a compact disc read-only memory (Compact Disc ROM, CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0134] The embodiments of the present disclosure also provide a vehicle, which includes electronic equipment and can implement the various processes and effects in the above embodiments of the present disclosure, which will not be elaborated here.

[0135] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0136] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A human attribute recognition method, characterized in that: The method comprises: Acquire the image to be detected; Obtaining a basic confidence and an auxiliary confidence corresponding to the target human body in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence; Determining a basic target confidence that meets a target condition, and adjusting the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence; Adding the adjusted basic target confidence and its corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, and adding the basic confidences other than the basic target confidence to their corresponding auxiliary confidences to obtain a second confidence corresponding to the target human body; The first confidence level is compared with the second confidence level, and the attribute information of the target human body in the image to be detected is determined according to the comparison result.

2. The method according to claim 1, characterized in that The obtaining of the basic confidence and the auxiliary confidence corresponding to the target human body in the image to be detected includes: Inputting the image to be detected into a human body detection and classification network, the human body detection and classification network detects and classifies the image to be detected, and obtains a first detection result corresponding to a target human body in the image to be detected, wherein the first detection result includes a detection frame corresponding to the target human body and a basic adult confidence and a basic child confidence corresponding to the target human body; Determine a region image corresponding to the target human body based on a detection frame corresponding to the target human body, input the region image into a human key point detection network, and use the human key point detection network to perform human key point recognition on the region image to obtain a second detection result, wherein the second detection result includes key point information corresponding to the target human body; The second detection result is input into the preset adult key point confidence mapping function to obtain the auxiliary adult confidence corresponding to the target human body. At the same time, the second detection result is input into the preset child key point confidence mapping function to obtain the auxiliary child confidence corresponding to the target human body.

3. The method according to claim 1, characterized in that The determining of the basic target confidence level that satisfies the target condition includes: Comparing the basic adult confidence with a preset adult confidence threshold to obtain a first comparison result, and when the first comparison result is that the basic adult confidence is greater than or equal to the preset adult confidence threshold, determining that the basic adult confidence is the basic target confidence; When the basic target confidence is the basic adult confidence, adjusting the basic target confidence includes: Inputting the preset adult confidence threshold and the basic adult confidence into a first exponential function corresponding to a preset first step function to obtain a first step coefficient, wherein the first step coefficient obtained based on the first exponential function is greater than 1; The first step coefficient is multiplied by the basic adult confidence to obtain an adjusted basic adult confidence.

4. The method according to claim 1, characterized in that The determining of the basic target confidence level that satisfies the target condition includes: Comparing the basic child confidence with a preset child confidence threshold to obtain a second comparison result, and when the second comparison result is that the basic child confidence is greater than or equal to the preset child confidence threshold, determining the basic child confidence to be the basic target confidence; When the basic target confidence is the basic child confidence, adjusting the basic target confidence includes: Inputting the preset child confidence threshold and the basic child confidence into a second exponential function corresponding to a preset second step function to obtain a second step coefficient, wherein the second step coefficient obtained based on the second exponential function is greater than 1; The second step coefficient is multiplied by the basic child confidence to obtain an adjusted basic child confidence.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining a sequence of images to be detected; Based on the attribute information of the target human body in each image to be detected in the image sequence to be detected, the attributes of the target human body in the image sequence to be detected are voted to obtain a human body attribute result of the target human body.

6. The method according to claim 5, characterized in that The step of voting on the attributes of the target human body in the sequence of images to be detected based on the attribute information of the target human body in each image to be detected in the sequence of images to be detected to obtain the human body attribute result of the target human body includes: Determine the number of first images in which the target human body is an adult and the number of second images in which the target human body is a child in the sequence of images to be detected; The number of the first images is compared with the number of the second images. When the number of the first images is greater than or equal to the number of the second images, the human attribute result of the target human body is determined to be an adult. When the number of the second images is greater than the number of the first images, the human attribute result of the target human body is determined to be a child.

7. A human attribute recognition device, characterized in that: include: An image acquisition module, used for acquiring an image to be detected; A confidence acquisition module, used to acquire a basic confidence and an auxiliary confidence corresponding to the target human body in the image to be detected, wherein the basic confidence includes a basic adult confidence and a basic child confidence, the auxiliary confidence includes an auxiliary adult confidence and an auxiliary child confidence, the basic adult confidence corresponds to the auxiliary adult confidence, and the basic child confidence corresponds to the auxiliary child confidence; a confidence adjustment module, configured to determine a basic target confidence that satisfies a target condition, and adjust the basic target confidence so that the adjusted basic target confidence is greater than the basic target confidence before the adjustment, wherein the target condition is that the basic confidence is greater than or equal to a preset confidence threshold, the basic target confidence is any one of the basic adult confidence and the basic child confidence, and the adjustment includes assigning a step coefficient to the basic target confidence; A confidence calculation module, used for adding the adjusted basic target confidence and the corresponding auxiliary confidence to obtain a first confidence corresponding to the target human body, and adding the basic confidence other than the basic target confidence and the corresponding auxiliary confidence to obtain a second confidence corresponding to the target human body; The attribute determination module is used to compare the first confidence level with the second confidence level, and determine the attribute information of the target human body in the image to be detected according to the comparison result.

8. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 8.