Human body attribute recognition method, device and equipment, medium and vehicle

By comparing the confidence of adults and children of each image to be detected in the sequence of image to be detected, the human body attribute information is determined, and the number of images is voted on, the problem of low accuracy of human body attribute recognition in complex environments is solved, and a higher recognition accuracy is achieved.

CN119964194APending Publication Date: 2025-05-09BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311481953.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing human attribute recognition scheme has the problem of low accuracy in complex environments, especially when there is occlusion or interference in the interior of the vehicle cockpit, it is difficult to accurately identify human attributes.

Method used

By obtaining the adult confidence and child confidence of the target human body in each image to be detected in the sequence of images to be detected, the attribute information is determined after comparison, including being an adult, being a child or a blurred state, and voting is made based on the number of images to determine the human body attribute results.

Benefits of technology

It improves the accuracy of human body attribute recognition, avoids black and white strategies, and can more accurately identify human body attributes in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964194A_ABST
    Figure CN119964194A_ABST
Patent Text Reader

Abstract

The invention relates to a human body attribute recognition method, device and equipment, a medium and a vehicle, and the method comprises the steps: obtaining a to-be-detected image sequence; for each to-be-detected image in the to-be-detected image sequence, obtaining an adult confidence coefficient and a child confidence coefficient respectively corresponding to a target human body in the to-be-detected image; comparing the adult confidence coefficient, the child confidence coefficient, a preset adult confidence coefficient threshold value and a preset child confidence coefficient threshold value, and determining attribute information of a target human body in the to-be-detected image according to a ratio of comparison results; voting is performed on the attributes of the target human body in the to-be-detected image sequence to obtain a human body attribute result of the target human body, so that the attribute information of the target human body can be divided into three types of adults, children and fuzzy states, and then voting is performed on the attributes of the target human body in the to-be-detected image sequence based on the three types of attribute information; and the accuracy of human body attribute recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image recognition technology, and in particular to a method, device, equipment, medium and vehicle for human attribute recognition. Background Art

[0002] Human attributes are the main basis for identifying the human body. By identifying human attributes, tracking and positioning of the human body can be achieved. In existing human attribute recognition schemes, human attributes are mainly identified and classified based on machine learning models used for detection and classification. For example, when judging whether a human body is an adult or a child, the image data corresponding to the human body is usually input into the machine learning model used for detection and classification. The machine learning model identifies and classifies the image data and outputs the result of whether the human body is an adult or a child.

[0003] However, when the human body is in a complex environment, such as when the human body is inside a vehicle cabin, there are often many obstructions or interferences, as well as drastic changes in environmental factors. In existing human attribute recognition schemes, when the machine learning model identifies the image data and determines that a human body does not belong to an adult, the human body is directly identified as a child. Similarly, when it is determined that a human body does not belong to a child, the human body is directly identified as an adult, that is, a black-and-white strategy is adopted, which has the problem of low accuracy in human attribute recognition. Summary of the invention

[0004] In order to solve the above technical problems, the present disclosure provides a human attribute recognition method, device, equipment, medium and vehicle.

[0005] A first aspect of an embodiment of the present disclosure provides a method for identifying human attributes, the method comprising:

[0006] Obtaining a sequence of images to be detected;

[0007] For each image to be detected in the sequence of images to be detected, obtaining the adult confidence and the child confidence corresponding to the target human body in the image to be detected;

[0008] Comparing the adult confidence with the child confidence to obtain a first comparison result, and at the same time comparing the adult confidence with a preset adult confidence threshold to obtain a second comparison result, and comparing the child confidence with a preset child confidence threshold to obtain a third comparison result;

[0009] Determine attribute information of a target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result, and the third comparison result, wherein the attribute information includes whether the target human body belongs to an adult, a child, or a fuzzy state, and the fuzzy state is used to indicate that it is uncertain whether the target human body belongs to an adult or a child;

[0010] According to the attribute information of the target human body in the image to be detected, the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states is determined, and the human body attribute results of the target human body are determined based on the number of images.

[0011] A second aspect of an embodiment of the present disclosure provides a human attribute recognition device, the device comprising:

[0012] An image acquisition module, used for acquiring a sequence of images to be detected;

[0013] A confidence acquisition module is used to acquire, for each image to be detected in the image sequence to be detected, an adult confidence and a child confidence corresponding to the target human body in the image to be detected;

[0014] A confidence comparison module is used to compare the adult confidence with the child confidence to obtain a first comparison result, and to compare the adult confidence with a preset adult confidence threshold to obtain a second comparison result, and to compare the child confidence with a preset child confidence threshold to obtain a third comparison result;

[0015] A first determination module is used to determine attribute information of a target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result and the third comparison result, wherein the attribute information includes belonging to an adult, belonging to a child, and a fuzzy state, wherein the fuzzy state is used to indicate that it is uncertain whether the target human body belongs to an adult or a child;

[0016] The second determination module is used to determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states according to the attribute information of the target human body in the image to be detected, and determine the human body attribute results of the target human body based on the number of images.

[0017] A third aspect of an embodiment of the present disclosure provides an electronic device, the device comprising:

[0018] Memory;

[0019] Processor; and

[0020] A computer program, wherein the computer program is stored in a memory and is configured to be executed by a processor to implement the human attribute recognition method of the first aspect as described above.

[0021] A fourth aspect of the embodiments of the present disclosure provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the human attribute recognition method as described in the first aspect above is implemented.

[0022] A fifth aspect of an embodiment of the present disclosure provides a vehicle, comprising the electronic device of the third aspect described above.

[0023] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0024] The human attribute recognition method, device, equipment, medium and vehicle provided by the embodiments of the present disclosure can obtain a sequence of images to be detected, and for each image to be detected in the sequence of images to be detected, obtain the adult confidence and child confidence respectively corresponding to the target human body in the image to be detected, compare the adult confidence and the child confidence to obtain a first comparison result, and at the same time compare the adult confidence with a preset adult confidence threshold to obtain a second comparison result, and compare the child confidence with a preset child confidence threshold to obtain a third comparison result, and determine the attribute information of the target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result and the third comparison result. The attribute information includes belonging to an adult, belonging to a child, and being fuzzy. The fuzzy state is used to represent that it is uncertain whether the target human body belongs to an adult or a child. According to the attribute information of the target human body in the image to be detected, the number of images in the image sequence to be detected whose attribute information respectively belongs to an adult, belongs to a child, and is in a fuzzy state is determined, and the human attribute results of the target human body are determined based on the number of images. Thus, the attribute information of the target human body can be divided into three categories: belonging to an adult, belonging to a child, and being fuzzy. Then, based on these three categories of attribute information, the attributes of the target human body in the image sequence to be detected are voted, thereby avoiding the use of a black-and-white strategy to judge and identify the attribute information of the human body, thereby improving the accuracy of human attribute recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0027] Figure 1 is a flow chart of a human attribute recognition method provided by an embodiment of the present disclosure;

[0028] Figure 2 is a flow chart of another human attribute recognition method provided by an embodiment of the present disclosure;

[0029] Figure 3 is a structural schematic diagram of a human attribute recognition device provided by an embodiment of the present disclosure;

[0030] Figure 4 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0033] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0034] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0036] Figure 1 It is a flowchart of a human attribute recognition method provided by an embodiment of the present disclosure. The method can be executed by a human attribute recognition device. The human attribute recognition device can be implemented in software and / or hardware. The human attribute recognition device can be configured in an electronic device, such as a server or a terminal or a server cluster, wherein the terminal can specifically include a mobile phone, a computer or a tablet computer, or a vehicle-mounted terminal.

[0037] like Figure 1 As shown, the human attribute recognition method provided by the embodiment of the present disclosure includes the following steps.

[0038] S110: Acquire a sequence of images to be detected.

[0039] In the embodiment of the present disclosure, the electronic device may acquire a sequence of images to be detected based on an image acquisition device.

[0040] The image sequence to be detected may be an image sequence of the interior of the cabin captured by a vehicle image acquisition device over a period of time, wherein each image to be detected in the image sequence to be detected may include at least one target human body, for example, the target human body may be a passenger in the vehicle cabin.

[0041] The image acquisition device may be at least one camera installed inside the vehicle cabin.

[0042] In some embodiments of the present disclosure, the electronic device may control the image acquisition device to acquire images to be detected in real time, and upload the acquired images to be detected to the electronic device based on the image acquisition device, so that the electronic device acquires a sequence of images to be detected.

[0043] In some other embodiments of the present disclosure, after receiving the human attribute recognition instruction, the electronic device may obtain the to-be-detected image sequence corresponding to the human attribute recognition instruction by loading from a local memory based on the human attribute recognition instruction.

[0044] S120 . For each image to be detected in the sequence of images to be detected, obtain an adult confidence level and a child confidence level corresponding to a target human body in the image to be detected.

[0045] In the disclosed embodiment, the adult confidence can be understood as the probability that the target person is an adult, and the child confidence can be understood as the probability that the target person is a child.

[0046] Specifically, after acquiring the sequence of images to be detected, the electronic device identifies each image to be detected in the sequence of images to be detected based on a preset machine learning model, and obtains the adult confidence and child confidence corresponding to the target human body in the image to be detected based on the recognition result.

[0047] The preset machine learning model may be a model trained using sample images containing a target human body, and used to identify a target human body in an image. In the disclosed embodiment, the preset machine learning model may be composed of a human body detection classification network and a human body key point detection network, wherein the human body detection classification network is used to detect and classify the target human body in the image to be detected, and the human body key point detection network is used to identify the key points of the target human body.

[0048] S130. Compare the adult confidence and the child confidence to obtain a first comparison result, and compare the adult confidence and a preset adult confidence threshold to obtain a second comparison result, and compare the child confidence and a preset child confidence threshold to obtain a third comparison result.

[0049] In the embodiment of the present disclosure, the preset adult confidence threshold is a pre-set confidence threshold for determining whether the target person is an adult.

[0050] Among them, the preset adult confidence threshold is obtained by collecting multiple images of adults in various positions and postures, calculating the adult confidence for each image, arranging the multiple adult confidences corresponding to the multiple images in order from small to large, and determining the preset adult confidence threshold as the target value corresponding to the preset proportion in which the adult confidence is higher than a certain target value in the arrangement result and the number of images in which the adult confidence is higher than a certain target value is greater than the preset proportion.

[0051] Exemplarily, the number of images of adults in various positions and postures collected is 100. When the number of images whose adult confidence values ​​corresponding to the 100 images are greater than 80 accounts for more than 90% of the total number of images, the adult confidence threshold is determined to be 80. It should be noted that this is just an example for ease of understanding and does not limit this solution.

[0052] The preset child confidence threshold is a preset confidence threshold used to determine whether the target person is a child.

[0053] The preset child confidence threshold is obtained by collecting multiple images of children in various positions and postures, calculating the child confidence for each image, arranging the multiple child confidences corresponding to the multiple images in order from small to large, and determining the preset child confidence threshold as the target value corresponding to the child confidence higher than a certain target value in the arrangement result and the proportion of images with child confidence higher than a certain target value greater than the preset proportion.

[0054] S140. Determine attribute information of the target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result, and the third comparison result. The attribute information includes belonging to an adult, belonging to a child, and being in a fuzzy state. The fuzzy state is used to indicate that it is uncertain whether the target human body belongs to an adult or a child.

[0055] In the embodiment of the present disclosure, after obtaining the first comparison result, the second comparison result and the third comparison result, the electronic device determines whether the target human body in the image to be detected is an adult, a child or in a blurred state according to the ratio of the first comparison result, the second comparison result and the third comparison result.

[0056] The condition for belonging to an adult is that the adult confidence in the first comparison result is greater than the child confidence, and the adult confidence is greater than a preset adult confidence threshold.

[0057] The condition for belonging to a child is that the child confidence in the first comparison result is greater than the adult confidence, and the child confidence is greater than a preset child confidence threshold.

[0058] When the first comparison result, the second comparison result, and the third comparison result satisfy neither the condition of belonging to an adult nor the condition of belonging to a child, they are determined to be in an ambiguous state.

[0059] S150, according to the attribute information of the target human body in the image to be detected, determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states, and determine the human body attribute results of the target human body based on the number of images.

[0060] In an embodiment of the present disclosure, after obtaining the attribute information of the target human body in each image to be detected, the electronic device determines the number of images to be detected belonging to adults, the number of images to be detected belonging to children, and the number of images to be detected in a fuzzy state in the image sequence to be detected according to the attribute information of the target human body in the image to be detected, compares the number of images to be detected belonging to adults, the number of images to be detected belonging to children, and the number of images to be detected in a fuzzy state, and determines the human body attribute result of the target human body according to the comparison result and a preset attribute determination rule.

[0061] In some embodiments of the present disclosure, the preset attribute determination rule may be to ignore the number of images to be detected belonging to blurry days, and determine the human attribute results of the target person only based on the comparison result of the number of images to be detected belonging to adults and the number of images to be detected belonging to children.

[0062] In the embodiment of the present disclosure, the implementation method of determining the human attribute result of the target human body according to the number of to-be-detected images belonging to adults and the number of to-be-detected images belonging to children can be set according to specific needs.

[0063] In some examples of the present disclosure, when the number of images to be detected belonging to adults is greater than the number of images to be detected belonging to children, the human attribute results of the target person are determined to be adult; when the number of images to be detected belonging to adults is less than the number of images to be detected belonging to children, the human attribute results of the target person are determined to be child; when the number of images to be detected belonging to adults is equal to the number of images to be detected belonging to children, the human attribute results of the target person are not output.

[0064] In other examples of the present disclosure, when the number of images to be detected belonging to adults is greater than or equal to the number of images to be detected belonging to children, the human attribute result of the target person is determined to be an adult; when the number of images to be detected belonging to adults is less than the number of images to be detected belonging to children, the human attribute result of the target person is determined to be a child.

[0065] In some further examples of the present disclosure, when the number of images to be detected belonging to adults is greater than the number of images to be detected belonging to children, the human attribute result of the target person is determined to be an adult; when the number of images to be detected belonging to adults is less than or equal to the number of images to be detected belonging to children, the human attribute result of the target person is determined to be a child.

[0066] In some other embodiments of the present disclosure, the preset attribute determination rule may be to determine the attribute information corresponding to the images to be detected with the largest number of images among the number of images to be detected belonging to adults, the number of images to be detected belonging to children, and the number of images to be detected in a fuzzy state as the human attribute result of the target human body.

[0067] In some further embodiments of the present disclosure, the preset attribute determination rule may be to directly determine the fuzzy state as the human attribute result of the target human body when the number of images to be detected that are determined to be in a fuzzy state reaches a preset ratio threshold of the total number of images in the image sequence to be detected.

[0068] In the embodiment of the present disclosure, a sequence of images to be detected can be obtained, and for each image to be detected in the sequence of images to be detected, the adult confidence and the child confidence corresponding to the target human body in the image to be detected are obtained, and the adult confidence and the child confidence are compared to obtain a first comparison result. At the same time, the adult confidence and the preset adult confidence threshold are compared to obtain a second comparison result, and the child confidence and the preset child confidence threshold are compared to obtain a third comparison result. According to the ratio of the first comparison result, the second comparison result and the third comparison result, the attribute information of the target human body in the image to be detected is determined, and the attribute information includes the attribute information belonging to an adult, Belonging to children, fuzzy state, the fuzzy state is used to characterize that it is uncertain whether the target human body belongs to an adult or a child. According to the attribute information of the target human body in the image to be detected, the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states respectively is determined, and the human attribute results of the target human body are determined based on the number of images. Thus, the attribute information of the target human body can be divided into three categories: belonging to adults, belonging to children, and fuzzy states. Then, based on these three types of attribute information, the attributes of the target human body in the image sequence to be detected are voted, avoiding the use of a black-and-white strategy to judge and identify the attribute information of the human body, thereby improving the accuracy of human attribute recognition.

[0069] On the basis of the above-mentioned embodiments of the present disclosure, obtaining the adult confidence and child confidence respectively corresponding to the target human body in the image to be detected in S120 may specifically include: inputting the image to be detected into a human body detection and classification network, and detecting and classifying the image to be detected by the human body detection and classification network to obtain a first detection result corresponding to the target human body in the image to be detected, and the first detection result includes a detection frame corresponding to the target human body; determining a regional image corresponding to the target human body based on the detection frame corresponding to the target human body, and inputting the regional image into a human body key point detection network, and performing human body key point recognition on the regional image by the human body key point detection network to obtain a second detection result, and the second detection result includes key point information corresponding to the target human body; inputting the first detection result and the second detection result into a preset adult confidence mapping function to obtain the adult confidence corresponding to the target human body, and at the same time inputting the first detection result and the second detection result into a preset child confidence mapping function to obtain the child confidence corresponding to the target human body.

[0070] In the embodiments of the present disclosure, the human body detection and classification network may be a pre-trained model for detecting and classifying a target human body in an image to be detected, and the human body key point detection network may be a pre-trained network model for identifying and predicting key points of a target human body in an image to be detected.

[0071] In some embodiments of the present disclosure, a human body detection classification network may include a first backbone network, a first shoulder network, and a first head network. Based on the first backbone network, the first shoulder network, and the first head network, backbone features, shoulder features, and head features are extracted from the image to be detected, respectively, to obtain a feature map corresponding to the image to be detected. The feature map after feature extraction is traversed in a candidate frame manner to obtain a detection frame corresponding to each target human body in the image to be detected, and attribute recognition is performed on the detection frame corresponding to each target human body to obtain a detection frame corresponding to each target human body in the image to be detected and basic attributes of the human body, and then the detection frame corresponding to the target human body is input into a human body key point detection network, wherein the human body key point detection network may include a second backbone network, a second shoulder network, and a second head network. Based on the second backbone network, the second shoulder network, and the second head network, backbone features, shoulder features, and head features are extracted from the target human body in the detection frame, respectively, to obtain a feature map corresponding to the target human body in the detection frame, and then the key points of the target human body are identified, and the identification results are matched with preset key points to obtain the human body key point information corresponding to the target human body.

[0072] In the embodiment of the present disclosure, the confidence mapping function is a function that maps the human body detection classification results obtained by the human body detection classification network and the human body key point information obtained by the human body key point detection network into human body attribute confidence.

[0073] The preset adult confidence mapping function and the preset child confidence mapping function are respectively generated based on the test data, using mathematical statistics to calculate the area position of the human body detection frame, the human body trunk length, the human body trunk length kernel density, etc., and are used to calculate the mapping functions for calculating adult confidence and child confidence respectively.

[0074] Specifically, after obtaining the image to be detected, the electronic device first preprocesses the image to be detected, wherein the preprocessing may include image cropping, image normalization processing, image filtering processing, etc., and inputs the preprocessed image to be detected into the human body detection classification network, and the human body detection classification network performs feature extraction on the preprocessed image to be detected to obtain a feature map corresponding to the image to be detected, and then the human body detection classification network performs human body recognition and human body attribute recognition on the feature map to obtain a detection frame corresponding to the target human body in the image to be detected and the human body basic attributes corresponding to the target human body, and performs a regional image corresponding to the detection frame corresponding to the target human body, that is, the regional image of the target human body. The human body is segmented, and the regional image of the target human body is input into the human body key point detection network. The human body key point detection network recognizes the human body key points of the regional image, and matches the recognition result with the preset key points to obtain the key point information corresponding to the target human body. The information of the detection frame corresponding to the target human body, the basic human attributes corresponding to the target human body, and the human body key point information are input into the preset adult confidence mapping function to obtain the adult confidence corresponding to the target human body. At the same time, the information of the detection frame corresponding to the target human body, the basic human attributes corresponding to the target human body, and the human body key point information are input into the preset child confidence mapping function to obtain the child confidence corresponding to the target human body.

[0075] Among them, the basic human attributes include basic adult confidence and basic child confidence. The basic human attributes are the attributes of the target human body obtained by the human body detection and classification network based on the image features of the target human body. Therefore, when determining the human attributes of the target human body, the image features of the target human body can be combined to improve the accuracy of the obtained human body attribute results.

[0076] In the embodiment of the present disclosure, the first detection result obtained by the human body detection classification network and the second detection result obtained by the human body key point detection network are respectively input into a preset adult confidence mapping function and a preset child confidence mapping function, so as to obtain the adult confidence and child confidence corresponding to the target human body, thereby improving the accuracy of the obtained adult confidence and child confidence corresponding to the target human body.

[0077] In an embodiment of the present disclosure, in S140, attribute information of the target human body in the image to be detected is determined based on the ratio of the first comparison result, the second comparison result and the third comparison result, which may specifically include: when the first comparison result is that the adult confidence is greater than the child confidence, and the second comparison result is that the adult confidence is greater than the preset adult confidence threshold, determining that the target human body belongs to an adult; when the first comparison result is that the child confidence is greater than the adult confidence, and the third comparison result is that the child confidence is greater than the preset child confidence threshold, determining that the target human body belongs to a child; when it does not meet other than the above two results, determining that the target human body is in a fuzzy state.

[0078] In the embodiments of the present disclosure, situations other than the above two results may include but are not limited to: the first comparison result is that the adult confidence is greater than the child confidence, but the second comparison result is that the adult confidence is less than or equal to the preset adult confidence threshold; the first comparison result is that the child confidence is greater than the adult confidence, but the third comparison result is that the child confidence is less than or equal to the preset child confidence threshold; the first comparison result is that the adult confidence is less than or equal to the child confidence, but the second comparison result is that the adult confidence is greater than the preset adult confidence threshold; the first comparison result is that the child confidence is less than or equal to the adult confidence, but the third comparison result is that the child confidence is greater than the preset child confidence threshold.

[0079] In the disclosed embodiment, the human attribute information of the target human body can be classified as belonging to an adult, belonging to a child, and being fuzzy according to the first comparison result, the second comparison result, and the third comparison result, thereby improving the accuracy of the obtained human attribute results.

[0080] In an embodiment of the present disclosure, before determining the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and is in a blurred state, and determining the human attribute results of the target human body based on the number of images, the human attribute recognition method may also include: determining the first image to be detected in the image sequence to be detected whose attribute information of the target human body is in a blurred state; removing the first image to be detected from the image sequence to be detected to obtain a second image to be detected.

[0081] Specifically, after obtaining the attribute information of the target human body in each image to be detected, the electronic device removes the first image to be detected whose attribute information of the target human body is blurred according to the attribute information of the target human body in the image to be detected, and obtains the second image to be detected.

[0082] Furthermore, determining the number of images in the image sequence to be detected whose attribute information belongs to an adult, a child, and a fuzzy state, respectively, and determining the human attribute result of the target person based on the number of images can specifically include: determining the number of first images in which the target person belongs to an adult and the number of second images in which the target person belongs to a child in the second image to be detected; comparing the number of first images with the second number of images, and when the number of first images is greater than the number of second images, determining the human attribute result of the target person is an adult, and when the number of second images is greater than the number of first images, determining the human attribute result of the target person is a child.

[0083] In the disclosed embodiments, it is possible to determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states, respectively. Before determining the human attribute results of the target human body based on the number of images, the first image to be detected that affects the human attribute results can be removed, and the human attribute results can be determined based on the second image to be detected after the first image to be detected is removed, thereby improving the rate of voting on the attributes of the target human body in the image sequence to be detected, and at the same time improving the accuracy of obtaining the human attribute results of the target human body.

[0084] In some other embodiments of the present disclosure, the number of images in a sequence of images to be detected whose attribute information belongs to adults, children, and in a blurred state, respectively, is determined, and the human attribute results of the target human body are determined based on the number of images. This may also specifically include: when determining the human attribute results, directly ignoring or skipping the images to be detected whose human attributes are in a blurred state, calculating the number of first images belonging to adults and the number of second images belonging to children, and determining the human attribute results of the target human body based on the size of the number of the first images and the second number of images.

[0085] The specific implementation of determining the human attribute result of the target human body based on the size of the first image number and the second image number is similar to the implementation in the above embodiment, which will not be described in detail here.

[0086] Figure 2 is a flow chart of another method for identifying human attributes provided by an embodiment of the present disclosure. Figure 2 As shown, the human attribute recognition method may specifically include the following steps:

[0087] S210: Acquire a sequence of images to be detected.

[0088] S220 . For each image to be detected in the sequence of images to be detected, obtain an adult confidence level and a child confidence level corresponding to a target human body in the image to be detected.

[0089] S230, comparing the adult confidence and the child confidence to obtain a first comparison result, and at the same time comparing the adult confidence and a preset adult confidence threshold to obtain a second comparison result, and comparing the child confidence and a preset child confidence threshold to obtain a third comparison result.

[0090] S240: When the first comparison result shows that the adult confidence is greater than the child confidence, and the second comparison result shows that the adult confidence is greater than a preset adult confidence threshold, it is determined that the target person is an adult.

[0091] S250: When the first comparison result is that the child confidence is greater than the adult confidence, and the third comparison result is that the child confidence is greater than a preset child confidence threshold, it is determined that the target person is a child.

[0092] S260. In cases other than S240 and S250, it is determined that the target human body is in a fuzzy state.

[0093] S270, determining the first image to be detected whose attribute information of the target human body in the sequence of images to be detected is in a fuzzy state, removing the first image to be detected from the sequence of images to be detected, and obtaining a second image to be detected.

[0094] S280, determining the number of first images in which the target human body is an adult and the number of second images in which the target human body is a child in the second image to be detected.

[0095] S290. Compare the number of first images with the number of second images. When the number of first images is greater than or equal to the number of second images, determine that the human attribute result of the target person is an adult. When the number of second images is greater than the number of first images, determine that the human attribute result of the target person is a child.

[0096] It should be noted that the specific implementation of steps S210-S290 is similar to the specific implementation of the above embodiment and will not be repeated here.

[0097] In the disclosed embodiment, the adult confidence and child confidence of the target human body in the image to be detected can be obtained, and the attribute information of the target human body in the image to be detected is determined based on the adult confidence, the child confidence, the preset adult confidence threshold and the preset child confidence threshold. At the same time, the attribute information of the target human body in the image to be detected is divided into three categories: belonging to adults, belonging to children, and fuzzy. Then, based on these three categories of attribute information, the attributes of the target human body in the image sequence to be detected are voted, thereby avoiding the use of a black-and-white strategy to judge and identify the attribute information of the human body, thereby improving the accuracy of human attribute recognition.

[0098] In an embodiment of the present disclosure, the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and is in a blurred state is determined, and the human attribute results of the target human body based on the number of images can also specifically include: determining the ratio of the number of third images of the first images to be detected whose attribute information of the target human body in the image sequence to be detected is in a blurred state to the total number of images in the image sequence to be detected; when the ratio is greater than a preset ratio threshold, it is determined that the human attribute results of the target human body in the image sequence to be detected are in a blurred state.

[0099] Specifically, the number of the third images of the first image to be detected is determined. When the ratio of the number of the third images to the total number of images in the image sequence to be detected is greater than a preset ratio threshold, such as 70%, it is directly determined that the human attribute results of the target human body in the image sequence to be detected are in a fuzzy state; when the ratio of the number of the third images to the total number of images in the image sequence to be detected is less than or equal to the preset ratio threshold, the first image sequence is removed from the image sequence to be detected to obtain a second image to be detected, and based on the number of first images of the target human body belonging to adults and the number of second images of the target human body belonging to children in the second image sequence to be detected, the human attribute results of the target human body in the image sequence to be detected are determined.

[0100] In the embodiment of the present disclosure, after obtaining the human body attribute results of the target human body, the human body attribute recognition method may also include: obtaining the next frame of image to be detected that is located after the sequence of images to be detected, and deleting the first frame of image to be detected in the sequence of images to be detected; obtaining target attribute information corresponding to the target human body in the next frame of image to be detected; and determining the target human body attribute results corresponding to the target human body based on the attribute information and target attribute information respectively corresponding to the target human body in the images to be detected except the first frame of image in the sequence of images to be detected.

[0101] Among them, the specific implementation method of obtaining the next frame of image to be detected is similar to the implementation method of obtaining the image sequence to be detected in the above embodiment, which will not be repeated here; obtaining the target attribute information corresponding to the target human body in the next frame of image to be detected is similar to the implementation method of determining the attribute information of the target human body in the image to be detected in the above embodiment, which will not be repeated here.

[0102] Specifically, the electronic device obtains the next frame of image to be detected after the sequence of images to be detected, deletes the first frame of image to be detected in the sequence of images to be detected, obtains target attribute information corresponding to the target human body in the next frame of image to be detected, and based on the attribute information and target attribute information respectively corresponding to the target human body in the images to be detected other than the first frame of image to be detected in the sequence of images to be detected, votes on the attributes of the target human body in the target image sequence composed of the images to be detected other than the first frame of image to be detected in the sequence of images to be detected and the next frame of image to be detected, to determine the target human body attribute results corresponding to the target human body, wherein the specific implementation method of voting on the attributes of the target human body in the target image sequence to determine the target human body attribute results corresponding to the target human body is similar to that of voting on the attributes of the target human body in the image sequence to be detected to obtain the human body attribute results of the target human body in the above-mentioned embodiment, and will not be elaborated here.

[0103] In the disclosed embodiment, after acquiring the next frame of the image to be detected, the image sequence to be detected can be updated, and the attributes of the target human body in the updated image sequence to be detected can be voted on to determine the human attribute results of the target human body, so as to ensure the real-time update of the human attributes of the target human body and improve the accuracy of the obtained human attribute results of the target human body.

[0104] Figure 3 : is a schematic diagram of the structure of a human attribute recognition device provided in an embodiment of the present disclosure. The human attribute recognition device in the embodiment of the present disclosure can be set in an electronic device, and the electronic device can be a server or a terminal or a server cluster, wherein the terminal can specifically include a mobile phone, a computer or a tablet computer, and a vehicle-mounted terminal, which is not limited here.

[0105] like Figure 3 As shown, the human attribute recognition device 300 may include an image acquisition module 310 , a confidence acquisition module 320 , a confidence comparison module 330 , a first determination module 340 and a second determination module 350 .

[0106] The image acquisition module 310 may be used to acquire a sequence of images to be detected.

[0107] The confidence acquisition module 320 may be used to acquire, for each image to be detected in the sequence of images to be detected, an adult confidence and a child confidence corresponding to the target human body in the image to be detected.

[0108] The confidence comparison module 330 can be used to compare the adult confidence and the child confidence to obtain a first comparison result, and to compare the adult confidence and the preset adult confidence threshold to obtain a second comparison result, and to compare the child confidence and the preset child confidence threshold to obtain a third comparison result.

[0109] The first determination module 340 can be used to determine the attribute information of the target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result and the third comparison result. The attribute information includes belonging to an adult, belonging to a child, and being in a fuzzy state. The fuzzy state is used to indicate that it cannot be determined whether the target human body belongs to an adult or a child.

[0110] The second determination module 350 can be used to determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states according to the attribute information of the target human body in the image to be detected, and determine the human body attribute results of the target human body based on the number of images.

[0111] In the embodiment of the present disclosure, a sequence of images to be detected can be obtained, and for each image to be detected in the sequence of images to be detected, the adult confidence and the child confidence corresponding to the target human body in the image to be detected are obtained, and the adult confidence and the child confidence are compared to obtain a first comparison result. At the same time, the adult confidence is compared with a preset adult confidence threshold to obtain a second comparison result, and the child confidence is compared with a preset child confidence threshold to obtain a third comparison result. According to the ratio of the first comparison result, the second comparison result and the third comparison result, the attribute confidence of the target human body in the image to be detected is determined. The attribute information includes belonging to adults, belonging to children, and being in a fuzzy state. According to the attribute information of the target human body in the image to be detected, the number of images in the image sequence to be detected whose attribute information respectively belongs to adults, belongs to children, and is in a fuzzy state is determined, and the human attribute results of the target human body are determined based on the number of images. Thus, the attribute information of the target human body can be divided into three categories: belonging to adults, belonging to children, and being in a fuzzy state. Then, based on these three categories of attribute information, the attributes of the target human body in the image sequence to be detected are voted, thereby avoiding the use of a black-and-white strategy to judge and identify the attribute information of the human body, thereby improving the accuracy of human attribute recognition.

[0112] In some embodiments of the present disclosure, the confidence acquisition module 320 can be specifically used to input the image to be detected into a human body detection and classification network, and the human body detection and classification network detects and classifies the image to be detected to obtain a first detection result corresponding to the target human body in the image to be detected, and the first detection result includes a detection frame corresponding to the target human body; based on the detection frame corresponding to the target human body, a regional image corresponding to the target human body is determined, and the regional image is input into a human body key point detection network, and the human body key point detection network performs human body key point recognition on the regional image to obtain a second detection result, and the second detection result includes key point information corresponding to the target human body; the first detection result and the second detection result are input into a preset adult confidence mapping function to obtain the adult confidence corresponding to the target human body, and at the same time, the first detection result and the second detection result are input into a preset child confidence mapping function to obtain the child confidence corresponding to the target human body.

[0113] In some embodiments of the present disclosure, the first determination module 340 may specifically include determining that the target human body belongs to an adult when the first comparison result is that the adult confidence is greater than the child confidence, and the second comparison result is that the adult confidence is greater than a preset adult confidence threshold; determining that the target human body belongs to a child when the first comparison result is that the child confidence is greater than the adult confidence, and the third comparison result is that the child confidence is greater than the preset child confidence threshold; and determining that the target human body belongs to a fuzzy state when neither of the above two results is met.

[0114] In some embodiments of the present disclosure, the human attribute recognition device 300 may further include an image screening module.

[0115] The image screening module can be used to determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and is in a fuzzy state, and before determining the human attribute results of the target human body based on the number of images, determine the first image to be detected in the image sequence to be detected whose attribute information of the target human body is in a fuzzy state; remove the first image to be detected from the image sequence to be detected to obtain the second image to be detected.

[0116] In some embodiments of the present disclosure, the second determination module 350 can be specifically used to determine the number of first images in which the target human body in the second image to be detected is an adult and the number of second images in which the target human body is a child; compare the number of first images with the number of second images, and when the number of first images is greater than or equal to the number of second images, determine that the human attribute result of the target human body is an adult, and when the number of second images is greater than the number of first images, determine that the human attribute result of the target human body is a child.

[0117] In some embodiments of the present disclosure, the human attribute recognition device 300 may further include a third determination module.

[0118] The third determination module can be used to determine the ratio of the number of third images of the first images to be detected whose attribute information of the target human body in the image sequence to be detected is in a fuzzy state to the total number of images in the image sequence to be detected; when the ratio is greater than a preset ratio threshold, it is determined that the human body attribute result of the target human body in the image sequence to be detected is in a fuzzy state.

[0119] In some embodiments of the present disclosure, the human attribute recognition device 300 may further include a fourth determination module.

[0120] The fourth determination module can be used to obtain the next frame of image to be detected that is located after the sequence of images to be detected after obtaining the human body attribute results of the target human body, and delete the first frame of image to be detected in the sequence of images to be detected; obtain the target attribute information corresponding to the target human body in the next frame of image to be detected; based on the attribute information and target attribute information corresponding to the target human body in the images to be detected except the first frame of image to be detected in the sequence of images to be detected, determine the target human body attribute results corresponding to the target human body.

[0121] It should be noted that Figure 3 The human attribute recognition device 300 shown can execute each step in the above method embodiment and realize each process and effect in the above method embodiment, which will not be described in detail here.

[0122] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown.

[0123] In the disclosed embodiment, Figure 4 The electronic device shown may be a server or a terminal or a server cluster, wherein the terminal may specifically include a mobile phone, a computer or a tablet computer, or a vehicle-mounted terminal, which is not limited here.

[0124] like Figure 4 As shown, the electronic device may include a processor 410 and a memory 420 storing computer program instructions.

[0125] Specifically, the processor 410 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0126] The memory 420 may include a large capacity memory for information or instructions. By way of example and not limitation, the memory 420 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 420 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 420 may be inside or outside the integrated gateway device. In a particular embodiment, the memory 420 is a non-volatile solid-state memory. In a particular embodiment, the memory 420 includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (Electrically Erasable Programmable ROM, EEPROM), an electrically rewritable ROM (EAROM) or a flash memory, or a combination of two or more of these.

[0127] The processor 410 reads and executes the computer program instructions stored in the memory 420 to perform the steps of the human attribute recognition method provided in the embodiment of the present disclosure.

[0128] In one example, the electronic device may further include a transceiver 430 and a bus 440. Figure 4 As shown, the processor 410, the memory 420 and the transceiver 430 are connected via a bus 440 and communicate with each other.

[0129] The bus 440 includes hardware, software, or both. For example, but not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 440 may include one or more buses.

[0130] The embodiments of the present disclosure further provide a computer-readable storage medium, which may store a computer program. When the computer program is executed by a processor, the processor implements the human attribute recognition method provided by the embodiments of the present disclosure.

[0131] The above-mentioned storage medium may, for example, include a memory 420 of computer program instructions, and the above-mentioned instructions may be executed by the processor 410 of the electronic device to complete the human attribute recognition method provided in the embodiment of the present disclosure. Optionally, the storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (Random Access Memory, RAM), a compact disc read-only memory (Compact Disc ROM, CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0132] The embodiments of the present disclosure also provide a vehicle, which includes electronic equipment and can implement the various processes and effects in the above embodiments of the present disclosure, which will not be elaborated here.

[0133] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0134] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A human attribute recognition method, characterized in that: The method comprises: Obtaining a sequence of images to be detected; For each image to be detected in the sequence of images to be detected, obtaining an adult confidence level and a child confidence level respectively corresponding to a target human body in the image to be detected; Comparing the adult confidence with the child confidence to obtain a first comparison result, and at the same time comparing the adult confidence with a preset adult confidence threshold to obtain a second comparison result, and comparing the child confidence with a preset child confidence threshold to obtain a third comparison result; Determine attribute information of the target human body in the image to be detected according to the ratio of the first comparison result, the second comparison result and the third comparison result, wherein the attribute information includes belonging to an adult, belonging to a child, and a fuzzy state, wherein the fuzzy state is used to indicate that it is not possible to determine whether the target human body belongs to an adult or a child; According to the attribute information of the target human body in the image to be detected, the number of images in the image sequence to be detected whose attribute information belongs to an adult, a child, and a fuzzy state is determined, and the human body attribute result of the target human body is determined based on the number of images.

2. The method according to claim 1, characterized in that The obtaining of the adult confidence level and the child confidence level respectively corresponding to the target human body in the image to be detected includes: Inputting the image to be detected into a human body detection and classification network, the human body detection and classification network detects and classifies the image to be detected, and obtains a first detection result corresponding to a target human body in the image to be detected, wherein the first detection result includes a detection frame corresponding to the target human body; Determine a region image corresponding to the target human body based on a detection frame corresponding to the target human body, input the region image into a human key point detection network, and use the human key point detection network to perform human key point recognition on the region image to obtain a second detection result, wherein the second detection result includes key point information corresponding to the target human body; The first detection result and the second detection result are input into a preset adult confidence mapping function to obtain the adult confidence corresponding to the target human body. At the same time, the first detection result and the second detection result are input into a preset child confidence mapping function to obtain the child confidence corresponding to the target human body.

3. The method according to claim 1, characterized in that The determining, according to the ratio of the first comparison result, the second comparison result and the third comparison result, the attribute information of the target human body in the image to be detected includes: When the first comparison result is that the adult confidence is greater than the child confidence, and the second comparison result is that the adult confidence is greater than the preset adult confidence threshold, it is determined that the target human body is an adult; When the first comparison result is that the child confidence is greater than the adult confidence, and the third comparison result is that the child confidence is greater than the preset child confidence threshold, it is determined that the target person is a child; When neither of the above two results is met, it is determined that the target human body is in a fuzzy state.

4. The method according to claim 1, characterized in that: Before determining the number of images in the to-be-detected image sequence whose attribute information belongs to an adult, a child, or a fuzzy state, and determining the human attribute result of the target human body based on the number of images, the method further includes: Determine the first image to be detected in the sequence of images to be detected whose attribute information of the target human body is in a fuzzy state; The first image to be detected is removed from the sequence of images to be detected to obtain a second image to be detected.

5. The method according to claim 4, characterized in that The step of determining the number of images in the to-be-detected image sequence whose attribute information respectively belongs to an adult, belongs to a child, and is in a fuzzy state, and determining the human attribute result of the target human body based on the number of images includes: Determine the number of first images in which the target human body is an adult and the number of second images in which the target human body is a child in the second image to be detected; The number of the first images is compared with the number of the second images. When the number of the first images is greater than or equal to the number of the second images, the human attribute result of the target human body is determined to be an adult. When the number of the second images is greater than the number of the first images, the human attribute result of the target human body is determined to be a child.

6. The method according to claim 1, characterized in that The step of determining the number of images in the to-be-detected image sequence whose attribute information respectively belongs to an adult, belongs to a child, and is in a fuzzy state, and determining the human attribute result of the target human body based on the number of images includes: Determine the ratio of the number of third images of the first to-be-detected images whose attribute information of the target human body is in a fuzzy state in the to-be-detected image sequence to the total number of images in the to-be-detected image sequence; When the ratio is greater than a preset ratio threshold, it is determined that the human attribute result of the target human body in the to-be-detected image sequence is in a fuzzy state.

7. The method according to claim 1, characterized in that After obtaining the human attribute result of the target human body, the method further includes: Acquire the next frame of the image to be detected after the sequence of images to be detected, and delete the first frame of the image to be detected in the sequence of images to be detected; Obtaining target attribute information corresponding to the target human body in the next frame of the image to be detected; Based on the attribute information respectively corresponding to the target human body in the images to be detected except the first frame of the image to be detected in the sequence of images to be detected and the target attribute information, a target human body attribute result corresponding to the target human body is determined.

8. A human attribute recognition device, characterized in that: include: An image acquisition module, used for acquiring a sequence of images to be detected; A confidence acquisition module, for acquiring, for each image to be detected in the sequence of images to be detected, an adult confidence and a child confidence corresponding to a target human body in the image to be detected; A confidence comparison module, configured to compare the adult confidence with the child confidence to obtain a first comparison result, and to compare the adult confidence with a preset adult confidence threshold to obtain a second comparison result, and to compare the child confidence with a preset child confidence threshold to obtain a third comparison result; A first determination module, configured to determine attribute information of the target human body in the image to be detected according to a ratio of the first comparison result, the second comparison result, and the third comparison result, wherein the attribute information includes whether the target human body belongs to an adult, a child, or a fuzzy state, wherein the fuzzy state is used to indicate that it is not possible to determine whether the target human body belongs to an adult or a child; The second determination module is used to determine the number of images in the image sequence to be detected whose attribute information belongs to adults, children, and fuzzy states according to the attribute information of the target human body in the image to be detected, and determine the human body attribute results of the target human body based on the number of images.

9. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 9.

Citation Information

Cited By

  • Method for correcting attribute detection result of object and electronic equipment

    CN121170420A