Method, device and equipment for determining area of external frame of human head, medium and vehicle
By detecting key points of the human body and adjusting the mapping coefficients, the problem of inaccurate detection of the area of the external frame of the human head in the cockpit is solved, and more accurate depth estimation is achieved.
Patent Information
- Application Number
- CN202410339624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-23
AI Technical Summary
The existing technology for detecting the area of the external frame of a person's head in the cockpit is inaccurate due to differences in height, affecting the accuracy of tasks such as anti-collision head alarms.
By using the key points of the human body of the person to be predicted to match the human body with the seat, the mapping coefficient and scaling coefficient are obtained, the initial torso length and the area of the head external frame are adjusted, and the scaling coefficient lookup table is used for accurate calculation.
Improves the accuracy of the bounding box area of the human head, ensuring accurate depth estimation for tasks such as head collision avoidance.
Smart Images

Figure CN120689389A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, electronic device, storage medium, and vehicle for determining the area of a human head circumference frame. Background Art
[0002] Inside cabins like cars and airplanes, passenger safety is paramount. To prevent collisions and head injuries, real-time monitoring of the passenger's head's external frame area is crucial. This monitoring can be used in a variety of applications, such as head collision warnings and seatbelt reminders.
[0003] Currently, related technologies within the cockpit typically use the head bounding box area derived from a human head detection and classification network as the final result for tasks like head collision avoidance alarms. However, when a head is at the same actual depth, two people with significantly different heights will have completely different head bounding box areas. The human head detection and classification network will predict completely different depths for these two people, resulting in inaccurate head bounding box area results at the same actual depth, significantly different from the actual head area. Summary of the Invention
[0004] The present disclosure provides a method, device, electronic device, storage medium, and vehicle for determining the area of a human head circumference frame.
[0005] According to the first aspect of the present disclosure, a method for determining the area of a human head circumscribed frame is provided, comprising: using the human body key points of a person object to be predicted to perform body seat pairing on the person object to be predicted, determining the seat position of the person object to be predicted, and obtaining a mapping coefficient corresponding to the seat position of the person object to be predicted, the mapping coefficient representing the mapping relationship between the torso length and the height of the person object to be predicted at the seat position; adjusting the initial torso length of the person object to be predicted to a standard torso length in a normal sitting posture, and using the mapping coefficient to determine the actual height corresponding to the standard torso length, the initial torso length being obtained using the human body key points; calling a scaling coefficient lookup table, and using the mapping relationship between height and head area in the scaling coefficient lookup table to determine the scaling coefficient corresponding to the actual height; multiplying the scaling coefficient by the initial human head circumscribed frame area of the person object to be predicted to determine the target human head circumscribed frame area of the person object to be predicted.
[0006] In some embodiments of the present disclosure, before the step of performing body seat pairing on the person to be predicted using the body key points of the person to be predicted, the method includes: obtaining the vertex coordinates of the initial body circumference frame of the person to be predicted in the input image; using the vertex coordinates of the initial body circumference frame to crop the input image according to the initial body circumference frame to obtain the body image of the person to be predicted in the input image; performing key point detection on the body image to obtain the body key points of the person to be predicted; using the body key points of the person to be predicted to perform body seat pairing on the person to be predicted, including: determining the hip point among the body key points of the person to be predicted; and performing body seat pairing on the person to be predicted according to the hip point among the body key points.
[0007] In some embodiments of the present disclosure, the initial torso length of the person object to be predicted is adjusted to the standard torso length in a normal sitting posture, and the mapping coefficient is used to determine the actual height corresponding to the standard torso length, including: determining the head vertex and crotch point among the key points of the human body of the person object to be predicted, and determining the key point straight-line distance between the head vertex and the crotch point, and using the key point straight-line distance as the initial torso length of the person object to be predicted; using the torso length repair mapping function to correct the initial torso length to the standard torso length in a normal sitting posture; multiplying the standard torso length by the mapping coefficient to obtain the actual height corresponding to the standard torso length.
[0008] In some embodiments of the present disclosure, using a torso length repair mapping function to correct the initial torso length to a standard torso length in a normal sitting posture includes: using the human body key points and human body key point coordinates of the human body object to be predicted to determine the sitting posture of the human body object to be predicted; determining the torso repair mapping function corresponding to the sitting posture; and using the torso repair mapping function corresponding to the sitting posture of the human body object to be predicted to correct the initial torso length to the standard torso length in a normal sitting posture.
[0009] In some embodiments of the present disclosure, a method for determining a mapping coefficient corresponding to a seat position includes: obtaining a preset number of subjects to be trained, and determining the standard trunk lengths and actual heights of the preset number of subjects to be trained in a normal sitting posture at the seat position to be trained; determining the ratio of the standard trunk length to the actual height of each subject to be trained in the preset number of subjects to be trained, using the ratio as the mapping coefficient of each subject to be trained, and counting the mapping coefficient corresponding to each subject to be trained to obtain a statistical value of the mapping coefficient; averaging the statistical values of the mapping coefficients to determine the mapping coefficient at the seat position to be trained.
[0010] In some embodiments of the present disclosure, the process of generating a scaling coefficient lookup table includes: determining subjects to be trained with different actual heights at preset height intervals within a preset height range, and obtaining the actual heights and actual head areas corresponding to subjects to be trained with different actual heights; counting the to-be-trained statistical values of the actual head areas of the subjects to be trained at the same actual height within the preset height range, the to-be-trained statistical values including the target statistical values of the actual head areas at the target height; respectively determining the ratio of the target statistical value of the actual head areas at each actual height in different actual heights to the to-be-trained statistical value, and using the ratio as the scaling coefficient of the actual height corresponding to the to-be-trained statistical value of the actual head area at each actual height; and tabulating the scaling coefficients of the actual heights corresponding to the to-be-trained statistical value of the actual head area at each actual height to generate a scaling coefficient lookup table.
[0011] According to a second aspect of the present disclosure, a device for determining the area of a human head circumscribed frame is provided, comprising:
[0012] a matching unit, configured to use the key points of the human body of the person to be predicted to perform body-seat pairing on the person to be predicted, determine the seat position of the person to be predicted, and obtain a mapping coefficient corresponding to the seat position of the person to be predicted, the mapping coefficient representing a mapping relationship between the torso length and height of the person to be predicted at the seat position;
[0013] a height determination unit, configured to adjust the initial torso length of the person to be predicted to a standard torso length in a normal sitting posture, and determine the actual height corresponding to the standard torso length using a mapping coefficient, wherein the initial torso length is obtained using key points of the human body;
[0014] A calling unit, configured to call a scaling factor lookup table and determine a scaling factor corresponding to the actual height by using a mapping relationship between height and head area in the scaling factor lookup table;
[0015] The area determination unit is used to multiply the scaling factor by the initial head circumference area of the person object to be predicted to determine the target head circumference area of the person object to be predicted.
[0016] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0020] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0021] According to a fifth aspect of the present disclosure, a vehicle is provided, comprising the apparatus for determining the area of a circumscribed frame of a human head as described in the second aspect or the electronic device as described in the third aspect.
[0022] The present disclosure provides a method, device, electronic device, storage medium and vehicle for determining the area of a head external frame. By using the key points of the human body of the person to be predicted, the person to be predicted is paired with the seat of the person to be predicted, and the seat position of the person to be predicted is determined. The mapping coefficient corresponding to the seat position of the person to be predicted is obtained. The mapping coefficient represents the mapping relationship between the torso length and height of the person to be predicted at the seat position. The initial torso length of the person to be predicted is adjusted to the standard torso length in a normal sitting position, and the actual height corresponding to the standard torso length is determined using the mapping coefficient. The initial torso length is the standard torso length calculated using the key points of the human body. Obtained; call the scaling factor lookup table, and use the mapping relationship between height and head area in the scaling factor lookup table to determine the scaling factor corresponding to the actual height; multiply the scaling factor by the initial head bounding box area of the person to be predicted to determine the target head bounding box area of the person to be predicted, and adjust the initial head bounding box area of the person to be predicted by using the height and torso length of the person to be predicted, so as to avoid the problem of inaccurate detection of the initial head bounding box area due to different heights at the same depth, improve the accuracy of the head bounding box area, and facilitate more accurate depth estimation for subsequent tasks such as head collision avoidance.
[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0025] Figure 1 A schematic flow chart of a method for determining the area of a human head circumference frame provided in an embodiment of the present disclosure;
[0026] Figure 2 A flowchart of another method for determining the area of a human head circumference frame provided by an embodiment of the present disclosure;
[0027] Figure 3A schematic diagram of obtaining the vertex coordinates of an initial human body circumference frame and an initial head circumference frame provided by an embodiment of the present disclosure;
[0028] Figure 4 A schematic diagram of obtaining coordinates of key points of a human body provided in an embodiment of the present disclosure;
[0029] Figure 5 A schematic diagram of the distribution of key points of a human body provided in an embodiment of the present disclosure;
[0030] Figure 6 A schematic structural diagram of a device for determining the area of a human head circumference frame provided in an embodiment of the present disclosure;
[0031] Figure 7 A schematic block diagram of an exemplary electronic device 700 provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0033] In order to solve the problems in the related art, the present invention adjusts the initial head circumference frame area of the person to be predicted by using the height and torso length of the person to be predicted, thereby avoiding the problem of inaccurate initial head circumference frame area detection due to different heights at the same depth, improving the accuracy of the head circumference frame area, and facilitating more accurate depth estimation for subsequent tasks such as head collision avoidance.
[0034] The following describes a method, device, electronic device, storage medium, and vehicle for determining the area of a human head circumference frame according to embodiments of the present disclosure with reference to the accompanying drawings.
[0035] Figure 1 This is a flow chart of a method for determining the area of a human head's external frame provided in an embodiment of the present disclosure. This method can be executed by an electronic device, specifically, a PC, server, or other electronic device. This method can be applied to scenarios where a human head is recognized and located in an image or video, specifically, in a car's smart cockpit or security monitoring scenario, and is not limited in this embodiment of the present disclosure. Figure 1 As shown, the method includes:
[0036] Step 101: Using the key points of the human body of the person to be predicted, the person to be predicted is paired with the seat of the person to be predicted, the seat position of the person to be predicted is determined, and the mapping coefficient corresponding to the seat position of the person to be predicted is obtained. The mapping coefficient represents the mapping relationship between the torso length and height of the person to be predicted at the seat position.
[0037] In some embodiments, the person object to be predicted refers to any person that needs to be detected in the input image collected by the acquisition device. In the present disclosure, the person object to be predicted can be a person who can be identified in front of a surveillance camera, or a user who is using a device that requires head tracking or anti-collision functions (such as a virtual reality helmet).
[0038] In this disclosure, the key points of the human body of the human object to be predicted refer to the tenting key points within the initial human body bounding box of the human object to be predicted. The human body bounding box refers to a geometric model used to represent the external contour of the human body and estimate and identify the size and position of the human body. The human body bounding box can be a simplified human body shape composed of a series of line segments, which is used to describe the basic shape and size of the human body. In this disclosure, the human body bounding box takes the form of an external rectangular box surrounding the human body.
[0039] In the present disclosure, the vertex coordinates of the initial human body bounding box can be obtained by detecting the input image through a human head detection and classification network. The initial human body bounding box can be directly formed using the detected vertex coordinates of the initial human body bounding box. The human key points in the initial human body bounding box are obtained by cropping the input image using the initial human body bounding box to obtain a human image, and then detecting the human image using a human key point detection network.
[0040] Using the detected key points of the person to be predicted, the person can be paired with a seat, thereby determining the current seat position of the person to be predicted. Since each seat position has a corresponding mapping coefficient, the mapping coefficient of the person to be predicted at the current seat position can be determined based on the seat position of the person to be predicted. The mapping coefficients are pre-trained, and their specific values are not limited in the present embodiments.
[0041] Step 102: Adjust the initial torso length of the person to be predicted to the standard torso length in a normal sitting posture, and use the mapping coefficient to determine the actual height corresponding to the standard torso length. The initial torso length is obtained using the key points of the human body.
[0042] In some embodiments, the human key points detected by the human key point detection network can be used to determine a first initial torso length, where the initial torso length refers to the distance between the crotch point and the head vertex of the human object to be predicted.
[0043] It should be noted that, since the position of the camera is fixed, and the person to be predicted may have sitting postures such as bending forward, lying down, etc., the initial torso length directly obtained using the key points of the human body may be inaccurate. Therefore, the present disclosure can use the torso length repair mapping function to adjust the initial torso length to obtain the standard torso length of the person to be predicted in a normal sitting posture. The normal sitting posture refers to the sitting posture in which the head of the person to be predicted is kept upright, the eyes are looking straight ahead, the waist is straight, and the person is in close contact with the seat back. After obtaining the adjusted standard torso length, the mapping coefficient obtained above under the current seat position can be used, and the mapping relationship between the torso length and the height in the mapping coefficient can be used to determine the actual height corresponding to the standard torso length of the current person to be predicted.
[0044] Step 103: Calling a scaling factor lookup table, and using the mapping relationship between height and head area in the scaling factor lookup table to determine the scaling factor corresponding to the actual height.
[0045] In some embodiments, the scaling factor lookup table is pre-trained. The present disclosure retrieves the scaling factor corresponding to the actual height of the person being predicted by calling the scaling factor lookup table and querying the relationship table between height and scaling factor. This scaling factor is used to adjust the initial head bounding box area.
[0046] Step 104: multiply the scaling factor by the initial head bounding box area of the person object to be predicted to determine the target head bounding box area of the person object to be predicted.
[0047] In some embodiments, the scaling factor obtained above is combined with the initial head bounding box area to obtain a target head bounding box area for the person to be predicted. The target head bounding box area is the head bounding box area obtained by combining the actual height of the person to be predicted, and the target head bounding box area is similar to the actual head area of the person to be predicted. The vertex coordinates of the initial head bounding box can be obtained by detecting the input image through a human head detection and classification network. The initial head bounding box area can be directly calculated using the detected vertex coordinates of the initial head bounding box.
[0048] In summary, the technical solution provided by the present disclosure utilizes the key points of the human body of the person to be predicted to perform body seat pairing on the person to be predicted, determine the seat position of the person to be predicted, and obtain the mapping coefficient corresponding to the seat position of the person to be predicted, where the mapping coefficient represents the mapping relationship between the torso length and height of the person to be predicted in the seat position; adjusts the initial torso length of the person to be predicted to the standard torso length in a normal sitting position, and uses the mapping coefficient to determine the actual height corresponding to the standard torso length, where the initial torso length is obtained using the key points of the human body; calls a scaling coefficient lookup table, and uses the mapping relationship between height and head area in the scaling coefficient lookup table to determine the scaling coefficient corresponding to the actual height; multiplies the scaling coefficient by the initial head bounding box area of the person to be predicted to determine the target head bounding box area of the person to be predicted, thereby adjusting the initial head bounding box area of the person to be predicted using the height and torso length of the person to be predicted, avoiding the problem of inaccurate detection of the initial head bounding box area due to different heights at the same depth, improving the accuracy of the head bounding box area, and facilitating more accurate depth estimation for subsequent tasks such as head collision avoidance.
[0049] Figure 2 This is a flow chart of another method for determining the area of a human head circumference frame provided in an embodiment of the present disclosure. The method can be executed by an electronic device, specifically, by an electronic device such as a PC or a server. Figure 2 based on Figure 1 In the embodiment shown, step 102 is further defined. Figure 2 In the embodiment shown, step 102 includes step 202, step 203 and step 204. Figure 2 As shown, the method may include:
[0050] Step 201: using the key points of the human body of the person to be predicted, performing body-seat pairing on the person to be predicted, determining the seat position of the person to be predicted, and obtaining the mapping coefficient corresponding to the seat position.
[0051] In some embodiments, before using the human body key points of the person to be predicted to perform body seat pairing on the person to be predicted, the method disclosed herein includes: obtaining vertex coordinates of an initial body bounding box of the person to be predicted in an input image; cropping the input image according to the initial body bounding box using the vertex coordinates of the initial body bounding box to obtain a body image of the person to be predicted in the input image; and performing key point detection on the body image to obtain the human body key points of the person to be predicted. Using the human body key points of the person to be predicted to perform body seat pairing on the person to be predicted includes: determining a crotch point among the human body key points of the person to be predicted; and performing body seat pairing on the person to be predicted based on the crotch point among the human body key points.
[0052] Among them, while obtaining the vertex coordinates of the initial human body circumscribed frame of the person object to be predicted in the input image, the present disclosure can also obtain the vertex coordinates of the initial head circumscribed frame of the person object to be predicted in the input image, so as to use the vertex coordinates of the initial head circumscribed frame to determine the area of the initial head circumscribed frame of the person object to be predicted.
[0053] Among them, the vertex coordinates of the initial human body external bounding box and the vertex coordinates of the initial human body external bounding box can be the vertex coordinates of the upper left corner and the vertex coordinates of the lower right corner of the external bounding box of the human body and head, and a rectangular frame can be formed by the vertex coordinates of the upper left corner and the vertex coordinates of the lower right corner.
[0054] Specifically, the human head detection and classification network is a deep learning network built for detection and classification scenarios, usually composed of a backbone network (backbone), a shoulder network (neck), and a head network (head), which is used to detect whether there is a human body or a human head and return the vertex coordinates of the human body's bounding box and the vertex coordinates of the head's bounding box. In the present disclosure, before processing the input image through the human head detection and classification network, it is necessary to train a human head detection and classification network through a large amount of training data. These training data can be various image or video data sets, which contain different human and head postures, different backgrounds and lighting conditions, etc. Figure 3The diagram shown is a schematic diagram of obtaining the vertex coordinates of the initial human body external bounding box and the vertex coordinates of the initial head external bounding box. Before the input image is input into the human head detection and classification network, the input image (original image) needs to be preprocessed. The preprocessing may include scaling, cropping, normalization and other operations to facilitate the human head detection and classification network to better extract features. After that, the human head detection and classification network is used to extract features from the preprocessed input image through the backbone network (backbone), shoulder network (neck), and head network (head) in sequence, and the output of the head network is parsed and processed through post-processing, such as non-maximum suppression (NMS) and other operations to eliminate overlapping detection frames, and return the vertex coordinates of the initial human body external bounding box (human body frame coordinates) and the vertex coordinates of the initial head external bounding box (head frame coordinates).
[0055] Since a rectangular frame can be formed by the vertex coordinates of the initial human body circumference frame, the present disclosure can use a preset cropping function to input the input image to be cropped and the obtained vertex coordinates of the initial human body circumference frame into this cropping function, thereby directly obtaining the cropped human body image.
[0056] Specifically, the present disclosure can determine the image area to be cropped using the vertex coordinates of the initial human body circumscribed frame through a cropping function set in OpenCV, and read the input image to be cropped to perform image cropping.
[0057] It should be noted that the cropping function preset in the present disclosure can be set according to the actual situation of the input image and the actual work experience of the staff, and is not limited in the embodiments of the present disclosure.
[0058] After obtaining the cropped human body image, the present disclosure can obtain multiple human body key points and coordinates of the multiple human body key points in the human body image through a human body key point detection network.
[0059] The human key point detection network is a deep learning network built for detection scenarios. It can be composed of a backbone network, a shoulder network (neck), and a head network (head). It is used to detect whether there are human key points and return the coordinates of the human key points. In the present disclosure, before processing human images through the human key point detection network, it is necessary to train a human key point detection network through a large amount of training data. These training data can be various human images, including annotated and unannotated human images of different scales, lighting conditions, and backgrounds. Figure 4As shown, a schematic diagram of obtaining the coordinates of human key points is shown. Before the human key point detection network detects the human image (original image), the input image needs to be cropped according to the size of the human body external frame obtained above to obtain the human body image, and the human body image is preprocessed. The preprocessing may include operations such as scaling, cropping, and normalization to facilitate the human key point detection network to better extract features. The preprocessed human body image is then sequentially subjected to feature extraction through the backbone network (backbone), shoulder network (neck), and head network (head). The output of the head network is then parsed and processed through post-processing, such as non-maximum suppression (NMS) and other operations to eliminate overlapping key points, thereby determining the human body key points and returning the human body key point coordinates corresponding to the human body key points.
[0060] like Figure 5 A schematic diagram of the distribution of key points of the human body is shown, wherein point 1 is the vertex of the head in the present disclosure, point 17 is the crotch point in the present disclosure, and points 16 and 18 are the hip points in the present disclosure.
[0061] Since the human body key points include the hip point, the human body key points in the initial human body circumscribed frame of the person to be predicted are used to perform body seat pairing on the person to be predicted, determine the seat position of the person to be predicted, and obtain the mapping coefficient corresponding to the seat position, including: determining the hip point of the human body key point in the initial human body circumscribed frame; determining the seat area corresponding to the hip point and the seat position corresponding to the seat area. In other words, the present disclosure can determine the hip area of the person to be predicted by the hip point of the person to be predicted, and determine the seat position of the person to be predicted by judging the seat area where the hip area is located. Among them, for the determination of the seat position, the present disclosure is not limited to using only the hip point as the human body key point for determination, and can also use other surrounding human body key points for auxiliary judgment to more accurately determine the seat position of the person to be predicted. The specific implementation method is based on the actual situation and is not limited in the embodiments of the present disclosure.
[0062] Before utilizing the key points of the human body of the person to be predicted, performing body-seat pairing on the person to be predicted, determining the seat position of the person to be predicted, and obtaining the mapping coefficient corresponding to the seat position, the present disclosure includes a method for determining the mapping coefficient corresponding to the seat position, comprising: obtaining a preset number of persons to be trained, and determining the standard trunk length and actual height of the preset number of persons to be trained in a normal sitting posture at the seat position to be trained; respectively determining the ratio of the standard trunk length to the actual height of each person to be trained in the preset number of persons to be trained, using the ratio as the mapping coefficient of each person to be trained, and counting the mapping coefficient corresponding to each person to be trained to obtain the statistical value of the mapping coefficient; averaging the statistical values of the mapping coefficient to determine the mapping coefficient at the seat position to be trained.
[0063] In an optional embodiment of the present disclosure, a preset number of 1000 training subjects is used as an example. The present disclosure recruits 1000 people, including tall, short, fat, thin, male, female, old and young. Data is collected from the 1000 training subjects in a normal sitting position on each seat in the actual vehicle. The standard torso length b of each training subject in a normal sitting position is calculated through the human head detection network. The mapping coefficient between the standard torso length in a normal sitting position and the actual height of the model is: Where h is the actual height of the training subject. The mapping coefficients of the normal sitting torso length and actual height of these 1000 training subjects are averaged to obtain the mapping coefficient R, which represents the standard torso length in the normal sitting position for that seat. Since there may be multiple seats in the smart cockpit, the mapping coefficient R, which represents the standard torso length in the normal sitting position for each seat, must be calculated separately. The specific values may be the same or different and are not limited in this embodiment.
[0064] Step 202: Determine the head vertex and crotch point among the key points of the human body of the person object to be predicted, and determine the key point straight-line distance between the head vertex and the crotch point, and use the key point straight-line distance as the initial torso length of the person object to be predicted.
[0065] In some embodiments, since the key points of the human body determined by the human body key point detection network may include the head vertex and the crotch point, Figure 5 Based on the coordinates of the head vertex and the crotch point, the distance calculation formula can be used to calculate the key point straight-line distance between the head vertex and the crotch point, and the key point straight-line distance can be used as the initial torso length b of the person to be predicted. The distance calculation formula can be shown in Formula 1:
[0066] \begin{equation}b=\sqrt{{(x1-x2)^2}+{(y1-y2)^2}}\end{equation}Formula 1
[0067] Among them, (x1, y1) is the coordinate of the head vertex, and (x2, y2) is the coordinate of the crotch point.
[0068] It can be understood that the distance calculation formula is a formula that can calculate the straight-line distance between two key points, and the distance calculation formula in this disclosure is only used as an example.
[0069] Step 203: Using the trunk length repair mapping function, the initial trunk length is corrected to the standard trunk length in a normal sitting posture.
[0070] In some embodiments, due to the position angle of the camera and the different distances of the person to be predicted from the camera, the torso length of the same person may be significantly different in normal sitting posture and other postures. In order to obtain the standard torso length of the person in normal sitting posture, the torso length in other postures needs to be repaired and mapped. The function used in this process is the torso length repair mapping function.
[0071] In the present disclosure, using a torso length repair mapping function to correct an initial torso length to a standard torso length in a normal sitting position includes: determining the sitting posture of the person to be predicted using the key points and coordinates of the key points of the person to be predicted; determining a torso repair mapping function corresponding to the sitting posture; and using the torso repair mapping function corresponding to the sitting posture of the person to be predicted, correcting the initial torso length to a standard torso length in a normal sitting position. The torso repair mapping function is shown in Formula 2:
[0072] B=F(b) Formula 2
[0073] Among them, B is the standard torso length, b is the initial torso length, and F() is the torso repair mapping function.
[0074] Step 204: Multiply the standard torso length by the mapping coefficient to obtain the actual height corresponding to the standard torso length.
[0075] In an optional embodiment of the present disclosure, the standard torso length B of the person to be predicted in a normal sitting position is mapped using the mapping coefficient corresponding to the seat position of the person to be predicted, to obtain the actual height H corresponding to the person to be predicted, as shown in Formula 3:
[0076] H = B * R Formula 3
[0077] Among them, H is the actual height of the person to be predicted, B is the standard torso length of the person to be predicted, and R is the currently determined mapping coefficient.
[0078] Step 205: Calling a scaling factor lookup table, and using the mapping relationship between height and head area in the scaling factor lookup table to determine the scaling factor corresponding to the actual height.
[0079] In some embodiments, a scaling factor P corresponding to the actual height is obtained by querying a scaling factor lookup table of height and head area according to the actual height H of the person object to be predicted.
[0080] In the present disclosure, a scaling factor lookup table is called. Before the step of determining the scaling factor corresponding to the actual height by using the mapping relationship between height and head area in the scaling factor lookup table, the present disclosure includes a process for generating the scaling factor lookup table, including: determining subjects to be trained with different actual heights at preset height intervals within a preset height range, and obtaining the actual heights and actual head areas corresponding to the subjects to be trained with different actual heights; counting statistical values to be trained of the actual head areas of the subjects to be trained at the same actual height within the preset height range, the statistical values to be trained including target statistical values of the actual head areas at the target height; determining the ratio of the target statistical value of the actual head area at each actual height in different actual heights to the statistical value to be trained, and using the ratio as the scaling factor of the actual height corresponding to the statistical value to be trained of the actual head area at each actual height; and tabulating the scaling factors of the actual height corresponding to the statistical value to be trained of the actual head area at each actual height to generate a scaling factor lookup table.
[0081] Among them, list processing means that after obtaining each scaling factor, these scaling factors can be combined into a list to generate a scaling factor query table. If the actual height of the person to be predicted is outside the preset height range, the mapping coefficient of the person to be predicted is determined based on the size relationship between the height boundary value of the current preset height range and the actual height of the person to be predicted. Since the boundary value of the preset height range includes a maximum boundary value and a minimum boundary value, when the actual height of the person to be predicted is greater than the maximum boundary value of the preset height range, the mapping coefficient corresponding to the maximum boundary value is determined as the mapping coefficient of the actual height of the person to be predicted; similarly, when the actual height of the person to be predicted is less than the minimum boundary value of the preset height range, the mapping coefficient corresponding to the minimum boundary value is determined as the mapping coefficient P of the actual height of the person to be predicted.
[0082] In an optional embodiment of the present disclosure, a height and head area scaling coefficient lookup table is established by taking 1000 subjects to be trained, with a preset height range of 0.7m to 1.9m and a preset height interval of 3cm as an example.
[0083] In the height range of 0.7m-1.9m, 1000 subjects to be trained are gathered every 3cm, including men, women, young and old, fat and thin as much as possible. In a fixed camera position, fixed seat, fixed normal sitting posture, the head is straightened for shooting, and the head area of all subjects to be trained is calculated and saved. Next, based on the actual experience and actual situation of the staff, this disclosure selects a target height of 1.7m. For example, in the case of a height of 0.7m, the head areas of 1000 subjects to be trained are collected, and the average of these 1000 head areas is obtained to obtain the head area mean S1 of the subjects to be trained with a height of 0.7m. Similarly, the head area mean S2 of the subjects with a height of 1.7m can be obtained. Then the head area scaling coefficient corresponding to a height of 0.7m is By analogy, we can derive head area scaling factors for heights between 0.7m and 1.9m. If the predicted person's actual height is less than 0.7m, a scaling factor of 0.7m is used. If the predicted person's actual height is greater than 1.9m, a scaling factor of 1.9m is used. At this point, these head area scaling factors are tabulated and combined, completing the lookup table for scaling factors for actual height and head area.
[0084] It should be noted that the scaling factor lookup table has a preset height interval of 3cm and cannot fully cover the height. Therefore, when the actual height is within this 3cm preset height interval, the scaling factor lookup table is queried using the nearest neighbor principle to determine the scaling factor corresponding to the actual height of the person to be predicted (for example, 1.014m corresponds to a scaling factor of 1.0m, and 1.016m corresponds to a scaling factor of 1.03m).
[0085] Step 206: Multiply the scaling factor by the initial head bounding box area of the person object to be predicted to determine the target head bounding box area of the person object to be predicted.
[0086] In some embodiments, since the human head detection network can detect the vertex coordinates of the initial head circumscribed frame of the person object to be predicted, the present disclosure can calculate the circumscribed frame length and circumscribed frame width of the initial head circumscribed frame based on the vertex coordinates of the initial head circumscribed frame, and multiply the circumscribed frame length and circumscribed frame width to obtain the initial head circumscribed frame area S_.
[0087] The present disclosure multiplies the initial head bounding box area by the scaling factor obtained above, and uses Formula 4 to correct the initial head bounding box area of the predicted person object and normalize it to the target head bounding box area size of the target height (e.g., 1.7 meters):
[0088] S = S_ * P Formula 4
[0089] Where S is the target head bounding box area, S_ is the initial head bounding box area, and P is the currently determined scaling factor.
[0090] In summary, the technical solution provided by the present invention uses a lookup table of actual height and head area scaling coefficients, as well as a mapping coefficient between the standard torso length and actual height in a normal sitting position, and combines human head pairing and human seat pairing to correct the initial head external frame area according to the torso length, and unify the initial head external frame area of the predicted human objects of different heights to the target head external frame area corresponding to the target height. Under this same scale standard, the accuracy of the head external frame area is improved, which facilitates more accurate depth estimation for subsequent tasks such as head collision avoidance, thereby improving the recall rate of these tasks.
[0091] Corresponding to the aforementioned method for determining the area of a human head's circumscribed frame, the present invention also provides a device for determining the area of a human head's circumscribed frame. Since the device embodiments of the present invention correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to in the aforementioned method embodiments and will not be further elaborated in this invention.
[0092] Figure 6 This is a schematic diagram of a device for determining the area of a human head circumference frame provided by an embodiment of the present disclosure. Figure 6 Shown, including:
[0093] Matching unit 610 is configured to use the key points of the human body of the predicted person to perform body-seat pairing on the predicted person, determine the seat position of the predicted person, and obtain a mapping coefficient corresponding to the seat position of the predicted person, where the mapping coefficient represents a mapping relationship between the torso length and height of the predicted person at the seat position.
[0094] A height determination unit 620 is configured to adjust the initial torso length of the person to be predicted to a standard torso length in a normal sitting position, and determine the actual height corresponding to the standard torso length using a mapping coefficient, wherein the initial torso length is obtained using key points of the human body;
[0095] A calling unit 630 is configured to call a scaling factor lookup table and determine a scaling factor corresponding to the actual height by using a mapping relationship between height and head area in the scaling factor lookup table;
[0096] The area determination unit 640 is configured to multiply the scaling factor by the initial head bounding box area of the person object to be predicted to determine the target head bounding box area of the person object to be predicted.
[0097] In some embodiments of the present disclosure, the matching unit 610 is used to: 2. obtain the vertex coordinates of the initial human body bounding box of the person to be predicted in the input image before the step of performing body seat pairing on the person to be predicted using the body key points of the person to be predicted; use the vertex coordinates of the initial human body bounding box to crop the input image according to the initial human body bounding box to obtain the body image of the person to be predicted in the input image; perform key point detection on the body image to obtain the body key points of the person to be predicted; determine the hip point among the body key points of the person to be predicted; and perform body seat pairing on the person to be predicted based on the hip point among the body key points.
[0098] In some embodiments of the present disclosure, the height determination unit 620 is used to: determine the head vertex and crotch point among the key points of the human body of the person object to be predicted, and determine the key point straight-line distance between the head vertex and the crotch point, and use the key point straight-line distance as the initial torso length of the person object to be predicted; use the torso length repair mapping function to correct the initial torso length to the standard torso length in a normal sitting posture; multiply the standard torso length by the mapping coefficient to obtain the actual height corresponding to the standard torso length.
[0099] In some embodiments of the present disclosure, the height determination unit 620 is used to: determine the sitting posture of the person object to be predicted using the human body key points and human body key point coordinates; determine the torso repair mapping function corresponding to the sitting posture; and use the torso repair mapping function corresponding to the sitting posture of the person object to be predicted to correct the initial torso length to the standard torso length in a normal sitting posture.
[0100] In some embodiments of the present disclosure, the matching unit 610 is further used to: obtain a preset number of subjects to be trained, and determine the standard trunk length and actual height of the preset number of subjects to be trained in a normal sitting posture at the seat position to be trained, before using the human body key points of the subject to be predicted to perform body seat pairing on the subject to be predicted, determining the seat position of the subject to be predicted, and obtaining the mapping coefficient corresponding to the seat position; determine the ratio of the standard trunk length to the actual height of each subject to be trained in the preset number of subjects to be trained, use the ratio as the mapping coefficient of each subject to be trained, and count the mapping coefficient corresponding to each subject to be trained to obtain the statistical value of the mapping coefficient; average the statistical values of the mapping coefficient to obtain the mapping coefficient at the seat position to be trained.
[0101] In some embodiments of the present disclosure, the calling unit 630 is used to: before calling the scaling coefficient lookup table and determining the scaling coefficient corresponding to the actual height using the mapping relationship between height and head area in the scaling coefficient lookup table, determine subjects to be trained with different actual heights at preset height intervals within a preset height range, and obtain the actual heights and actual head areas corresponding to the subjects to be trained with different actual heights; count the to-be-trained statistical values of the actual head areas of the subjects to be trained at the same actual height within the preset height range, the to-be-trained statistical values including the target statistical values of the actual head areas at the target height; determine the ratio of the target statistical value of the actual head areas at each actual height for different actual heights to the to-be-trained statistical value, and use the ratio as the scaling coefficient of the actual height corresponding to the to-be-trained statistical value of the actual head area at each actual height; list the scaling coefficients of the actual height corresponding to the to-be-trained statistical value of the actual head area at each actual height to generate a scaling coefficient lookup table.
[0102] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.
[0103] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a vehicle.
[0104] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded from a storage unit 708 into a RAM (Random Access Memory) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.
[0106] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0107] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the method for determining the area of a human head circumscribed frame. For example, in some embodiments, the method for determining the area of a human head circumscribed frame can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the aforementioned method for determining the area of the circumscribed frame of a human head in any other appropriate manner (for example, by means of firmware).
[0108] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0112] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0113] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0114] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0115] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0116] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for determining the area of a human head circumference frame, characterized in that: The method comprises: Using the key points of the human body of the person to be predicted, performing body-seat pairing on the person to be predicted, determining the seat position of the person to be predicted, and obtaining a mapping coefficient corresponding to the seat position of the person to be predicted, the mapping coefficient representing a mapping relationship between the torso length and height of the person to be predicted at the seat position; Adjusting the initial torso length of the person to be predicted to a standard torso length in a normal sitting posture, and determining the actual height corresponding to the standard torso length using the mapping coefficient, wherein the initial torso length is obtained using the human body key points; Calling a scaling factor lookup table, and using the mapping relationship between height and head area in the scaling factor lookup table to determine the scaling factor corresponding to the actual height; The scaling factor is multiplied by the initial head circumscribed frame area of the person object to be predicted to determine the target head circumscribed frame area of the person object to be predicted.
2. The method according to claim 1, characterized in that Before the step of performing body seat pairing on the person to be predicted using the key points of the person to be predicted, the method includes: Obtain the vertex coordinates of the initial human body bounding box of the person object to be predicted in the input image; Using the vertex coordinates of the initial human body bounding box, the input image is cropped according to the initial human body bounding box to obtain a human body image of the person object to be predicted in the input image; Performing key point detection on the human body image to obtain the human body key points of the human object to be predicted; The method of performing body seat pairing on the person to be predicted by utilizing the key points of the person to be predicted includes: Determine a crotch point among key points of the human body of the person object to be predicted; According to the hip point in the human body key points, the human body seat pairing is performed on the human object to be predicted.
3. The method according to claim 2, characterized in that The adjusting the initial torso length of the person to be predicted to a standard torso length in a normal sitting posture, and using the mapping coefficient to determine the actual height corresponding to the standard torso length includes: Determine the head vertex and crotch point among the key points of the human body of the to-be-predicted person object, and determine the key point straight-line distance between the head vertex and the crotch point, and use the key point straight-line distance as the initial torso length of the to-be-predicted person object; Using a trunk length repair mapping function, the initial trunk length is corrected to a standard trunk length in a normal sitting posture; The standard trunk length is multiplied by the mapping coefficient to obtain the actual height corresponding to the standard trunk length.
4. The method according to claim 3, characterized in that The method of using the trunk length repair mapping function to correct the initial trunk length to the standard trunk length in a normal sitting posture includes: Determining the sitting posture of the person object to be predicted using the body key points and the body key point coordinates of the person object to be predicted; determining a trunk restoration mapping function corresponding to the sitting posture; The initial torso length is corrected to a standard torso length in a normal sitting posture by using a torso repair mapping function corresponding to the sitting posture of the human object to be predicted.
5. The method according to claim 1, wherein The method for determining the mapping coefficient corresponding to the seat position includes: Obtaining a preset number of subjects to be trained, and determining the standard torso lengths and actual heights of the preset number of subjects to be trained in a normal sitting posture in the seat positions to be trained; Determining the ratio of the standard trunk length to the actual height of each of the preset number of subjects to be trained, using the ratio as a mapping coefficient of each subject to be trained, and statistically calculating the mapping coefficient corresponding to each subject to be trained to obtain a statistical value of the mapping coefficient; The statistical values of the mapping coefficients are averaged to determine the mapping coefficients for the seat position to be trained.
6. The method according to claim 1, wherein The process of generating the scaling factor lookup table includes: Within a preset height range, determine subjects to be trained with different actual heights at preset height intervals, and obtain the actual heights and actual head areas corresponding to the subjects to be trained with different actual heights; Counting the to-be-trained statistical values of the actual head area of the to-be-trained subject at the same actual height within the preset height range, the to-be-trained statistical values including the target statistical value of the actual head area at the target height; Determining the ratio of the target statistical value of the actual head area for each actual height to the statistical value to be trained, and using the ratio as a scaling factor for the actual height corresponding to the statistical value to be trained for the actual head area for each actual height; The scaling coefficients of the actual heights corresponding to the to-be-trained statistical values of the actual head areas under each actual height are tabulated to generate the scaling coefficient lookup table.
7. A device for determining the area of a human head external frame, characterized in that: include: a matching unit, configured to perform body-seat pairing on the person to be predicted using the key points of the person to be predicted, determine the seat position of the person to be predicted, and obtain a mapping coefficient corresponding to the seat position of the person to be predicted, the mapping coefficient representing a mapping relationship between the torso length and height of the person to be predicted at the seat position; a height determination unit, configured to adjust an initial torso length of the human subject to be predicted to a standard torso length in a normal sitting posture, and determine an actual height corresponding to the standard torso length using the mapping coefficients, wherein the initial torso length is obtained using the human body key points; A calling unit, configured to call a scaling factor lookup table, and determine a scaling factor corresponding to the actual height by using a mapping relationship between height and head area in the scaling factor lookup table; The area determination unit is used to multiply the scaling factor by the initial head circumscribed frame area of the person object to be predicted to determine the target head circumscribed frame area of the person object to be predicted.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A vehicle, characterized in that: It includes the device for determining the area of the external frame of a human head as described in claim 5 or the electronic device as described in claim 6.
Citation Information
Patent Citations
Method and device for detecting vehicle occupancy using passengers keypoint detection
CN111507156A
Image evaluation device, image evaluation method, and computer program
JP2022136278A
Method and device for estimating height and weight of passengers using body part length and face information based on human's status recognition
US10643085B1
Personal protective equipment fitting device and method
US20200050836A1
Systems, devices and methods for vehicle post-crash support
WO2020136658A1