Method and device for determining area of external frame of human head, electronic equipment, medium and vehicle

By obtaining the associated face bounding box and determining the head angle using a face pose estimation algorithm, and combining the mapping function of the seat information to correct the area of ​​the head bounding box, the problem of inaccurate head bounding box area under different viewing angles or postures is solved, thereby improving the accuracy and reliability of the anti-collision head alarm.

CN120689390APending Publication Date: 2025-09-23BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410339634.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the existing technology, the head bounding box area obtained by directly using the human head detection and classification network has a large difference between the head bounding box area under different viewing angles or postures and the actual head area, which affects the accuracy and reliability of tasks such as anti-head collision alarm.

Method used

By obtaining the associated face bounding box of the person to be predicted, the head yaw angle and head pitch angle are determined using the facial key point coordinates and the facial pose estimation algorithm, the mapping function of the seat information is called, and the area of ​​the head bounding box is corrected in combination with the head angle correction coefficient to determine the actual area of ​​the head bounding box.

Benefits of technology

The accuracy of determining the external frame of the human head is improved, thereby improving the accuracy and reliability of subsequent alarm tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689390A_ABST
    Figure CN120689390A_ABST
Patent Text Reader

Abstract

The invention provides a head external frame area determination method and device, electronic equipment, a storage medium and a vehicle. The method comprises the following steps: acquiring an associated face bounding box of a person object to be predicted; determining a head yaw angle and a head pitch angle corresponding to the associated face external frame according to the face key point coordinates in the associated face external frame and a face attitude estimation algorithm; calling a mapping function corresponding to the seat information of the to-be-predicted figure object, and determining a head angle correction coefficient corresponding to a head external frame of the to-be-predicted figure object by using the mapping function; the head angle correction coefficient is utilized to correct the head external frame area, the actual head external frame area is determined, the head external frame area of the to-be-predicted person object is corrected for different seats based on the head yaw angle and the head pitch angle of the to-be-predicted person object, the determination accuracy of the head external frame is improved, and the prediction accuracy of the to-be-predicted person object is improved. And the accuracy and reliability of a subsequent alarm task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, electronic device, storage medium, and vehicle for determining the area of ​​a human head circumference frame. Background Art

[0002] Inside the cabin, tasks like head collision avoidance are extremely sensitive to parameters related to head area, which is the primary information used for depth prediction. Therefore, accurately determining the area of ​​the human head's bounding box becomes a key issue.

[0003] Currently, most related technologies directly use the head bounding box area obtained from a human head detection and classification network as the final result for tasks such as head collision avoidance alarms. However, directly using the head bounding box area obtained from a human head detection and classification network in related technologies may result in significant discrepancies between the head bounding box area and the actual head area under different viewing angles or postures, affecting the accuracy and reliability of tasks such as head collision avoidance alarms. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, storage medium, and vehicle for determining the area of ​​a human head circumference frame.

[0005] According to a first aspect of the present disclosure, a method for determining the area of ​​a head circumscribed frame is provided, the method comprising: obtaining a face circumscribed frame associated with a person object to be predicted; determining a head yaw angle and a head pitch angle corresponding to the associated face circumscribed frame based on the coordinates of facial key points in the associated face circumscribed frame in combination with a face posture estimation algorithm; calling a mapping function corresponding to seat information of the person object to be predicted, and determining a head angle correction coefficient corresponding to the head circumscribed frame of the person object to be predicted by utilizing a mapping relationship among the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function, wherein the seat information is obtained by pairing the person object to be predicted with a human body seat; and correcting the area of ​​the head circumscribed frame of the person object to be predicted by utilizing the head angle correction coefficient to determine the actual area of ​​the head circumscribed frame.

[0006] In some embodiments of the present disclosure, obtaining an associated face bounding box of a person to be predicted includes: obtaining an image to be predicted, and performing human head and face classification detection and human key point detection on the image to be predicted, determining an initial head bounding box and human key points of the person to be predicted in the image to be predicted; determining a key point straight line and a key point straight line distance between the chin point and the head vertex of the human key point in the initial head bounding box; determining a tangent value of a head roll angle corresponding to an intersection straight line intersecting with the initial head bounding box, the intersection straight line passing through the chin point, and the intersection straight line being perpendicular to the key point straight line; determining the intersection straight line distance of the intersection straight line in the initial head bounding box using the tangent value, the vertex coordinates of the initial head bounding box, and the chin point coordinates of the chin point; using the key point straight line as the length of the predicted head bounding box, and the intersection straight line as the width of the predicted head bounding box to obtain a predicted head bounding box; performing head and face pairing on the person to be predicted, and obtaining an associated face bounding box that uniquely matches the predicted head bounding box.

[0007] In some embodiments of the present disclosure, the head angle correction coefficient is used to correct the head external frame area of ​​the predicted person object, and determining the actual head external frame area includes: multiplying the straight-line distance of the key points and the straight-line distance of the intersection points to determine the predicted head external frame area; using the predicted head external frame area as the head external frame area of ​​the predicted person object, and using the head angle correction coefficient to correct the head external frame area to determine the actual head external frame area.

[0008] In some embodiments of the present disclosure, obtaining an associated face bounding box of a person object to be predicted includes: obtaining the head key points of the human body key points and the head key point coordinates corresponding to the head key points, and determining the head key point bounding box corresponding to the head key point coordinates; determining the head intersection-over-union ratio between the head key point bounding box and the initial head bounding box; if the head intersection-over-union ratio is greater than or equal to a preset head intersection-over-union ratio threshold, determining that the initial human body bounding box of the person object to be predicted matches the initial head bounding box, and the initial human body bounding box is obtained by performing human head and face classification detection on the image to be predicted; when the initial human body bounding box of the person object to be predicted matches the initial head bounding box, obtaining an associated face bounding box of the person object to be predicted.

[0009] In some embodiments of the present disclosure, performing head-face pairing on a predicted person object and obtaining an associated face bounding box that uniquely matches the predicted head bounding box includes: determining a face intersection-over-union (IoU) ratio between an initial face bounding box and the predicted head bounding box, the initial face bounding box being obtained by performing human head and face classification detection on the predicted image; and using an initial face bounding box having a face intersection-over-union ratio greater than or equal to a preset face intersection-over-union ratio threshold as an associated face bounding box that uniquely matches the predicted head bounding box.

[0010] In some embodiments of the present disclosure, before calling the mapping function corresponding to the seat information of the person object to be predicted and determining the head angle correction coefficient corresponding to the head external frame of the person object to be predicted by using the mapping relationship between the head yaw angle, the head pitch angle and the head angle correction coefficient in the mapping function, the method includes: determining the standard head external frame area of ​​multiple subjects to be trained under the seat to be trained when the head roll angle, the head yaw angle and the head pitch angle meet the target conditions, the target conditions including the head roll angle being the target head roll angle, the head yaw angle being the target head yaw angle and the head pitch angle being the target head pitch angle; obtaining the initial positions of multiple subjects to be trained in turn according to the head rotation directions of the multiple subjects to be trained The vertex coordinates of the initial head external bounding box, the vertex coordinates of the initial human body external bounding box, the coordinates of the human body key points, and the coordinates of the facial key points are obtained; the vertex coordinates of the initial head external bounding box, the vertex coordinates of the initial human body external bounding box, the coordinates of the human body key points, and the coordinates of the facial key points corresponding to multiple objects to be trained are used, combined with the face posture estimation algorithm, to sequentially obtain the head yaw angle, the head pitch angle, and the predicted head external bounding box area corresponding to multiple objects to be trained in each input image when the head roll angle is the target head roll angle; determine the ratio between the predicted head external bounding box area and the standard head external bounding box area of ​​multiple objects to be trained, obtain the initial head angle correction coefficient, and use the initial head angle correction coefficient to obtain the mapping function corresponding to the seat to be trained.

[0011] In some embodiments of the present disclosure, using the initial head angle correction coefficient to obtain the mapping function corresponding to the seat to be trained includes: combining the initial head angle correction coefficient, the head yaw angle, and the head pitch angle in a preset order, obtaining a data group for each of the multiple objects to be trained, and storing the data group in a database; summarizing the data group of each object to be trained, determining similar data groups with the same head yaw angle and head pitch angle in the data group of each object to be trained, and averaging the initial head angle correction coefficients in the similar data groups to obtain the head angle correction coefficient; using the head yaw angle, the head pitch angle, and the head angle correction coefficient as inputs of a three-dimensional function, fitting to obtain the mapping function corresponding to the seat to be trained, the mapping function including the mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient under the seat to be trained.

[0012] In some embodiments of the present disclosure, the head angle correction coefficient is used to correct the head circumference frame area of ​​the person object to be predicted, and obtaining the actual head circumference frame area includes: multiplying the head angle correction coefficient by the head circumference frame area of ​​the person object to be predicted to obtain the actual head circumference frame area of ​​the person object to be predicted.

[0013] In some embodiments of the present disclosure, after using the head angle correction coefficient to correct the head circumference frame area of ​​the predicted person object and determining the actual head circumference frame area, the method includes: determining the distance between the actual head circumference frame area and the camera; if the distance is less than or equal to the preset safety distance, sending an alarm message to the predicted person object.

[0014] According to a second aspect of the present disclosure, a device for determining the area of ​​a human head circumscribed frame is provided, comprising:

[0015] A pairing unit, used to obtain the associated face bounding box of the person object to be predicted;

[0016] A first determining unit is configured to determine the head yaw angle and head pitch angle corresponding to the associated face circumscribed frame based on the coordinates of the facial key points in the associated face circumscribed frame in combination with a face pose estimation algorithm;

[0017] a second determining unit, configured to call a mapping function corresponding to seat information of a person to be predicted, and determine a head angle correction coefficient corresponding to a head bounding box of the person to be predicted using a mapping relationship among a head yaw angle, a head pitch angle, and a head angle correction coefficient in the mapping function, wherein the seat information is obtained by pairing the person to be predicted with a human body;

[0018] The correction unit is used to correct the head circumference frame area of ​​the predicted person object using the head angle correction coefficient to determine the actual head circumference frame area.

[0019] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0020] at least one processor; and

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.

[0023] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.

[0024] According to a fifth aspect of the present disclosure, a vehicle is provided, comprising the device as described in the second aspect or the electronic device as described in the third aspect.

[0025] The present invention provides a method, device, electronic device, storage medium and vehicle for determining the area of ​​a head external bounding box, which obtains a face external bounding box associated with a person to be predicted; determines the head yaw angle and head pitch angle corresponding to the associated face external bounding box based on the coordinates of facial key points in the associated face external bounding box and in combination with a face posture estimation algorithm; calls a mapping function corresponding to the seat information of the person to be predicted, and uses the mapping relationship between the head yaw angle, the head pitch angle and the head angle correction coefficient in the mapping function to determine the head angle correction coefficient corresponding to the head external bounding box of the person to be predicted, wherein the seat information is obtained by pairing the person to be predicted with a human body; uses the head angle correction coefficient to correct the area of ​​the head external bounding box of the person to be predicted to determine the actual area of ​​the head external bounding box, and realizes that for different seats, the area of ​​the head external bounding box of the person to be predicted is corrected based on the head yaw angle and the head pitch angle of the person to be predicted, thereby improving the accuracy of the determination of the head external bounding box, and thereby improving the accuracy and reliability of subsequent alarm tasks.

[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0028] Figure 1 A schematic flow chart of a method for determining the area of ​​a human head circumference frame provided in an embodiment of the present disclosure;

[0029] Figure 2 A flowchart of another method for determining the area of ​​a human head circumference frame provided by an embodiment of the present disclosure;

[0030] Figure 3 A schematic diagram of obtaining the vertex coordinates of an initial human body circumference frame, an initial head circumference frame, and an initial face circumference frame provided by an embodiment of the present disclosure;

[0031] Figure 4 A schematic diagram of obtaining coordinates of key points of a human body provided in an embodiment of the present disclosure;

[0032] Figure 5 A schematic diagram of obtaining the coordinates of key points of a face provided in an embodiment of the present disclosure;

[0033] Figure 6 A schematic diagram of the distribution of key points of a human body provided in an embodiment of the present disclosure;

[0034] Figure 7A schematic diagram of the distribution of key points of a face provided in an embodiment of the present disclosure;

[0035] Figure 8 A schematic diagram of a head rotation angle provided by an embodiment of the present disclosure;

[0036] Figure 9 A schematic diagram of a face angle calibration plate provided in an embodiment of the present disclosure;

[0037] Figure 10 A schematic structural diagram of a device for determining the area of ​​a human head circumference frame provided in an embodiment of the present disclosure;

[0038] Figure 11 A schematic block diagram of an exemplary electronic device 1100 is provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0040] In order to solve the problems in the related art, the present disclosure regresses the head external frame area of ​​the task object to be predicted to the state where the head roll angle is 0, determines the predicted head external frame area in this state, and then regresses all predicted head external frame areas with different head yaw angles and head pitch angles to the standard head external frame area when the head yaw angle and head pitch angle are both 0, thereby improving the accuracy of the actual head external frame area finally determined, facilitating more accurate depth estimation of alarm tasks (such as anti-collision head alarm tasks), and improving the reliability and accuracy of alarm tasks.

[0041] The following describes a method, device, electronic device, storage medium, and vehicle for determining the area of ​​a human head circumference frame according to embodiments of the present disclosure with reference to the accompanying drawings.

[0042] Figure 1 This is a flow chart of a method for determining the area of ​​a human head's external frame provided in an embodiment of the present disclosure. This method can be applied to scenarios where a human head is recognized and located in an image or video, and can be specifically applied to scenarios such as smart cockpits or security monitoring in automobiles, which are not limited in the embodiments of the present disclosure. Figure 1 As shown, the method includes:

[0043] Step 101: Obtain the associated facial bounding box of the person object to be predicted.

[0044] In some embodiments, the person to be predicted refers to any person to be detected in an input image captured by a capture device. In the present disclosure, the person to be predicted can be a person facing a surveillance camera and identifiable, or a user using a device requiring head tracking or collision avoidance (e.g., a virtual reality headset).

[0045] A head bounding box is a geometric model used to represent the external contours of a head and to estimate and identify its size and position. The head bounding box can be a simplified head shape composed of a series of line segments that describe the basic shape and size of the head. In this disclosure, the head bounding box takes the form of an outer rectangular box surrounding the head.

[0046] The associated face bounding box in the present disclosure refers to the initial face bounding box that uniquely matches the predicted head bounding box, wherein the predicted head bounding box is obtained by adjusting the initial head bounding box using the head roll angle of the person object to be predicted, and the initial head bounding box is obtained by directly detecting the human head and face classification detection network in the present disclosure.

[0047] Since there can only be one face on a head, there may be a situation where the predicted head bounding box has multiple initial face bounding boxes when the human head and face classification detection network is detecting. Therefore, the present disclosure requires head and face matching to obtain the initial face bounding box that uniquely matches the predicted head bounding box, and use the initial face bounding box as the associated face bounding box.

[0048] Step 102: Determine the head yaw angle and head pitch angle corresponding to the associated face bounding box based on the coordinates of the facial key points in the associated face bounding box and in combination with a face pose estimation algorithm.

[0049] In some embodiments, facial pose estimation is performed based on the coordinates of facial key points in the associated face bounding box detected by the detection network, thereby determining the head yaw angle and head pitch angle of the person to be predicted. The head yaw angle represents the rotation of the head in the horizontal plane, and the head pitch angle represents the rotation of the head in the vertical direction.

[0050] Face pose estimation is an image processing technique used to estimate the pose and angle of a face relative to a camera or a reference coordinate system. It infers the face pose and head rotation angles, primarily including head yaw, pitch, and roll, by analyzing the coordinates of facial landmarks within a bounding box.

[0051] Step 103: Call the mapping function corresponding to the seat information of the person to be predicted, and use the mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function to determine the head angle correction coefficient corresponding to the head bounding box of the person to be predicted.

[0052] In the present disclosure, the seat information is obtained by pairing the human body and seat of the person object to be predicted.

[0053] In some embodiments, the seat information of the person to be predicted refers to the current seat position of the person to be predicted within the cabin. The seat information of the person to be predicted can be determined using the coordinates of key points of the person to be predicted. After determining the seat information of the person to be predicted, a mapping function for that seat information can be determined based on the seat information. This means that the mapping functions corresponding to different seat information may be the same or different, which is not a limitation in the presently disclosed embodiments.

[0054] After determining the mapping function, the head angle correction coefficient corresponding to the head bounding box of the current person object to be predicted can be determined based on the mapping relationship between the head yaw angle, the head pitch angle and the body angle correction coefficient in the mapping function.

[0055] The head circumscribed frame of the person object to be predicted may be a predicted head circumscribed frame obtained by adjusting the head roll angle of the person object to be predicted.

[0056] It should be noted that the mapping function in the present disclosure is pre-trained for different seat information and is not limited in the embodiments of the present disclosure.

[0057] Step 104: Correct the head bounding box area of ​​the person to be predicted using the head angle correction coefficient to obtain the actual head bounding box area.

[0058] In some embodiments, after determining the head angle correction coefficient of the head external frame of the person object to be predicted, the head angle correction coefficient is multiplied by the head external frame area of ​​the person object to be predicted, and the product is used as the actual head external frame area of ​​the person object to be predicted to avoid the problem of inaccurate head external frame area due to different head yaw angles and head pitch angles.

[0059] The actual head bounding box area refers to the area of ​​the head bounding box of the person being predicted in a standard sitting posture, that is, the area of ​​the head bounding box when the person is looking directly at the camera. This means that the head pitch, yaw, and roll angles are all zero. In this case, the error between the actual head bounding box area and the actual head is minimized.

[0060] In summary, the technical solution provided by the present invention obtains the associated face bounding box of the person object to be predicted; determines the head yaw angle and head pitch angle corresponding to the associated face bounding box based on the coordinates of the facial key points in the associated face bounding box and the face posture estimation algorithm; calls the mapping function corresponding to the seat information of the person object to be predicted, and uses the mapping relationship between the head yaw angle, the head pitch angle and the head angle correction coefficient in the mapping function to determine the head angle correction coefficient corresponding to the head bounding box of the person object to be predicted, where the seat information is obtained by pairing the human body and seat of the person object to be predicted; uses the head angle correction coefficient to correct the area of ​​the head bounding box of the person object to be predicted, determines the actual area of ​​the head bounding box, and realizes that for different seats, the area of ​​the head bounding box of the person object to be predicted is corrected based on the head yaw angle and the head pitch angle of the person object to be predicted, thereby improving the accuracy of the determination of the head bounding box, and thereby improving the accuracy and reliability of subsequent alarm tasks.

[0061] Figure 2 This is a flow chart of another method for determining the area of ​​a human head circumference frame provided in an embodiment of the present disclosure. The method can be executed by an electronic device, specifically, by an electronic device such as a PC or a server. Figure 2 based on Figure 1 In the embodiment shown, step 101 and step 104 are further defined. Figure 2 In the embodiment shown, step 101 includes step 201 and step 202, and step 104 includes step 205. Figure 2 As shown, the method may include:

[0062] Step 201: Acquire an image to be predicted, perform human head and face classification detection and human key point detection on the image to be predicted, and determine the head external bounding box of the person to be predicted in the image to be predicted.

[0063] In some embodiments, after obtaining the image to be predicted, human head and face classification detection can be performed on the image to be predicted, and the initial human body bounding box, initial head bounding box and initial face bounding box of the person object to be predicted in the image to be predicted can be obtained; the vertex coordinates of the initial human body bounding box are used to crop the image to be predicted according to the initial human body bounding box to obtain the human body image in the image to be predicted; human body key point detection is performed on the human body image to obtain the human body key points in the human body image and the human body key point coordinates corresponding to the human body key points; the vertex coordinates of the initial face bounding box are used to crop the image to be predicted according to the initial face bounding box to obtain the face image in the image to be predicted; key point detection is performed on the face image to obtain the face key points in the face image and the face key point coordinates of the face key points.

[0064] The initial head circumscribed frame and the key points of the human body obtained above can be used to determine the predicted head circumscribed frame of the person object to be predicted, and the predicted head circumscribed frame is used as the head circumscribed frame of the person object to be predicted.

[0065] In an optional embodiment of the present disclosure, the human head and face classification detection in the present disclosure can be implemented by a human head and face detection classification network, and the human key point detection can be implemented by a human key point detection network. The input image to be predicted is processed by the human head and face detection classification network to obtain the vertex coordinates of the initial human body external bounding box, the vertex coordinates of the initial head external bounding box, and the vertex coordinates of the initial face external bounding box of the person to be predicted in the image to be predicted. The image to be predicted is obtained by an image acquisition device, such as a camera.

[0066] Among them, the vertex coordinates of the initial human body external bounding box, the vertex coordinates of the initial human body external bounding box, and the vertex coordinates of the initial face external bounding box can be the vertex coordinates of the upper left corner and the vertex coordinates of the lower right corner of the external bounding box of the human body, head, and face. A rectangular frame can be formed by the vertex coordinates of the upper left corner and the vertex coordinates of the lower right corner.

[0067] The human head and face detection and classification network is a deep learning network built for detection and classification scenarios, usually composed of a backbone network (backbone), a shoulder network (neck), and a head network (head). It is used to detect whether there is a human body, head, or face and return the vertex coordinates of the initial human body external frame, the vertex coordinates of the initial head external frame, and the vertex coordinates of the initial face external frame. In the present disclosure, before processing the image to be predicted through the human head and face detection and classification network, it is necessary to train a human head and face detection and classification network through a large amount of training data. These training data can be various image or video data sets, which contain different human bodies, heads, and face postures, different backgrounds and lighting conditions, etc. Figure 3The diagram shown is a schematic diagram of obtaining the vertex coordinates of the initial human body external bounding box, the vertex coordinates of the initial head external bounding box, and the vertex coordinates of the initial face external bounding box. Before the image to be predicted is input into the human head and face detection and classification network, the image to be predicted (original image) needs to be preprocessed. The preprocessing may include operations such as scaling, cropping, and normalization to facilitate the human head and face detection and classification network to better extract features. Then, the human head and face detection and classification network is used to extract features from the preprocessed image to be predicted through the backbone network (backbone), shoulder network (neck), and head network (head). The output of the head network is parsed and processed through post-processing, such as non-maximum suppression (NMS) operations, to eliminate overlapping detection frames, and return the vertex coordinates of the initial human body external bounding box (human body frame coordinates), the vertex coordinates of the initial head external bounding box (head frame coordinates), and the vertex coordinates of the initial face external bounding box (face frame coordinates).

[0068] Since a rectangular frame can be formed by the vertex coordinates of the initial human body circumscribed frame and the vertex coordinates of the initial face circumscribed frame, the present disclosure can use a preset cropping function to input the predicted image to be cropped and the obtained vertex coordinates of the initial human body circumscribed frame and the vertex coordinates of the initial face circumscribed frame into this cropping function, thereby directly obtaining the cropped human body image and face image.

[0069] Specifically, the present disclosure can determine the image area to be cropped by using the vertex coordinates of the initial human body circumference frame and the vertex coordinates of the initial face circumference frame through the cropping function set in OpenCV, and read the predicted image to be cropped to perform image cropping.

[0070] It should be noted that the preset cropping function in the present disclosure can be set according to the actual situation of the image to be predicted and the actual working experience of the staff, and is not limited in the embodiments of the present disclosure.

[0071] Afterwards, the present disclosure can process the cropped human body image and facial image through the human body key point detection network and the facial key point detection network respectively, thereby obtaining multiple human body key points and multiple human body key point coordinates in the human body image and multiple facial key points and multiple facial key point coordinates in the facial image.

[0072] The human key point detection network and the face key point detection network are both deep learning networks built for detection scenarios, which can be composed of a backbone network (backbone), a shoulder network (neck), and a head network (head), and are used to detect whether human key points and face key points exist and return the coordinates of the human key points and the coordinates of the face key points. In the present disclosure, before processing human images and face images through the human key point detection network and the face key point detection network, it is necessary to train a human key point detection network and a face key point detection network through a large amount of training data. These training data can be various human images and face images, including labeled and unlabeled human images and face images of different scales, lighting conditions and backgrounds. As Figure 4 A schematic diagram of obtaining the coordinates of key points of the human body is shown as follows Figure 5 The schematic diagram of obtaining the coordinates of facial key points shown in the figure shows that before the human key point detection network and the face key point detection network detect the human image and the face image, it is necessary to crop the predicted image (original image) according to the size of the initial human external frame and the initial face external frame obtained above to obtain the human image and the face image, and preprocess the human image and the face image. The preprocessing may include operations such as scaling, cropping, and normalization, so that the human key point detection network and the face key point detection network can better extract features. The preprocessed human image is then sequentially subjected to feature extraction by the backbone network (backbone), the shoulder network (neck), and the head network (head). The output of the head network is then parsed and processed by post-processing, such as non-maximum suppression (NMS) and other operations, to eliminate overlapping key points, thereby determining the human key points and the face key points, and returning the human key point coordinates corresponding to the human key points and the face key point coordinates corresponding to the face key points. Specifically, it can be as follows Figure 6 A schematic diagram of the distribution of key points of the human body as shown in Figure 7 A schematic diagram of the distribution of facial key points is shown.

[0073] In some embodiments, the above method can be used to determine the initial head bounding box and human body key points of the person to be predicted in the image to be predicted; and the initial head bounding box and human body key points are used to determine the predicted head bounding box. Specifically, the method includes: determining the key point straight line and the key point straight line distance between the chin point and the head vertex of the human body key point in the initial head bounding box; determining the tangent value of the head roll angle corresponding to the intersection straight line intersecting with the initial head bounding box, where the intersection straight line passes through the chin point and is perpendicular to the key point straight line; determining the intersection straight line distance of the intersection straight line in the initial head bounding box using the tangent value, the vertex coordinates of the initial head bounding box, and the chin point coordinates of the chin point; using the key point straight line as the length of the predicted head bounding box and the intersection straight line as the width of the predicted head bounding box to obtain the predicted head bounding box.

[0074] After the predicted head bounding box is obtained, the predicted head bounding box of the person object to be predicted may be used as the head bounding box of the person object to be predicted for subsequent pairing processing.

[0075] Step 202: performing head and face matching on the person to be predicted, and obtaining an associated face bounding box that uniquely matches the head bounding box of the person to be predicted.

[0076] In some embodiments, since the head bounding box of the person object to be predicted may be the predicted head bounding box of the person object to be predicted, the present disclosure needs to obtain an associated face bounding box that uniquely matches the predicted head bounding box.

[0077] Among them, the head and face pairing of the predicted person object and obtaining the associated face bounding box that uniquely matches the predicted head bounding box specifically include: determining the face intersection-over-union (IoU) between the initial face bounding box and the predicted head bounding box, the initial face bounding box is obtained by performing human head and face classification detection on the predicted image; and taking the initial face bounding box with a face intersection-over-union (IoU) greater than or equal to a preset face intersection-over-union (IoU) threshold as the associated face bounding box that uniquely matches the predicted head bounding box.

[0078] Intersection over Union (IOU) is a metric that measures the similarity between two rectangular boxes (commonly used for object detection). The preset IOU threshold for human heads is pre-set based on actual manual experience and is not limited in the disclosed embodiments.

[0079] In an optional embodiment of the present disclosure, since there may be a situation where a predicted head bounding box corresponds to multiple initial face bounding boxes, the present disclosure can screen multiple initial face bounding boxes based on the facial intersection and comparison between the initial face bounding box and the predicted head bounding box to find the associated face bounding box that uniquely corresponds to the current predicted head bounding box.

[0080] It should be noted that, since it is necessary to ensure that the head bounding box of the person to be predicted matches the body bounding box when obtaining the associated face bounding box of the person to be predicted, the present disclosure includes: obtaining the head key points of the body key points and the head key point coordinates corresponding to the head key points, and determining the head key point bounding box corresponding to the head key point coordinates; determining the head intersection-over-union ratio between the head key point bounding box and the initial head bounding box; if the head intersection-over-union ratio is greater than or equal to the preset head intersection-over-union ratio threshold, determining that the initial body bounding box of the person to be predicted matches the initial head bounding box, and the initial body bounding box is obtained by performing human head and face classification detection on the image to be predicted; when the initial body bounding box of the person to be predicted matches the initial head bounding box, obtaining the associated face bounding box of the person to be predicted.

[0081] Step 203: Determine the head yaw angle and head pitch angle corresponding to the associated face bounding box based on the coordinates of the facial key points in the associated face bounding box and in combination with a face pose estimation algorithm.

[0082] In some embodiments, as Figure 8 The schematic diagram of a head rotation angle shown in FIG. 1 is a diagram of a head rotation angle, where the head pitch angle is the rotation angle about the x-axis, corresponding to Figure 8 Pitch in; the head yaw angle is the rotation angle about the y-axis, corresponding to Figure 8 Yaw in; the head roll angle is the rotation angle about the z axis, corresponding to Figure 8 Roll in.

[0083] In the present disclosure, the facial pose estimation algorithm is combined with the feature extraction of the coordinates of the key points of the face to obtain the head yaw angle (yaw angle) and head pitch angle (pitch angle) corresponding to the face external frame.

[0084] Step 204: Call the mapping function corresponding to the seat information of the person to be predicted, and use the mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function to determine the head angle correction coefficient corresponding to the head bounding box of the person to be predicted.

[0085] In the present disclosure, the seat information is obtained by pairing the human body and seat of the person object to be predicted.

[0086] In some embodiments, before calling the mapping function corresponding to the seat information of the person object to be predicted, the present disclosure needs to first determine the seat information of the person object to be predicted. Specifically, this can be done by determining the hip point among the key points of the human body of the first person object to be predicted; determining the seat area corresponding to the hip point and the seat information corresponding to the seat area.

[0087] After obtaining the seat information of the person to be predicted, the mapping function F of the seat information can be called to obtain the head angle correction coefficient R=F(yaw, pitch) of the head bounding box of the current person to be predicted.

[0088] In addition, it should be noted that the mapping function is pre-trained, and the specific training process includes: determining the standard head external frame area of ​​multiple subjects to be trained under the seat to be trained with respect to the head roll angle, head yaw angle and head pitch angle under the target conditions, and the target conditions include the head roll angle being the target head roll angle, the head yaw angle being the target head yaw angle and the head pitch angle being the target head pitch angle; according to the head rotation directions of multiple subjects to be trained, sequentially obtaining the vertex coordinates of the initial head external frame, the vertex coordinates of the initial human external frame, the coordinates of the human body key points, and the coordinates of the human face key points. The method comprises the following steps: using the vertex coordinates of the initial head bounding box, the vertex coordinates of the initial human body bounding box, the coordinates of the key points of the human body, and the coordinates of the key points of the human face corresponding to the multiple subjects to be trained, and combining the face pose estimation algorithm to sequentially obtain the head yaw angle, head pitch angle, and predicted head bounding box area corresponding to the multiple subjects to be trained in each input image when the head roll angle is the target head roll angle; determining the ratio between the predicted head bounding box area and the standard head bounding box area of ​​the multiple subjects to be trained to obtain the initial head angle correction coefficient, and using the initial head angle correction coefficient, obtaining the mapping function corresponding to the seat to be trained. In the present disclosure, the angle values ​​corresponding to the target head roll angle, target head yaw angle, and target head pitch angle are all 0.

[0089] Among them, using the initial head angle correction coefficient to obtain the mapping function corresponding to the seat to be trained includes: combining the initial head angle correction coefficient, the head yaw angle and the head pitch angle in a preset order, obtaining a data group for each of the multiple objects to be trained, and storing the data group in a database; summarizing the data group of each object to be trained, determining similar data groups with the same head yaw angle and head pitch angle in the data group of each object to be trained, and averaging the initial head angle correction coefficients in the similar data groups to obtain the head angle correction coefficient; using the head yaw angle, the head pitch angle and the head angle correction coefficient as inputs of a three-dimensional function, fitting to obtain the mapping function corresponding to the seat to be trained, the mapping function including the mapping relationship between the head yaw angle, the head pitch angle and the head angle correction coefficient under the seat to be trained.

[0090] In an optional embodiment of the present disclosure, due to the particularity of the cabin space, its complex internal environment, and the tilt of the camera, it is necessary to recalibrate the head yaw angle (yaw angle) and head pitch angle (pitch angle) to establish a mapping relationship between them and the predicted head bounding box area and the standard head bounding box area:

[0091] For different seats, the mapping relationship is also different, so the present disclosure needs to process each seat separately. Here, one of the seats, that is, the seat to be trained, is taken as an example for explanation.

[0092] First, the present invention can collect calibration data, specifically using a face angle calibration plate to collect calibration data, such as Figure 9 The schematic diagram of the face angle calibration plate shown is shown, where l represents the horizontal line and r represents the vertical line.

[0093] In this disclosure, a certain number of subjects to be trained (e.g. 100 people) can be recruited, including those of different heights, weights, and sizes, males, females, young and old, and data collection is performed on a real vehicle. During the collection process, each subject to be trained maintains a head roll angle of 0. Initially, the subject to be trained needs to look straight ahead. Figure 9 The intersection of l4 and r9 is the center of the entire face angle calibration plate. This corresponds to a state where both the head yaw angle and the head pitch angle are 0. Combined with the premise that the head roll angle is 0, the head is in a standard posture at this time, and the head external frame area Ss under the camera is the standard head external frame area. Next, the line of sight of the subject to be trained moves slowly along the head rotation direction of l1->l7, r1->r17, first completing all the red lines and then all the blue lines. During this process, the human head face detection and classification network and the facial key point detection network can be called, and the face posture estimation algorithm can be used to obtain the head yaw angle, head pitch angle and predicted head external frame area Sp under each frame of the captured image, and calculate the head angle correction coefficient at this time. And according to the preset order of the head yaw angle, the head pitch angle and the head angle correction coefficient, a data group (yaw, pitch, R) is obtained and the data group (yaw, pitch, R) is saved in the database.

[0094] After all data is collected, each subject to be trained will have many data sets (yaw, pitch, R). These data sets are aggregated, and for data sets with the same head yaw and head pitch angles (yaw, pitch), all the head angle correction coefficients R are averaged. After the above operation, each (yaw, pitch) corresponds to only one head angle correction coefficient R. Therefore, MATLAB can be used to fit a three-dimensional function, where the head angle correction coefficient R can be mapped to the z-axis, and the head yaw and head pitch angles can be mapped to the x-axis and y-axis respectively. This successfully establishes the mapping relationship between the head yaw angle, head pitch angle, and head angle correction coefficient R, denoted as R = F(yaw, pitch).

[0095] Step 205: Multiply the head angle correction coefficient by the head bounding box area of ​​the person to be predicted to determine the actual head bounding box area of ​​the person to be predicted.

[0096] In some embodiments, after obtaining the predicted head bounding box, the straight-line distance between the key points and the straight-line distance between the intersection points can be multiplied to determine the predicted head bounding box area; the predicted head bounding box area is used as the head bounding box area of ​​the person object to be predicted, and the head angle correction coefficient is used to correct the head bounding box area to determine the actual head bounding box area.

[0097] Among them, before determining the predicted head circumscribed frame area, the present invention can first obtain the initial head circumscribed frame area of ​​the person object to be predicted, that is, determine the width and height of the initial head circumscribed frame according to the vertex coordinates of the initial head circumscribed frame, and multiply the width and height to determine the initial head circumscribed frame area.

[0098] In an optional embodiment of the present disclosure, since the width W and height H of the initial head circumference frame can be determined according to the vertex coordinates of the initial head circumference frame, the present disclosure can calculate the area Sp_ of the initial head circumference frame according to Formula 1.

[0099] Sp_ = W * H Formula 1

[0100] Wherein, Sp_ represents the calculated area of ​​the initial head bounding box, W represents the width of the initial head bounding box, and H represents the height of the initial head bounding box.

[0101] After obtaining the initial head bounding box area, the predicted head bounding box area can be calculated using the obtained key point straight-line distance and intersection straight-line distance, and the predicted head bounding box area is recorded as Sp.

[0102] The present disclosure can use the head angle correction coefficient to perform a secondary angle correction on the predicted head circumference frame area Sp, wherein the first angle correction refers to the process of adjusting the initial head circumference frame area to the predicted head circumference frame area based on the head roll angle.

[0103] Specifically, the actual head bounding box area of ​​the current person to be predicted can be obtained according to Formula 2:

[0104] S = Sp * R Formula 2

[0105] Where S represents the actual head bounding box area, Sp represents the predicted head bounding box area, and R represents the head angle correction coefficient.

[0106] The actual head circumscribed frame area S calculated at this time has been corrected to roll=0 and yaw=pitch=0, and is closest to the actual head area.

[0107] Furthermore, after obtaining the actual head bounding box area that most closely matches the actual head area, the present disclosure can also use the actual head bounding box area to execute a corresponding alarm task. This includes determining the distance between the actual head bounding box area and the camera; if the distance is less than or equal to a preset safety distance, sending an alarm message to the person being predicted. The alarm message can be delivered in the form of a voice broadcast, a pop-up window display, or other forms, which are not limited in the present embodiment.

[0108] In summary, the technical solution provided by the present disclosure can calibrate the head yaw angle and head pitch angle in the interior space of the cabin, establish a mapping relationship between the head yaw angle, head pitch angle and the head angle correction coefficient R, thereby determining the actual head external frame area of ​​the predicted person object when the head roll angle, head yaw angle and head pitch angle are all 0, improving the accuracy of the actual head external frame area, and facilitating more accurate depth estimation of subsequent anti-collision head warning tasks, thereby improving the recall rate of these tasks.

[0109] Corresponding to the aforementioned area determination method, the present invention also provides a device for determining the area of ​​a human head circumference frame. Since the device embodiment of the present invention corresponds to the aforementioned method embodiment, details not disclosed in the device embodiment can be referred to the aforementioned method embodiment and will not be further described in this invention.

[0110] Figure 10 This is a schematic diagram of a device for determining the area of ​​a human head circumference frame provided by an embodiment of the present disclosure. Figure 10 As shown, the device includes:

[0111] A pairing unit 1010 is configured to obtain a facial bounding box associated with the person object to be predicted;

[0112] The first determining unit 1020 is configured to determine the head yaw angle and head pitch angle corresponding to the associated face circumscribed frame based on the coordinates of the facial key points in the associated face circumscribed frame in combination with a face pose estimation algorithm;

[0113] A second determining unit 1030 is configured to call a mapping function corresponding to the seat information of the person to be predicted, and determine a head angle correction coefficient corresponding to the head bounding box of the person to be predicted using a mapping relationship among the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function, wherein the seat information is obtained by pairing the person to be predicted with the person's body;

[0114] The correction unit 1040 is configured to correct the head circumference area of ​​the person object to be predicted by using the head angle correction coefficient to determine the actual head circumference area.

[0115] In some embodiments of the present disclosure, the pairing unit 1010 is used to: obtain an image to be predicted, and perform human head and face classification detection and human key point detection on the image to be predicted, determine the initial head external bounding box and human key points of the person to be predicted in the image to be predicted; determine the key point straight line and the key point straight line distance between the chin point and the head vertex of the human key point in the initial head external bounding box; determine the tangent value of the head roll angle corresponding to the intersection straight line intersecting with the initial head external bounding box, the intersection straight line passes through the chin point, and the intersection straight line is perpendicular to the key point straight line; use the tangent value, the vertex coordinates of the initial head external bounding box, and the chin point coordinates of the chin point to determine the intersection straight line distance of the intersection straight line in the initial head external bounding box; use the key point straight line as the length of the predicted head external bounding box, and the intersection straight line as the width of the predicted head external bounding box to obtain the predicted head external bounding box; perform head and face pairing on the predicted person object, and obtain an associated face external bounding box that uniquely matches the predicted head external bounding box.

[0116] In some embodiments of the present disclosure, the correction unit 1040 is used to: multiply the straight-line distance of the key points and the straight-line distance of the intersection points to determine the predicted head circumference frame area; use the predicted head circumference frame area as the head circumference frame area of ​​the person object to be predicted, and use the head angle correction coefficient to correct the head circumference frame area to determine the actual head circumference frame area.

[0117] In some embodiments of the present disclosure, the pairing unit 1010 is used to: obtain the head key points of the human body key points and the head key point coordinates corresponding to the head key points, and determine the head key point bounding box corresponding to the head key point coordinates; determine the head intersection-over-union ratio between the head key point bounding box and the initial head bounding box; if the head intersection-over-union ratio is greater than or equal to a preset head intersection-over-union ratio threshold, determine that the initial human body bounding box of the person object to be predicted matches the initial head bounding box, and the initial human body bounding box is obtained by performing human head and face classification detection on the image to be predicted; when the initial human body bounding box of the person object to be predicted matches the initial head bounding box, obtain the associated face bounding box of the person object to be predicted.

[0118] In some embodiments of the present disclosure, the pairing unit 1010 is used to: determine a face intersection-over-union (FIU) between an initial face bounding box and a predicted head bounding box, where the initial face bounding box is obtained by performing human head and face classification detection on the predicted image; and use the initial face bounding box whose face intersection-over-union (FIU) is greater than or equal to a preset face intersection-over-union (FIU) threshold as an associated face bounding box that uniquely matches the predicted head bounding box.

[0119] In some embodiments of the present disclosure, the device 1000 further includes: a training unit for determining, before the step of calling the mapping function corresponding to the seat information of the person object to be predicted and determining the head angle correction coefficient corresponding to the head external bounding box of the person object to be predicted by using the mapping relationship between the head yaw angle, the head pitch angle and the head angle correction coefficient in the mapping function, a standard head external bounding box area of ​​multiple persons to be trained under the seat to be trained when the head roll angle, the head yaw angle and the head pitch angle meet the target conditions, the target conditions including the head roll angle being the target head roll angle, the head yaw angle being the target head yaw angle and the head pitch angle being the target head pitch angle; obtaining, in turn, multiple persons to be trained according to the head rotation directions of the multiple persons to be trained The vertex coordinates of the initial head bounding box, the vertex coordinates of the initial human body bounding box, the coordinates of the human body key points, and the coordinates of the facial key points of the training object; using the vertex coordinates of the initial head bounding box, the vertex coordinates of the initial human body bounding box, the coordinates of the human body key points, and the coordinates of the facial key points corresponding to multiple objects to be trained, combined with the face posture estimation algorithm, sequentially obtain the head yaw angle, head pitch angle, and predicted head bounding box area corresponding to multiple objects to be trained in each input image when the head roll angle is the target head roll angle; determine the ratio between the predicted head bounding box area and the standard head bounding box area of ​​multiple objects to be trained, obtain the initial head angle correction coefficient, and use the initial head angle correction coefficient to obtain the mapping function corresponding to the seat to be trained.

[0120] In some embodiments of the present disclosure, the training unit is used to: combine the initial head angle correction coefficient, the head yaw angle, and the head pitch angle in a preset order, obtain a data group for each of the multiple subjects to be trained, and store the data group in a database; summarize the data group of each subject to be trained, determine similar data groups with the same head yaw angle and head pitch angle in the data group of each subject to be trained, and average the initial head angle correction coefficients in the similar data groups to obtain the head angle correction coefficient; use the head yaw angle, the head pitch angle, and the head angle correction coefficient as inputs of a three-dimensional function, and fit a mapping function corresponding to the seat to be trained, wherein the mapping function includes a mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient under the seat to be trained.

[0121] In some embodiments of the present disclosure, the correction unit 1040 is configured to multiply the head angle correction coefficient by the head circumscribed frame area of ​​the person object to be predicted to obtain the actual head circumscribed frame area of ​​the person object to be predicted.

[0122] In some embodiments of the present disclosure, the device 1000 also includes: an alarm unit, which is used to determine the distance between the actual head external frame area and the camera after correcting the head external frame area of ​​the predicted person object using the head angle correction coefficient and determining the actual head external frame area; if the distance is less than or equal to the preset safety distance, send an alarm message to the predicted person object.

[0123] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.

[0124] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a vehicle.

[0125] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0126] like Figure 11 As shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1102 or a computer program loaded from a storage unit 1108 into a RAM (Random Access Memory) 1103. Various programs and data required for the operation of the device 1100 can also be stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An I / O (Input / Output) interface 1105 is also connected to the bus 1104.

[0127] Various components in device 1100 are connected to I / O interface 1105, including an input unit 1106, such as a keyboard and mouse; an output unit 1107, such as various types of displays and speakers; a storage unit 1108, such as a magnetic disk and optical disk; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0128] The computing unit 1101 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the method for determining the area of ​​a human head circumscribed frame. For example, in some embodiments, the method for determining the area of ​​a human head circumscribed frame can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the aforementioned method for determining the area of ​​the circumscribed frame of a human head in any other appropriate manner (for example, by means of firmware).

[0129] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0134] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0135] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0136] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0137] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for determining the area of ​​a human head circumference frame, characterized in that: The method comprises: Get the associated face bounding box of the person to be predicted; Determine the head yaw angle and head pitch angle corresponding to the associated face circumscribed frame based on the coordinates of the facial key points in the associated face circumscribed frame and in combination with a face pose estimation algorithm; calling a mapping function corresponding to the seat information of the person to be predicted, and determining a head angle correction coefficient corresponding to the head bounding box of the person to be predicted using a mapping relationship among the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function, wherein the seat information is obtained by pairing the person to be predicted with a human body seat; The head angle correction coefficient is used to correct the head circumference frame area of ​​the person object to be predicted to determine the actual head circumference frame area.

2. The method according to claim 1, characterized in that The step of obtaining the associated face bounding box of the person object to be predicted includes: Acquire an image to be predicted, and perform human head and face classification detection and human body key point detection on the image to be predicted, to determine an initial head circumscribed frame and human body key points of the person to be predicted in the image to be predicted; Determine the key point straight line and the key point straight line distance between the chin point and the head vertex of the human body key point in the initial head circumscribed frame; Determining a tangent value of a head roll angle corresponding to an intersection line intersecting the initial head circumscribed frame, where the intersection line passes through the chin point and is perpendicular to the key point line; Determine the intersection straight line distance of the intersection straight line in the initial head circumscribed frame using the tangent value, the vertex coordinates of the initial head circumscribed frame, and the chin point coordinates of the chin point; The key point straight line is used as the length of the predicted head circumference frame, and the intersection straight line is used as the width of the predicted head circumference frame to obtain the predicted head circumference frame; Perform head and face pairing on the person object to be predicted to obtain an associated face bounding box that uniquely matches the predicted head bounding box.

3. The method according to claim 2, characterized in that The correcting the head circumference frame area of ​​the to-be-predicted person object by using the head angle correction coefficient to determine the actual head circumference frame area includes: Multiplying the straight-line distance between the key points and the straight-line distance between the intersection points to determine the predicted area of ​​the head bounding box; The predicted head circumference frame area is used as the head circumference frame area of ​​the person object to be predicted, and the head circumference frame area is corrected using the head angle correction coefficient to determine the actual head circumference frame area.

4. The method according to claim 1, wherein The step of obtaining the associated face bounding box of the person object to be predicted includes: Obtaining a head key point of a human body key point and the head key point coordinates corresponding to the head key point, and determining a head key point circumscribed frame corresponding to the head key point coordinates; Determine the head intersection-over-union ratio between the head key point bounding box and the initial head bounding box; If the IoU is greater than or equal to a preset IoU threshold, determining that the initial human bounding box of the person to be predicted matches the initial head bounding box, where the initial human bounding box is obtained by performing human head and face classification detection on the image to be predicted; When the initial human body circumscribed frame of the to-be-predicted person object matches the initial head circumscribed frame, an associated face circumscribed frame of the to-be-predicted person object is obtained.

5. The method according to claim 2, characterized in that The step of performing head and face pairing on the person to be predicted to obtain an associated face bounding box that uniquely matches the predicted head bounding box includes: Determining a face intersection-over-union ratio between an initial face bounding box and the predicted head bounding box, wherein the initial face bounding box is obtained by performing human head and face classification detection on the image to be predicted; The initial face bounding box whose face intersection-over-universal ratio is greater than or equal to a preset face intersection-over-universal ratio threshold is used as the associated face bounding box that uniquely matches the predicted head bounding box.

6. The method according to claim 1, wherein Before the step of calling a mapping function corresponding to the seat information of the person to be predicted, and determining the head angle correction coefficient corresponding to the head bounding box of the person to be predicted by using the mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function, the method includes: determining the areas of standard human head circumscribed frames for a plurality of subjects to be trained under a seat to be trained, when the head roll angle, the head yaw angle, and the head pitch angle meet target conditions, wherein the target conditions include the head roll angle being a target head roll angle, the head yaw angle being a target head yaw angle, and the head pitch angle being a target head pitch angle; According to the head rotation directions of the multiple subjects to be trained, sequentially obtain the vertex coordinates of the initial head circumscribed frame, the vertex coordinates of the initial human body circumscribed frame, the coordinates of the human body key points, and the coordinates of the human face key points of the multiple subjects to be trained; Using the vertex coordinates of the initial head bounding box, the vertex coordinates of the initial human body bounding box, the coordinates of the human body key points, and the coordinates of the facial key points corresponding to the multiple subjects to be trained, in combination with a face pose estimation algorithm, sequentially obtain the head yaw angle, the head pitch angle, and the predicted head bounding box area corresponding to the multiple subjects to be trained in each input image when the head roll angle is the target head roll angle; Determine the ratio between the predicted head circumscribed frame area and the standard head circumscribed frame area of ​​the multiple subjects to be trained, obtain an initial head angle correction coefficient, and use the initial head angle correction coefficient to obtain a mapping function corresponding to the seat to be trained.

7. The method according to claim 6, characterized in that Using the initial head angle correction coefficient, obtaining the mapping function corresponding to the seat to be trained includes: Combining the initial head angle correction coefficient, the head yaw angle, and the head pitch angle in a preset order, respectively obtaining a data group for each of the multiple subjects to be trained, and storing the data group in a database; Summarizing the data groups of each subject to be trained, determining similar data groups with the same head yaw angle and head pitch angle in the data groups of each subject to be trained, and averaging the initial head angle correction coefficients in the similar data groups to obtain the head angle correction coefficient; The head yaw angle, the head pitch angle, and the head angle correction coefficient are used as inputs of a three-dimensional function, and a mapping function corresponding to the seat to be trained is fitted. The mapping function includes a mapping relationship between the head yaw angle, the head pitch angle, and the head angle correction coefficient under the seat to be trained.

8. The method according to claim 1, characterized in that The step of correcting the head bounding box area of ​​the to-be-predicted person object by using the head angle correction coefficient to obtain the actual head bounding box area includes: The head angle correction coefficient is multiplied by the head circumscribed frame area of ​​the person object to be predicted to obtain the actual head circumscribed frame area of ​​the person object to be predicted.

9. The method according to claim 1, characterized in that After the step of correcting the head circumference frame area of ​​the person object to be predicted by using the head angle correction coefficient to determine the actual head circumference frame area, the method includes: Determining the distance between the actual head circumference frame area and the camera; If the distance is less than or equal to the preset safety distance, an alarm message is sent to the person object to be predicted.

10. A device for determining the area of ​​a human head external frame, characterized in that: include: A pairing unit, used to obtain the associated face bounding box of the person object to be predicted; A first determining unit is configured to determine the head yaw angle and head pitch angle corresponding to the associated face circumscribed frame based on the coordinates of the facial key points in the associated face circumscribed frame in combination with a face pose estimation algorithm; a second determining unit, configured to call a mapping function corresponding to the seat information of the person to be predicted, and determine a head angle correction coefficient corresponding to the head bounding box of the person to be predicted by using a mapping relationship among the head yaw angle, the head pitch angle, and the head angle correction coefficient in the mapping function, wherein the seat information is obtained by pairing the person to be predicted with a human body and a seat; The correction unit is used to correct the head circumference frame area of ​​the person object to be predicted by using the head angle correction coefficient to determine the actual head circumference frame area.

11. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

13. A vehicle, characterized in that: It includes the device for determining the area of ​​the external frame of a human head as described in claim 10 or the electronic device as described in claim 11.