Detection Method, Device, Electronic Device and Storage Medium

By obtaining the head center point and image acquisition equipment calibration parameters to calculate the head size, the problems of high annotation cost and perspective effect of human head detection in the prior art are solved, and a higher accuracy and stable detection effect is achieved.

CN114677367BActive Publication Date: 2025-07-18BEIJING SENSETIME TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210411308.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-07-18
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

In the prior art, the human head detection technology has high labeling cost, low labeling efficiency, susceptibility to perspective effects and unstable detection accuracy, especially in different shooting scenes, which are difficult to maintain high accuracy.

Method used

By obtaining the center point of the head in the image to be detected, combining the calibration parameters of the image acquisition device, calculate the head size and determine the head area, reduce the impact on the perspective effect, and adapt to different shooting scenes.

Benefits of technology

It improves the accuracy and stability of human head detection, reduces the cost and error of manual labeling, and enhances the applicability of the detection method in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677367B_ABST
    Figure CN114677367B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a detection method, apparatus, electronic device, and storage medium. The method includes: obtaining a to-be-detected image, where the to-be-detected image is acquired by an image acquisition device; determining a head center point corresponding to a target object in the to-be-detected image; determining the head size of the target object in the to-be-detected image according to the position of the head center point in the to-be-detected image and the calibration parameters of the image acquisition device; and determining a head region corresponding to the target object in the to-be-detected image according to the head center point corresponding to the target object in the to-be-detected image and the corresponding head size, and using it as a detection result. The embodiments of the present disclosure can be adapted to different shooting scenarios, thereby improving the stability of detection accuracy. In addition, the embodiments of the present disclosure can also reduce the influence of the perspective effect on the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a detection method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of related technologies such as video security and augmented reality, the human head detection technology has gradually attracted the attention of developers. The human head detection technology can support the upper-layer tasks related thereto, such as pedestrian counting, pedestrian distance estimation, etc. in the scenarios of smart city and intelligent security, or the human body pose reconstruction task in the augmented reality scenario. And the detection accuracy of the human head detection technology will affect the execution effect of the upper-layer tasks. Therefore, how to perform human head detection more accurately is a technical problem that developers urgently need to solve. Summary of the Invention

[0003] The present disclosure proposes a detection technical solution.

[0004] According to one aspect of the present disclosure, there is provided a detection method, including: obtaining an image to be detected, where the image to be detected is collected by an image acquisition device; determining a head center point corresponding to a target object in the image to be detected; determining a head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device; and determining a head region corresponding to the target object in the image to be detected according to the head center point corresponding to the target object in the image to be detected and the corresponding head size, and using it as a detection result.

[0005] In a possible implementation manner, the determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device includes: determining a first size range corresponding to the target object in the image to be detected according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device; where the pixel coordinates corresponding to the head center point are used to represent the position of the head center point in the image to be detected; determining a second size range corresponding to the target object according to the first size range; where the measurement directions of the first size range and the second size range are different; and using the first size range and the second size range as the head size of the target object in the image to be detected.

[0006] In a possible implementation manner, determining the first size range corresponding to the target object in the to-be-detected image according to the pixel coordinates corresponding to the center point of the head and the calibration parameters of the image acquisition device includes: determining the conversion relationship between the pixel coordinates in the to-be-detected image and the world coordinates in the world coordinate system according to the calibration parameters corresponding to the image acquisition device; determining the first size range corresponding to the target object in the to-be-detected image according to the conversion relationship and the preset human body parameters; wherein, the preset human body parameters are measured based on the measurement standard of the world coordinate system.

[0007] In a possible implementation manner, determining the first size range corresponding to the target object in the to-be-detected image according to the conversion relationship and the preset human body parameters includes: determining the first world sub-coordinates of the center point of the head in the world coordinate system according to the installation height of the image acquisition device and the preset human body parameters; wherein, the first world sub-coordinates have the same measurement direction as the first size range; determining the second world sub-coordinates of the center point of the head in the world coordinate system according to the conversion relationship, the pixel coordinates corresponding to the center point of the head, and the first world sub-coordinates; wherein, the second world sub-coordinates have a different measurement direction from the first world sub-coordinates; determining the world coordinates of the two vertices of the head region corresponding to the center point of the head in the world coordinate system according to the first world sub-coordinates and the preset human body parameters; wherein, the numerical values of the world coordinates of the two vertices of the head region in the measurement direction of the first world sub-coordinates are different; determining the pixel coordinates corresponding to the two vertices of the head region according to the world coordinates of the two vertices of the head region, the second world sub-coordinates, and the conversion relationship; determining the first size range corresponding to the target object in the to-be-detected image based on the offset amount of the two vertices of the head region in the measurement direction of the first world sub-coordinates; wherein, the offset amount is determined based on the pixel coordinates corresponding to the two vertices of the head region.

[0008] In a possible implementation manner, determining the first size range corresponding to the target object in the to-be-detected image based on the offset amount of the two vertices of the head region in the measurement direction of the first world sub-coordinates includes: determining the offset amount of the pixel coordinates corresponding to the two vertices of the head region in the measurement direction of the first world sub-coordinates; in the case where the offset amount is greater than the preset offset amount, using the offset amount as the first size range corresponding to the target object in the to-be-detected image.

[0009] In a possible implementation, the first size range and the second size range are respectively one of the width and height of the head size, and have a preset proportional relationship. Determining the second size range corresponding to the target object according to the first size range includes: determining the second size range of the target object based on the first size range and the preset proportional relationship.

[0010] In a possible implementation, obtaining the image to be detected includes: obtaining at least two images to be detected with a target object through at least two image acquisition devices; determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device includes: determining the head size corresponding to the target object in each image to be detected according to the positions of the head center point in the at least two images to be detected and the calibration parameters of the at least two image acquisition devices.

[0011] In a possible implementation, the detection method further includes: determining the facial features corresponding to the target object based on the detection result; and generating a prompt message when it is determined that the facial features corresponding to the target object meet the preset conditions.

[0012] According to one aspect of the present disclosure, there is provided a detection device, including: an image to be detected acquisition module for acquiring an image to be detected, where the image to be detected is acquired through an image acquisition device; a head center point determination module for determining the head center point corresponding to the target object in the image to be detected; a head size determination module for determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device; and a head region determination module for determining the head region corresponding to the target object in the image to be detected according to the head center point corresponding to the target object in the image to be detected and the corresponding head size, and using it as the detection result.

[0013] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0014] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0015] In an embodiment of the present disclosure, a to-be-detected image can be obtained, and then the head center point corresponding to the target object in the to-be-detected image can be determined. Next, according to the position of the head center point in the to-be-detected image and the calibration parameters of the image acquisition device of the to-be-detected image, the head size of the target object in the to-be-detected image can be determined. Finally, according to the head center point corresponding to the target object in the to-be-detected image and the corresponding head size, the head region corresponding to the target object in the to-be-detected image can be determined and used as the detection result. The embodiment of the present disclosure can determine the head region of the target object according to the calibration parameters of the image acquisition device, so that the above detection method can be adapted to different shooting scenarios, and thus the stability of the detection accuracy can be improved. In addition, the above head region is generated based on the head center point and the head size. Compared with the method of directly generating the head region in the above related technology, the embodiment of the present disclosure can reduce the influence of the perspective effect on the detection accuracy.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0018] Figure 1 Shows a flowchart of a detection method provided according to an embodiment of the present disclosure.

[0019] Figure 2 Shows a flowchart of a detection method provided according to an embodiment of the present disclosure.

[0020] Figure 3 Shows a block diagram of a detection device provided according to an embodiment of the present disclosure.

[0021] Figure 4 Shows a block diagram of an electronic device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0023] The special term "exemplary" herein means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior to or better than other embodiments.

[0024] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0025] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can still be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.

[0026] In related technologies, human head detection technology is usually performed through a machine learning model. For example, developers first label the head region in the training image with a bounding box, and then train the machine learning model based on the training image and its corresponding bounding box so that the trained machine learning model can identify the head region in the image to be detected. However, this approach is prone to the following problems: 1. The bounding box is usually manually labeled, with a high labeling cost and poor labeling efficiency. 2. In a crowded situation, the bounding boxes are prone to overlap, which is not conducive to visual observation by humans and is prone to manual labeling errors. 3. This approach is easily affected by the perspective effect, that is, it is prone to a decrease in detection accuracy. 4. The generalization ability of the machine learning model is affected by different shooting scenarios, and for different shooting scenarios, the detection accuracy of the machine learning model is uncontrollable.

[0027] In view of this, the embodiments of the present disclosure provide a detection method. It can obtain an image to be detected, then determine the head center point corresponding to the target object in the image to be detected, and then determine the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device of the image to be detected. Finally, according to the head center point corresponding to the target object and the corresponding head size in the image to be detected, determine the head region corresponding to the target object in the image to be detected and use it as the detection result. The embodiments of the present disclosure can determine the head region of the target object according to the calibration parameters of the image acquisition device, so that the above detection method can be adapted to different shooting scenarios, and thus the stability of the detection accuracy can be improved. In addition, the above head region is generated based on the head center point and the head size. Compared with the method of directly generating the head region in the above related technologies, the embodiments of the present disclosure can reduce the influence of the perspective effect on the detection accuracy.

[0028] In a possible implementation, the detection method may be executed by an electronic device such as a terminal device or a server. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method may be implemented by a processor invoking computer-readable instructions stored in a memory. Alternatively, the method may be executed by a server.

[0029] Refer to Figure 1 as shown Figure 1 which shows a flowchart of the detection method provided according to an embodiment of the present disclosure. As Figure 1 shown, the above detection method includes:

[0030] Step S100, obtain an image to be detected. Among them, the image to be detected is obtained by an image acquisition device. Exemplarily, the above electronic device may be wired or wirelessly connected to the above image acquisition device so that the electronic device can obtain the image to be detected through the image acquisition device. The above image acquisition device may include: a visible light camera, a multi-camera, etc., and developers can flexibly set according to actual needs. In one implementation, the image to be detected is the original image acquired by the image acquisition device; in other implementations, the image to be detected may be a processed image, for example, an image processed by enhancement processing, screening, etc., or an image stitched after being acquired by multiple image acquisition devices, which is not specifically limited here.

[0031] Step S200: Determine the head center point corresponding to the target object in the image to be detected. The above target object can be any object that the developer wants to detect, such as pedestrians, animals, specific persons, specific animals, etc., which are not limited in the embodiments of the present disclosure. Exemplarily, the above head center point can be obtained through a trained machine learning model. For example, the training objects in the training images can be manually labeled. The above training objects have the same object category as the target object, such as both being pedestrians, both being animals, both being specific persons, etc., so that the trained machine learning model can determine the target object in the image to be detected. Exemplarily, the head center points of the training objects in the training images can be labeled. In other words, based on the head center points of the training objects, the trained machine learning model can identify the head center point corresponding to the target object. For example, after the machine learning model is trained, the image to be detected can be input into it to obtain the position information corresponding to the head center point of the target object in the image to be detected. In one example, the position information of the above head center point can be represented as pixel coordinates in the pixel coordinate system in the related art. The embodiments of the present disclosure only need to allow the machine learning model to output the head center point. Compared with the machine learning model in the related art that directly outputs the head region, the manual annotation cost of the embodiments of the present disclosure is lower and the manual annotation error is smaller. In other words, in the annotation stage of the training images, the developer only needs to annotate the head center points of the training objects. Compared with annotating the complete head region, the manual annotation cost of the embodiments of the present disclosure is lower. In addition, compared with the situation of overlapping annotation frames in the related art, the visual effect of annotating the head center point is better, that is, usually the distance between the annotation points is larger. Therefore, even when there are too many training objects in the training images, the developer can complete highly recognizable annotations, which is beneficial to improving the detection effect of the machine learning model trained based on the training images.

[0032] Step S300: Determine the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device. Exemplarily, when the calibration parameters of the image acquisition device are known, the above calibration parameters of the image acquisition device can be directly obtained. When the image acquisition device is initially set in the shooting area, the calibration parameters of the image acquisition device can be determined through the pose information of the image acquisition device. Exemplarily, the above calibration parameters can be represented as the external parameters of the image acquisition device, and the above pose information can include information such as the shooting angle of the image acquisition device. The above head size can be represented based on the pixel coordinate difference in the above pixel coordinate system.

[0033] Refer to Figure 2 shown Figure 2 shows a flowchart of the detection method provided according to an embodiment of the present disclosure. As Figure 2As shown, in a possible implementation, step S300 may include:

[0034] Step S310: Determine a first size range corresponding to the target object in the image to be detected according to the pixel coordinates corresponding to the center point of the head and the calibration parameters of the image acquisition device. Among them, the pixel coordinates corresponding to the center point of the head are used to represent the position of the center point of the head in the image to be detected, and the above pixel coordinates are also the coordinates generated based on the above pixel coordinate system. Exemplarily, if the head region that the developer wants to obtain is a rectangular region, the above first size range may be the height range or the width range of the target object. The specific shape of the head region and the meaning of the corresponding first size range are not limited in this embodiment of the present disclosure and can be set according to the needs of the developer. It is only necessary to obtain the head region based on the first size range. In the following, the head region will be taken as a rectangle for example in the embodiments of the present disclosure.

[0035] In a possible implementation, step S310 may include: When the calibration parameters of the image acquisition device are unknown (for example, when the image acquisition device is initially installed at a position and the external parameters have not been calibrated according to the measurement standard in the world coordinate system), the calibration parameters corresponding to the image acquisition device may be determined first according to the device angle of the image acquisition device. Exemplarily, the above device angle may be expressed as the camera pose angle in the related art (such as: pitch angle, yaw angle, roll angle). The above calibration parameters may include a rotation matrix. Considering the actual usage scenario, usually the pitch angle of the camera has a relatively high influence on the perspective relationship in the image to be detected. Therefore, the pitch angle of the camera may be selected as the above device angle to reduce the influence of the perspective effect on the detection accuracy. Then, according to the calibration parameters corresponding to the image acquisition device, determine the conversion relationship between the pixel coordinates in the image to be detected and the world coordinates in the world coordinate system. Exemplarily, there is a mapping relationship between the above world coordinate system and the pixel coordinate system corresponding to the image to be detected. The world coordinates in the above world coordinate system (which can be calibrated by the calibration algorithm in the related art) are used to represent the position of the target object in the objective world, and the above pixel coordinate system (which can be established with the upper left vertex of the captured image as the origin) is used to represent the position of the target object in the captured image.

[0036] Exemplarily, the rotation matrix R of the image acquisition device may be determined first according to the above pitch angle θ. For example, the relationship between the two may be set as follows:

[0037]

[0038] Based on this, according to the imaging model of the image acquisition device in the related technology, the corresponding relationship between the world coordinates (X, Y, Z) of any spatial point in the world coordinate system and the pixel coordinates (u, v, 1) (expressed as homogeneous coordinates here) can be obtained:

[0039]

[0040] Among them, Z ′ is used to represent the value of Z in the world coordinates (X, Y, Z) of the spatial point mapped to the camera coordinate system in the related technology, and can be used to represent the depth value in the camera coordinate system. u is used to represent the coordinate information of the spatial point in the vertical direction in the pixel coordinate system, and v is used to represent the coordinate information of the spatial point in the horizontal direction in the pixel coordinate system. is the internal parameter matrix of the image acquisition device, and the specific generation method can refer to the related technology. f x 、f y are used to represent the lengths of the focal lengths in the x and y directions (the unit can be the number of pixel points). The x direction is the above-mentioned vertical direction, and the y direction is the above-mentioned horizontal direction. c x 、c y are used to represent the offsets of the principal point in the camera coordinate system (such as the vertex at the upper left corner of the image) relative to the principal point in the pixel coordinate system in the x and y directions. R T is used to represent the transpose matrix of R. X, Y, and Z are used to represent the coordinate information of the spatial point in the vertical, horizontal, and depth directions in the world coordinate system.

[0041] After simplification, the following equations can be obtained and can be used as the above conversion relationship:

[0042] And

[0043] Then, according to the conversion relationship and the preset human body parameters, the first size range corresponding to the target object in the to-be-detected image can be determined. Among them, the preset human body parameters are measured based on the measurement standard of the world coordinate system, and the preset human body parameters may include at least one of body size and head size. Exemplarily, the above preset human body parameters may include: body size (for example: the height and width of the body, the unit can be meters), head size (for example: the height and width of the head, the unit can be meters).

[0044] In a possible implementation manner, the above determining the first size range corresponding to the target object in the to-be-detected image according to the conversion relationship and the preset human body parameters may include:

[0045] Determine the first world sub - coordinate of the head center point in the world coordinate system according to the installation height of the image acquisition device and the preset human body parameters. Wherein, the measurement direction of the first world sub - coordinate is the same as that of the first size range. Exemplarily, here, taking the spatial point as the head center point corresponding to the target object mentioned above, the first size range as the height range, and the coordinate of a1 as the vertical axis (used to measure height), if the coordinate of the head center point in the world coordinate system is (a1, b1, z1) and the coordinate in the pixel coordinate system is (u A , v A ), then the above - mentioned first world sub - coordinate can be expressed as a1. If the first size range is the width range and the coordinate of b1 is the horizontal axis (used to measure width), then the first world sub - coordinate can be expressed as b1. In other words, the measurement direction of the first world sub - coordinate is the same as that of the first size range. Combining the body height M, head height m in the preset human body parameters, and the installation height h of the image acquisition device, the above - mentioned first world sub - coordinate can be expressed as:

[0046]

[0047] Then, determine the second world sub - coordinate of the head center point in the world coordinate system according to the conversion relationship, the pixel coordinate corresponding to the head center point, and the first world sub - coordinate. Wherein, the measurement direction of the second world sub - coordinate is different from that of the first world sub - coordinate. Exemplarily, the above - mentioned second world sub - coordinate can be expressed as z1 in the above text. For example, it can be expressed as the coordinate of the depth axis (used to measure depth) in the world coordinate system. Continuing to refer to the above example, the second world sub - coordinate can be expressed as:

[0048]

[0049] Then, according to the first world sub - coordinate and the preset human body parameters, determine the world coordinates of the two vertices of the head region corresponding to the head center point in the world coordinate system. The numerical values of the world coordinates of the two vertices of the head region in the measurement direction of the first world sub - coordinate are different. Exemplarily, if the first measurement dimension is height, there only needs to be a height difference between the two vertices of the head region. In other words, in this case, the two vertices of the head region can be selected as: the upper - left vertex and the lower - left vertex of the head region, or the upper - right vertex and the lower - left vertex, etc. Here, taking the upper - left vertex and the lower - right vertex as an example: the world coordinate of the upper - left vertex can be expressed as: The pixel coordinate can be expressed as: The world coordinate of the lower - right vertex can be expressed as: The pixel coordinate can be expressed as Wherein, n is used to represent the head width.

[0050] Based on the world coordinates of the two vertices of the head region, the second world sub-coordinates, and the conversion relationship, determine the pixel coordinates corresponding to the two vertices of the head region. Exemplarily, substituting the world coordinates and pixel coordinates of the upper left corner vertex into the conversion relationship in the above text, it can be determined that:

[0051]

[0052] Substituting the world coordinates and pixel coordinates of the lower right corner vertex into the conversion relationship in the above text, it can be determined that:

[0053]

[0054] Then, based on the offset of the two vertices of the head region in the measurement direction of the first world sub-coordinates, determine the first size range corresponding to the target object in the image to be detected. Wherein, the offset is determined based on the pixel coordinates corresponding to the two vertices of the head region. Exemplarily, if Δu A is used to represent the first size range, it can be expressed as:

[0055]

[0056] Substitute the above various parameters into this formula to obtain the first size range.

[0057] In one example, the determining the first size range corresponding to the target object in the image to be detected based on the offset of the two vertices of the head region in the measurement direction of the first world sub-coordinates may include: determining the offset of the pixel coordinates corresponding to the two vertices of the head region in the measurement direction of the first world sub-coordinates, and then, when the offset is greater than a preset offset, using the offset as the first size range corresponding to the target object in the image to be detected. Exemplarily, taking the offset between the upper left corner vertex and the lower right corner vertex as an example, the preset offset can be set to 0. Since the of the upper left corner vertex is higher than that of the lower right corner vertex Therefore, when taking the difference between the two, Δu A should be greater than 0. If it is less than or equal to 0, it can be determined that there is an abnormal situation in the detection method. The detection of this target object can be abandoned, or an abnormal prompt can be generated, etc. The embodiments of the present disclosure do not limit this here.

[0058] Continue to refer to Figure 2, step S320, determine the second size range corresponding to the target object according to the first size range. Wherein, the measurement directions of the first size range and the second size range are different. Exemplarily, if the head region is rectangular, the first size range can be the height range or the width range, and the second size range can be the width range or the height range. In a possible implementation manner, the first size range and the second size range can be respectively one of the width and height of the head size, and can have a preset proportional relationship. On this basis, step S320 can include: determining the second size range of the target object based on the first size range and the preset proportional relationship. For example: If the second size range is represented as Δv A , and the ratio of the second size range to the first size range is two-thirds, then the following proportional relationship can be determined and used as the above preset proportional relationship:

[0059]

[0060] Step S330, use the first size range and the second size range as the head size corresponding to the target object in the to-be-detected image.

[0061] In a possible implementation, the above step S100 can include: obtaining at least two to-be-detected images with the target object through at least two image acquisition devices. Exemplarily, the shooting planes of the at least two image acquisition devices can be the same, that is, in the shooting perspectives of the multiple image acquisition devices, the perspective relationship of the target object is the same (that is, the sizes of the target objects captured are the same). Then, the electronic device determines the head size corresponding to the target object in each to-be-detected image according to the position of the head center point in the at least two to-be-detected images and the calibration parameters of the at least two image acquisition devices. Exemplarily, the head size corresponding to each to-be-detected image can be directly averaged or weighted averaged to obtain an average value. Finally, according to the head center points corresponding to the target object in different to-be-detected images and the above average value, the head region in each to-be-detected image is determined. The embodiments of the present disclosure can detect the head region of the target object through multiple image acquisition devices to improve the detection accuracy of the head region, and can be applied to application scenarios such as human pose reconstruction that require accurate acquisition of the head region.

[0062] Continue to refer to Figure 1 , step S400, determine the head region corresponding to the target object in the to-be-detected image according to the head center point corresponding to the target object in the to-be-detected image and the corresponding head size, and use it as the detection result. Exemplarily, the range of the head region can be determined by dividing the first size range and the second size range in the head size by 2 and adding them to the pixel coordinates in the corresponding measurement direction of the head center point.

[0063] In a possible implementation manner, the above detection method may further include: based on the detection result, determining the facial features corresponding to the target object. Then, when it is determined that the facial features corresponding to the target object meet the preset conditions, a prompt message is generated. Exemplarily, the above detection result may be input into a feature extraction model in the related art to determine the facial features corresponding to the target object. The above preset conditions can be used to screen whether the facial features of the target object are the facial features of a specific object, and the embodiments of the present disclosure do not limit this here. In the detection method provided by the embodiments of the present disclosure, since the detection accuracy of the head region is relatively high and the influence of the perspective effect is small, the accuracy of the facial detection result generated based on this head region can be improved, and it can be applied to scenarios such as security and personnel warning.

[0064] It can be understood that, without violating the principle logic, the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above method of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0065] In addition, the present disclosure also provides a detection device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any detection method provided by the present disclosure. For the corresponding technical solutions and descriptions, reference can be made to the corresponding records in the method part, and details will not be repeated here.

[0066] Figure 3 The block diagram of the detection device provided by the embodiment of the present disclosure is shown. As Figure 3 shown, the detection device 100 includes: a to-be-detected image acquisition module 110, configured to acquire a to-be-detected image, where the to-be-detected image is acquired by an image acquisition device. A head center point determination module 120, configured to determine the head center point corresponding to the target object in the to-be-detected image. A head size determination module 130, configured to determine the head size of the target object in the to-be-detected image according to the position of the head center point in the to-be-detected image and the calibration parameters of the image acquisition device. A head region determination module 140, configured to determine the head region corresponding to the target object in the to-be-detected image according to the head center point corresponding to the target object in the to-be-detected image and the corresponding head size, and use it as the detection result.

[0067] In a possible implementation, determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device includes: determining a first size range of the target object in the image to be detected according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device; wherein the pixel coordinates corresponding to the head center point are used to represent the position of the head center point in the image to be detected; determining a second size range corresponding to the target object according to the first size range; wherein the measurement directions of the first size range and the second size range are different; using the first size range and the second size range as the head size of the target object corresponding to the image to be detected.

[0068] In a possible implementation, determining the first size range of the target object in the image to be detected according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device includes: determining the conversion relationship between the pixel coordinates in the image to be detected and the world coordinates in the world coordinate system according to the calibration parameters corresponding to the image acquisition device; determining the first size range of the target object in the image to be detected according to the conversion relationship and the preset human body parameters; wherein the preset human body parameters are measured based on the measurement standard of the world coordinate system.

[0069] In a possible implementation, determining the first size range of the target object in the image to be detected according to the conversion relationship and the preset human body parameters includes: determining a first world sub-coordinate of the head center point in the world coordinate system according to the installation height of the image acquisition device and the preset human body parameters; wherein the first world sub-coordinate is in the same measurement direction as the first size range; determining a second world sub-coordinate of the head center point in the world coordinate system according to the conversion relationship, the pixel coordinates corresponding to the head center point, and the first world sub-coordinate; wherein the second world sub-coordinate is in a different measurement direction from the first world sub-coordinate; determining the world coordinates of two vertices of the head region corresponding to the head center point in the world coordinate system according to the first world sub-coordinate and the preset human body parameters; wherein the numerical values of the world coordinates of the two vertices of the head region in the measurement direction of the first world sub-coordinate are different; determining the pixel coordinates corresponding to the two vertices of the head region according to the world coordinates of the two vertices of the head region, the second world sub-coordinate, and the conversion relationship; determining the first size range of the target object in the image to be detected based on the offset between the two vertices of the head region in the measurement direction of the first world sub-coordinate; wherein the offset is determined based on the pixel coordinates corresponding to the two vertices of the head region.

[0070] In a possible implementation manner, determining the first size range corresponding to the target object in the to-be-detected image based on the offset of two vertices of the head region in the first world sub-coordinate measurement direction includes: determining the pixel coordinates corresponding to the two vertices of the head region and the offset in the first world sub-coordinate measurement direction; when the offset is greater than a preset offset, using the offset as the first size range corresponding to the target object in the to-be-detected image.

[0071] In a possible implementation manner, the first size range and the second size range are respectively one of the width and height of the head size and have a preset proportional relationship. Determining the second size range corresponding to the target object according to the first size range includes: determining the second size range corresponding to the target object based on the first size range and the preset proportional relationship.

[0072] In a possible implementation manner, acquiring the to-be-detected image includes: acquiring at least two to-be-detected images with a target object through at least two image acquisition devices; determining the head size of the target object in the to-be-detected image according to the position of the head center point in the to-be-detected image and the calibration parameters of the image acquisition device includes: determining the head size corresponding to the target object in each to-be-detected image according to the position of the head center point in the at least two to-be-detected images and the calibration parameters of the at least two image acquisition devices.

[0073] In a possible implementation manner, the detection device further includes: a prompt information generation module, configured to perform the following steps: determining the facial features corresponding to the target object based on the detection result; generating prompt information when it is determined that the facial features corresponding to the target object meet a preset condition.

[0074] This method has a specific technical association with the internal structure of the computer system and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing the data storage amount, reducing the data transmission amount, and increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in line with natural laws.

[0075] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0076] Embodiments of the present disclosure also provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the above method. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0077] Embodiments of the present disclosure also provide an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0078] Embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code runs on a processor of an electronic device, the processor in the electronic device executes the above method.

[0079] The electronic device may be provided as a server or other form of device.

[0080] Figure 4 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 4 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0081] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.

[0082] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.

[0083] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0084] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0085] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0086] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0087] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0088] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture comprising instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0089] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0090] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0091] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.

[0092] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be elaborated herein.

[0093] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of the steps does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0094] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent logo is set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious logos / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0095] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A detection method, characterized in that, The detection method includes: Obtain an image to be detected, where the image to be detected is obtained by an image acquisition device; Determine the head center point corresponding to the target object in the image to be detected; According to the calibration parameters corresponding to the image acquisition device, determine the conversion relationship between the pixel coordinates in the image to be detected and the world coordinates in the world coordinate system; According to the conversion relationship and preset human body parameters, determine the first size range corresponding to the target object in the image to be detected; where the preset human body parameters are measured based on the measurement standard of the world coordinate system; According to the first size range, determine the second size range corresponding to the target object; where the measurement directions of the first size range and the second size range are different; Use the first size range and the second size range as the head size corresponding to the target object in the image to be detected; According to the head center point corresponding to the target object in the image to be detected and the corresponding head size, determine the head region corresponding to the target object in the image to be detected and use it as the detection result.

2. The detection method according to claim 1, wherein The step of determining the first size range corresponding to the target object in the image to be detected according to the conversion relationship and preset human body parameters includes: According to the installation height of the image acquisition device and preset human body parameters, determine the first world sub-coordinate of the head center point in the world coordinate system; where the first world sub-coordinate has the same measurement direction as the first size range; According to the conversion relationship, the pixel coordinates corresponding to the head center point, and the first world sub-coordinate, determine the second world sub-coordinate of the head center point in the world coordinate system; where the second world sub-coordinate has a different measurement direction from the first world sub-coordinate; According to the first world sub-coordinate and the preset human body parameters, determine the world coordinates of the two vertices of the head region corresponding to the head center point in the world coordinate system; where the numerical values of the world coordinates of the two vertices of the head region in the measurement direction of the first world sub-coordinate are different; According to the world coordinates of the two vertices of the head region, the second world sub-coordinate, and the conversion relationship, determine the pixel coordinates corresponding to the two vertices of the head region; Based on the offset of the two vertices of the head region in the measurement direction of the first world sub-coordinate, determine the first size range corresponding to the target object in the image to be detected; where the offset is determined based on the pixel coordinates corresponding to the two vertices of the head region.

3. The detection method according to claim 2, wherein, The step of determining the first size range corresponding to the target object in the image to be detected based on the offset of the two vertices of the head region in the measurement direction of the first world sub-coordinate includes: Determine the offset of the pixel coordinates corresponding to the two vertices of the head region in the measurement direction of the first world sub-coordinate; When the offset is greater than the preset offset, use the offset as the first size range corresponding to the target object in the image to be detected.

4. The detection method according to any one of claims 1 to 3, characterized in that, The first size range and the second size range are respectively one of the width and height of the head size, and have a preset proportional relationship. Determining the second size range corresponding to the target object according to the first size range includes: Based on the first size range and the preset proportional relationship, determine the second size range of the target object.

5. The detection method according to any one of claims 1 to 3, characterized in that, Obtaining the image to be detected includes: obtaining at least two images to be detected with a target object through at least two image acquisition devices; Determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device includes: According to the positions of the head center point in the at least two images to be detected and the calibration parameters of the at least two image acquisition devices, determine the head size corresponding to the target object in each image to be detected.

6. The detection method according to any one of claims 1 to 3, characterized in that, The detection method further includes: Based on the detection result, determine the facial features corresponding to the target object; Generate a prompt message when it is determined that the facial features corresponding to the target object meet the preset conditions.

7. A detection device, characterized in that, The detection device includes: An image acquisition module for obtaining an image to be detected, where the image to be detected is acquired by an image acquisition device; A head center point determination module for determining the head center point corresponding to the target object in the image to be detected; A head size determination module for determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device; A head region determination module for determining the head region corresponding to the target object in the image to be detected according to the head center point corresponding to the target object in the image to be detected and the corresponding head size, and using it as the detection result; Among them, determining the head size of the target object in the image to be detected according to the position of the head center point in the image to be detected and the calibration parameters of the image acquisition device includes: Determine the first size range corresponding to the target object in the image to be detected according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device; where the pixel coordinates corresponding to the head center point are used to represent the position of the head center point in the image to be detected; According to the first size range, determine the second size range corresponding to the target object; where the measurement directions of the first size range and the second size range are different; Use the first size range and the second size range as the head size corresponding to the target object in the image to be detected; Determining the first size range corresponding to the target object in the image to be detected according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device includes: According to the calibration parameters corresponding to the image acquisition device, determine the conversion relationship between the pixel coordinates in the image to be detected and the world coordinates in the world coordinate system; Determine a first size range corresponding to the target object in the image to be detected according to the conversion relationship and preset human body parameters; wherein, the preset human body parameters are measured based on the measurement standard of the world coordinate system.

8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the detection method according to any one of claims 1 to 6 is implemented.