Detection method and device, electronic equipment and storage medium

By fusing the calibration parameters and mapping relationships of the image acquisition device to determine the human head size, the problem of high annotation cost and uncontrollable accuracy in human head detection in the prior art is solved, and higher accuracy and robustness detection are achieved.

CN114708556BActive Publication Date: 2025-12-16BEIJING SENSETIME TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210410112.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-12-16
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

Existing technologies for human head detection suffer from problems such as high annotation costs, low annotation efficiency, susceptibility to perspective effects, and uncontrollable detection accuracy.

Method used

By acquiring the image to be detected, using the calibration parameters of the image acquisition device and the preset mapping relationship, the first head size and the second head size of the target object are determined, and then fused to obtain the fused head size, and finally the head region is determined.

Benefits of technology

It reduces sensitivity to calibration parameters, improves the robustness and adaptability of detection accuracy, and enhances detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708556B_ABST
    Figure CN114708556B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a detection method and device, electronic equipment and storage medium. The method comprises: obtaining a to-be-detected image; wherein the to-be-detected image is obtained by an image acquisition device; determining a first head size corresponding to a target object in the to-be-detected image according to a position of the target object in the to-be-detected image and a calibration parameter of the image acquisition device; determining a second head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image and a preset mapping relationship; wherein the first head size and the second head size have the same measurement direction; and fusing the first head size and the second head size to obtain a fused head size corresponding to the target object. The present disclosure can reduce the parameter sensitivity of the first head size generated by different calibration parameters, and is beneficial to improving the robustness of detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and particularly relates to a detection method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of video security, augmented reality and other related technologies, human head detection technology gradually attracts the attention of developers. Human head detection technology can support related upper-layer tasks, such as pedestrian counting, pedestrian distance estimation and other tasks in the smart city and intelligent security scene, or human pose reconstruction tasks in the augmented reality scene. The detection accuracy of human head detection technology will affect the execution effect of the upper-layer tasks. Therefore, how to more accurately perform human head detection is a technical problem that developers urgently need to solve. SUMMARY

[0003] The present disclosure provides a detection technical solution.

[0004] According to an aspect of the present disclosure, a detection method is provided, including: obtaining a to-be-detected image; wherein the to-be-detected image is obtained by an image acquisition device; determining a first head size corresponding to a target object in the to-be-detected image according to the position of the target object in the to-be-detected image and the calibration parameters of the image acquisition device; determining a second head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image and a preset mapping relationship; wherein the first head size and the second head size are in the same measurement direction; fusing the first head size and the second head size to obtain a fused head size corresponding to the target object; and determining a head region corresponding to the target object in the to-be-detected image according to the fused head size and the position of the target object in the to-be-detected image, and taking the head region as a detection result.

[0005] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image includes: determining a head center point corresponding to the target object in the to-be-detected image; and determining the first head size corresponding to the target object in the to-be-detected image according to the pixel coordinates of the head center point and the calibration parameters of the image acquisition device; wherein the pixel coordinates of the head center point are used to represent the position of the head center point in the to-be-detected image.

[0006] In a possible implementation, the determining the second head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image and the preset mapping relationship comprises: determining a head center point corresponding to the target object in the to-be-detected image; and determining the second head size corresponding to the target object in the to-be-detected image according to a pixel coordinate corresponding to the head center point and the preset mapping relationship.

[0007] In a possible implementation, the detection method further comprises: obtaining a training image with a training object; wherein the training object is of the same object category as the target object; determining a body size corresponding to the training object according to the training image; determining a head size corresponding to the training object according to the body size; and determining the preset mapping relationship according to the head size corresponding to the training object and a pixel coordinate corresponding to the training object; wherein the pixel coordinate corresponding to the training object is used to represent the position of the training object in the training image.

[0008] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image according to the pixel coordinate corresponding to the head center point and the calibration parameter of the image acquisition device comprises: determining a conversion relationship between the pixel coordinate in the to-be-detected image and a world coordinate in a world coordinate system according to the calibration parameter corresponding to the image acquisition device; and determining the first head size corresponding to the target object in the to-be-detected image according to the conversion relationship, a preset human body parameter and the pixel coordinate corresponding to the head center point; wherein the preset human body parameter is measured based on a measurement standard of the world coordinate system.

[0009] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image according to the conversion relationship, the preset human body parameter, and the pixel coordinates corresponding to the head center point comprises: determining a first world sub-coordinate of the head center point in the world coordinate system according to the installation height of the image acquisition device and the preset human body parameter, wherein the first world sub-coordinate is in the same measurement direction as the first head size; determining a second world sub-coordinate of the head center point in the world coordinate system according to the conversion relationship, the pixel coordinates corresponding to the head center point, and the first world sub-coordinate, wherein the second world sub-coordinate is in a different measurement direction from the first world sub-coordinate; determining world coordinates of two vertices of a head region corresponding to the head center point in the world coordinate system according to the first world sub-coordinate and the preset human body parameter, wherein the world coordinates of the two vertices of the head region are different in the measurement direction of the first world sub-coordinate; determining pixel coordinates corresponding to the two vertices of the head region according to the world coordinates of the two vertices of the head region, the second world sub-coordinate, and the conversion relationship; and determining the first head size corresponding to the target object in the to-be-detected image based on an offset of the two vertices of the head region in the measurement direction of the first world sub-coordinate, wherein the offset is determined based on the pixel coordinates corresponding to the two vertices of the head region.

[0010] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image based on the offset of the two vertices of the head region in the measurement direction of the first world sub-coordinate comprises: determining the offset of the pixel coordinates corresponding to the two vertices of the head region in the measurement direction of the first world sub-coordinate; and in a case where the offset is greater than a preset offset, taking the offset as the first head size corresponding to the target object in the to-be-detected image.

[0011] In a possible implementation, the determining the head region corresponding to the target object in the to-be-detected image according to the fused head size and the position of the target object in the to-be-detected image comprises: determining a third head size corresponding to the target object in the to-be-detected image according to the fused head size and a preset proportional relationship between the fused head size and the third head size, wherein the fused head size and the third head size are one of a width and a height of a head region; and determining the head region corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image, the fused head size, and the third head size.

[0012] In a possible implementation, the fusing the first head size and the second head size to obtain the fused head size corresponding to the target object comprises: fusing the first head size, the second head size, a weight corresponding to the first head size, and a weight corresponding to the second head size, to obtain the fused head size corresponding to the target object; and the weights are related to a position of the target object in the to-be-detected image and a difference between the first head size and the second head size.

[0013] In a possible implementation, the obtaining the to-be-detected image comprises: obtaining at least two to-be-detected images with the target object by using at least two image acquisition devices; and the determining the head region corresponding to the target object in the to-be-detected image according to the fused head size and the position of the target object in the to-be-detected image comprises: fusing the fused head sizes obtained by using each image acquisition device to obtain a multi-device fusion result; and determining the head region corresponding to the target object in each to-be-detected image according to the multi-device fusion result and the position of the target object in the to-be-detected image.

[0014] According to an aspect of the present disclosure, a detection device is provided, which comprises: a to-be-detected image obtaining module configured to obtain a to-be-detected image; the to-be-detected image is obtained by using an image acquisition device; a first head size generating module configured to determine a first head size corresponding to a target object in the to-be-detected image according to a position of the target object in the to-be-detected image and a calibration parameter of the image acquisition device; different first head sizes have the same metric direction and different generation manners; a second head size generating module configured to determine a second head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image and a preset mapping relationship; the first head size and the second head size have the same metric direction; a fused head size determining module configured to fuse the first head size and the second head size to obtain a fused head size corresponding to the target object; and a head region determining module configured to determine a head region corresponding to the target object in the to-be-detected image according to the fused head size and the position of the target object in the to-be-detected image, and use the head region as a detection result.

[0015] According to an aspect of the present disclosure, an electronic device is provided, which comprises: a processor; a memory configured to store processor-executable instructions; and the processor is configured to invoke the instructions stored in the memory to execute the above method.

[0016] According to an aspect of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions. The computer program instructions are executed by a processor to implement the method described above.

[0017] In the embodiments of the present disclosure, the to-be-detected image can be acquired, and then a first head size corresponding to the target object in the to-be-detected image is determined according to the position of the target object in the to-be-detected image and the calibration parameters of the image acquisition device. Then, a second head size corresponding to the target object in the to-be-detected image is determined according to the position of the target object in the to-be-detected image and a preset mapping relationship. Then, the first head size and the second head size are fused to obtain a fused head size corresponding to the target object. Finally, a head region corresponding to the target object in the to-be-detected image is determined according to the fused head size and the position of the target object in the to-be-detected image, and is taken as a detection result. The embodiments of the present disclosure can fuse the first head size and the second head size determined based on different determination manners, and then obtain the fused head size, so that the parameter sensitivity of the first head size generated by different calibration parameters can be reduced, and the robustness of the detection accuracy can be improved.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure. Other features and aspects of the present disclosure will become apparent according to the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are incorporated into the specification and constitute part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the technical solutions of the present disclosure together with the specification.

[0020] Figure 1 A flowchart of a detection method according to an embodiment of the present disclosure is shown.

[0021] Figure 2 A flowchart of a detection method according to an embodiment of the present disclosure is shown.

[0022] Figure 3 A block diagram of a detection device according to an embodiment of the present disclosure is shown.

[0023] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0024] Various exemplary embodiments, features, and aspects of the present disclosure will be described below with reference to the accompanying drawings. The same or similar components are denoted by the same reference numerals throughout the drawings. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0025] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.

[0026] The term "and / or" used herein only means an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0027] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main idea of the present disclosure.

[0028] In the related art, the human head detection technology is usually executed through a machine learning model, such as a developer first annotates a head region in a training image with an annotation box, and then trains the machine learning model based on the training image and the corresponding annotation box, so that the trained machine learning model can recognize the head region in the to-be-detected image. However, this can cause the following problems: 1. The annotation box is usually annotated by artificial, which has high annotation cost and poor annotation efficiency. 2. In the case of a crowded crowd, the annotation boxes are easy to overlap, which is not conducive to visual observation by artificial, and is easy to cause artificial annotation error. 3. This is easy to be affected by the perspective effect, that is, it is easy to cause the detection accuracy to decrease. 4. The generalization ability of the machine learning model is affected by different shooting scenes, and the detection accuracy of the machine learning model is uncontrollable for different shooting scenes.

[0029] Therefore, the detection method provided in the embodiments of the present disclosure can obtain a to-be-detected image, determine a first head size corresponding to a target object in the to-be-detected image according to the position of the target object in the to-be-detected image and the calibration parameters of the image acquisition device, determine a second head size corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image and a preset mapping relationship, fuse the first head size and the second head size to obtain a fused head size corresponding to the target object, and finally determine a head region corresponding to the target object in the to-be-detected image according to the fused head size and the position of the target object in the to-be-detected image, and take the head region as a detection result. The embodiments of the present disclosure can fuse the first head size and the second head size determined based on different determination manners to obtain a fused head size, so that the parameter sensitivity of the first head size generated based on different calibration parameters can be reduced, and the robustness of the detection accuracy can be improved.

[0030] In a possible implementation, the detection method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, or the like. The method can be implemented by a processor invoking computer-readable instructions stored in a memory. Alternatively, the method can be executed by a server. In combination with an actual application scenario, the electronic device can collect a to-be-detected image through an image acquisition device connected thereto, and then automatically and accurately determine a head region in the to-be-detected image through the detection method described above. The electronic device can also generate a bounding box for the head region to facilitate user observation.

[0031] Referring to Figure 1 as shown in the accompanying drawings, Figure 1 a flowchart of a detection method provided by the embodiments of the present disclosure is shown, as Figure 1 as shown, the detection method includes the following steps:

[0032] In step S100, a to-be-detected image is obtained. The to-be-detected image is obtained through an image acquisition device. For example, the electronic device can be connected to the image acquisition device in a wired or wireless manner, so that the electronic device can obtain the to-be-detected image through the image acquisition device. The image acquisition device can include a visible light camera, a multi-view camera, or the like. The developer can flexibly set the image acquisition device according to actual needs.

[0033] At step S200, a first head size of the target object in the image to be detected is determined according to a position of the target object in the image to be detected and a calibration parameter of the image acquisition device. The target object can be any object that a developer wants to detect, such as a pedestrian, an animal, a specific person, a specific animal, and the like, which is not limited in the embodiments of the present disclosure. For example, the position can be represented as a pixel coordinate in a pixel coordinate system in the related art.

[0034] In a possible implementation, step S200 can include determining a head center point of the target object in the image to be detected. Then, the first head size of the target object in the image to be detected is determined according to a pixel coordinate corresponding to the head center point and the calibration parameter of the image acquisition device. The pixel coordinate corresponding to the head center point is used to represent the position of the head center point in the image to be detected. For example, the head center point can be obtained by using a trained machine learning model. For example, a training object in a training image can be manually labeled. The object category of the training object is the same as that of the target object, for example, both are pedestrians, both are animals, both are specific persons, and the like, so that the trained machine learning model can determine the target object in the image to be detected. For example, the head center point of the training object in the training image can be labeled, in other words, the head center point of the target object can be recognized by the trained machine learning model based on the head center point of the training object. For example, after the machine learning model is trained, the image to be detected can be input into the machine learning model to obtain position information of the head center point of the target object in the image to be detected. In an example, the position information of the head center point can be represented as a pixel coordinate in a pixel coordinate system in the related art. For example, the calibration parameter can include a shooting angle of the image acquisition device, a shooting height of the image acquisition device, and the like. The first head size can be represented based on a pixel coordinate difference of the pixel coordinate system.

[0035] Referring to Figure 2 illustrated, Figure 2 A flowchart of a detection method is shown, and in a possible implementation, step S200 can include:

[0036] At step S210, a conversion relationship between the pixel coordinates in the to-be-detected image and the world coordinates in the world coordinate system is determined according to the calibration parameters corresponding to the image acquisition device. In one example, the conversion relationship can be determined in the following manner: in the case where the calibration parameters of the image acquisition device are unknown (for example, when the image acquisition device is initially installed at a position, and the external parameters have not been calibrated according to the measurement standard in the world coordinate system), the external parameters corresponding to the image acquisition device can be determined according to the device angle of the image acquisition device. For example, the device angle can be represented by the camera pose angle (for example, the pitch angle, the yaw angle, and the roll angle) in the related art. The external parameters can include a rotation matrix. Then, the conversion relationship between the pixel coordinates in the to-be-detected image and the world coordinates in the world coordinate system is determined according to the calibration parameters corresponding to the image acquisition device. For example, there is a mapping relationship between the world coordinate system and the pixel coordinate system corresponding to the to-be-detected image. The world coordinates in the world coordinate system (which can be calibrated by the calibration algorithm in the related art) represent the position of the target object in the objective world, and the pixel coordinate system (which can be established with the top left corner of the captured image as the origin) represents the position of the target object in the captured image.

[0037] For example, the rotation matrix R of the image acquisition device can be determined according to the pitch angle θ. For example, the relationship between the two can be set as follows:

[0038]

[0039] On this basis, according to the imaging model of the image acquisition device in the related art, the corresponding relationship between the world coordinates (X, Y, Z) of any spatial point in the world coordinate system and the pixel coordinates (u, v, 1) (herein referred to as homogeneous coordinates) can be obtained:

[0040]

[0041] wherein Z' represents the value of Z in the world coordinates (X, Y, Z) of the spatial point mapped in the camera coordinate system in the related art, which can be used to represent the depth value in the camera coordinate system. u represents the coordinate information of the spatial point in the vertical direction in the pixel coordinate system, and v represents the coordinate information of the spatial point in the horizontal direction in the pixel coordinate system. is the intrinsic matrix of the image acquisition device, and the specific generation manner can refer to the related art. x y represents the length of the focal length in the x and y directions (the unit can be the number of pixels), the x direction is the vertical direction, and the y direction is the horizontal direction. x y ​​R is used to represent the offset of the principal point in the camera coordinate system (e.g. the top-left corner of the image) relative to the principal point in the pixel coordinate system in the x, y direction. T R is used to represent the transpose matrix of R. X, Y, Z are used to represent the coordinate information of the space point in the vertical direction, the horizontal direction, and the depth direction in the world coordinate system.

[0042] Through simplification, the following equation can be obtained, and can be used as the conversion relationship described above:

[0043] And

[0044] In step S220, a first head size corresponding to the target object in the image to be detected is determined according to the conversion relationship, the preset human body parameter, and the pixel coordinate corresponding to the head center point. The preset human body parameter is measured based on the measurement standard of the world coordinate system. For example, in step S220, a first world sub-coordinate of the head center point in the world coordinate system can be determined according to the installation height of the image acquisition device and the preset human body parameter. The first world sub-coordinate is in the same measurement direction as the first head size. For example, taking the space point as the head center point of the target object described above, the first head size as the height size, and a1 as the coordinate of the vertical axis (used to measure the height) as an example, if the coordinates of the head center point in the world coordinate system are (a1, b1, z1), and the coordinates in the pixel coordinate system are (u A ,v A ), the first world sub-coordinate can be represented as a1, and if the first head size is the width size and b1 is the coordinate of the horizontal axis (used to measure the width), the first world sub-coordinate can be represented as b1. In other words, the measurement direction of the first world sub-coordinate is the same as that of the first head size. In combination with the body height M, the head height m, and the installation height h of the image acquisition device in the preset human body parameter, the first world sub-coordinate can be represented as:

[0045]

[0046] Then, a second world sub-coordinate of the head center point in the world coordinate system is determined according to the conversion relationship, the pixel coordinate corresponding to the head center point, and the first world sub-coordinate. The measurement direction of the second world sub-coordinate is different from that of the first world sub-coordinate. For example, the second world sub-coordinate can be represented as z1 in the above description, for example, as the coordinate of the depth axis (used to measure the depth) in the world coordinate system. Continuing to refer to the above example, the second world sub-coordinate can be represented as:

[0047]

[0048] According to the first world sub-coordinate, the preset human body parameter, two vertices of the head region corresponding to the head center point are determined, and the world coordinates in the world coordinate system are determined. The world coordinates of the two vertices of the head region are different in the metric direction of the first world sub-coordinate. For example, if the first metric size is height, there is a height difference between the two vertices of the head region. In other words, in this case, the two vertices of the head region can be selected as: the left upper corner vertex and the left lower corner vertex of the head region, or the right upper corner vertex and the left lower corner vertex, etc. Here, the left upper corner vertex and the right lower corner vertex are taken as an example: the world coordinates of the left upper corner vertex can be represented as: The pixel coordinates can be represented as: The world coordinates of the right lower corner vertex can be represented as: The pixel coordinates can be represented as Wherein, n is used to represent the head width.

[0049] Then, according to the world coordinates of the two vertices of the head region, the second world sub-coordinate, and the conversion relationship, the pixel coordinates corresponding to the two vertices of the head region are determined. For example, the world coordinates and pixel coordinates of the left upper corner vertex are substituted into the conversion relationship above, and the following can be determined:

[0050]

[0051] The world coordinates and pixel coordinates of the right lower corner vertex are substituted into the conversion relationship above, and the following can be determined:

[0052]

[0053] Then, based on the offset of the two vertices of the head region in the metric direction of the first world sub-coordinate, the first head size corresponding to the target object in the to-be-detected image is determined. The offset is determined based on the pixel coordinates corresponding to the two vertices of the head region. For example, if Δu A represents the first head size, it can be represented as:

[0054]

[0055] Substituting the above parameters into the formula, the first head size can be obtained. For example, the generation method of the above first head size is sensitive to the pose information (such as shooting angle, installation height, etc.) of the image acquisition device, so the detection method provided by the embodiment of the present disclosure can be used to fuse the first head sizes generated by different methods to realize a detection process with higher robustness, thereby improving the adaptability to different environments and having strong practicality.

[0056] In a possible implementation, the determining of the first head size of the target object in the image to be detected based on the offset of the two vertices of the head region in the first world sub-coordinate measurement direction can include: determining pixel coordinates corresponding to the two vertices of the head region, determining the offset in the first world sub-coordinate measurement direction, and then, in a case where the offset is greater than a preset offset, taking the offset as the first head size of the target object in the image to be detected. For example, the preset offset can be 0 in a case where the offset between the top-left vertex and the bottom-right vertex is taken as an example. Since the top-left vertex is higher than the bottom-right vertex in the first world sub-coordinate measurement direction, the offset between the two vertices should be greater than 0. If the offset is less than or equal to 0, it can be determined that an abnormal situation occurs in the detection method, and the detection of the target object can be abandoned, or an abnormal prompt can be generated, which is not limited in the embodiments of the present disclosure. The first head size and the second head size are in the same measurement direction Therefore, when the two are subtracted, Δu A should be greater than 0. If it is less than or equal to 0, it can be determined that an abnormal situation occurs in the detection method, and the detection of the target object can be abandoned, or an abnormal prompt can be generated, which is not limited in the embodiments of the present disclosure.

[0057] Referring back to FIG. 3, Figure 1 the second head size of the target object in the image to be detected based on the position of the target object in the image to be detected and a preset mapping relationship. The first head size and the second head size are in the same measurement direction

[0058] In a possible implementation, the step S300 can include: determining a head center point corresponding to the target object in the image to be detected, and then determining the second head size of the target object in the image to be detected based on the pixel coordinates corresponding to the head center point and the preset mapping relationship.

[0059] In one example, the preset mapping relationship can be generated by obtaining a training image with a training object. The training object is of the same object category as the target object. The training image and the training object used in the mapping relationship establishment stage can be the same as or different from the training image and the training object used in the machine learning model training stage, which is not limited in the embodiments of the present disclosure. Then, the body size corresponding to the training object is determined according to the training image. The body size can include the body height and the body width of the training object. The body frame data (such as the frame top point position, the frame height, and the frame width) corresponding to the training object can be determined by a machine learning model. The height and the width of the body frame data can be used as the body height and the body width of the training object. For example, the machine learning model can be trained by a training image labeled with a body frame, so that the trained machine learning model can output the body frame data corresponding to the to-be-detected image. The electronic device determines the head size corresponding to the training object according to the body size. For example, the head size can be obtained by multiplying the body height by a preset ratio. For example, if the body height is ΔU and the ratio is one-seventh, the head size can be ΔU / 7. Finally, the preset mapping relationship is determined according to the head size corresponding to the training object and the position of the training object in the training image. For example, the pixel coordinates of the head center point of a plurality of training objects in the training image and the head size corresponding to the same training object can be used to perform polynomial fitting in related technologies to obtain a fitting curve between the two. Subsequently, the second head size corresponding to the target object can be obtained according to the pixel coordinates of the head center point of the target object and the fitting curve.

[0060] With reference to Figure 1 , step S400, the first head size and the second head size are fused to obtain a fused head size corresponding to the target object. For example, the first head size and the second head size can be averaged or weightedly averaged, and the average value can be used as the fused head size, which is not limited in the embodiments of the present disclosure.

[0061] In one possible implementation, step S400 can include: fusing the first head size, the weight corresponding to the first head size, the second head size, and the weight corresponding to the second head size to obtain the fused head size corresponding to the target object. The weight is related to the position of the target object in the to-be-detected image and the difference between the first head size and the second head size. For example, the fused head size can be determined by the following formula:

[0062] h A =α·h c +(1-α)·h f .

[0063] Among them, h A The first head size is used to represent the fusion head size, α is used to represent the weight, and the sum of the weights of the first head size and the second head size should be 1. c h f These are used to represent the first head size and the second head size, respectively. For example, combining the methods for determining the first head size and the second head size described above, then h... c h can be used to represent the first head size of the target object determined based on the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device. f It can be used to represent the second head size of the target object, determined by the pixel coordinates corresponding to the head center point and a preset mapping relationship.

[0064] For example, a fitting formula for the weight α can be obtained using polynomial fitting techniques in related technologies, as follows:

[0065]

[0066] Experimental data testing shows that when the weight α is constructed based on the above formula, it can improve h. A Regarding the accuracy, the generation method of weight α can also be set according to the actual required parameters in this embodiment, and this embodiment does not impose any restrictions here.

[0067] Continue reading Figure 1 As shown, in step S500, based on the fused head size and the position of the target object in the image to be detected, the head region corresponding to the target object in the image to be detected is determined and used as the detection result. For example, a corresponding bounding box can be generated based on the aforementioned head region to visualize the detection result.

[0068] Continue reading Figure 2 ,like Figure 2 As shown, in one possible implementation, step S500 may include:

[0069] Step S510: Based on the fused head size and the preset proportional relationship between the fused head size and the third head size, determine the third head size corresponding to the target object in the image to be detected. Here, the fused head size and the third head size are either the width or height of the head region. For example, the fused head size can be either the head height or the head width, and the corresponding second head size can be either the head width or the head height.

[0070] In one example, the second head size can be determined based on a preset ratio of head height to width. For example, if the second head size is denoted as Δv A and the ratio is two-thirds, the preset ratio relationship can be determined as follows:

[0071]

[0072] In step S520, a head region corresponding to the target object in the to-be-detected image is determined according to the position of the target object in the to-be-detected image, the fused head size, and the third head size. For example, the position of the target object in the to-be-detected image can be represented by the pixel coordinates of the head center point of the target object as described above. The range of the head region can be determined by dividing the fused head size and the third head size by 2 and adding the pixel coordinates of the head center point in the measurement direction.

[0073] In one possible implementation, step S100 can include acquiring at least two to-be-detected images with the target object by at least two image acquisition devices. For example, the shooting planes of the at least two image acquisition devices can be the same, that is, the perspective relationship of the target object in the shooting angles of the multiple image acquisition devices is the same (that is, the size of the target object captured is the same) (for example, the pitch angles of the multiple image acquisition devices can be set to be the same). In this case, step S500 can include fusing the fused head sizes obtained by each image acquisition device to obtain a multi-device fusion result, and then determining the head region corresponding to the target object in each to-be-detected image according to the multi-device fusion result and the position of the target object in the to-be-detected image. For example, the fused head size corresponding to each to-be-detected image can be directly averaged or weightedly averaged to obtain a mean value. Finally, the head region in each to-be-detected image is determined according to the head center point of the target object in the different to-be-detected images and the mean value. The embodiments of the present disclosure can detect the head region of the target object by multiple image acquisition devices to improve the detection accuracy of the head region, and can be applied to application scenarios such as human pose reconstruction that require accurate acquisition of the head region.

[0074] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length, the present disclosure will not be repeated. Those skilled in the art can understand that the specific execution order of each step in the above-mentioned method should be determined according to its function and possible internal logic.

[0075] In addition, this disclosure also provides a detection device, electronic device, computer-readable storage medium, and program, all of which can be used to implement any of the detection methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding records in the method section and will not be repeated here.

[0076] Figure 3 A block diagram of a detection apparatus provided according to an embodiment of the present disclosure is shown, such as Figure 3 As shown, the detection device 100 includes: a target image acquisition module 110 for acquiring a target image; wherein the target image is acquired by an image acquisition device. A first head size generation module 120 for determining a first head size corresponding to the target object in the target image based on the position of the target object in the target image and the calibration parameters of the image acquisition device; wherein different first head sizes have the same measurement direction but different generation methods. A second head size generation module 130 for determining a second head size corresponding to the target object in the target image based on the position of the target object in the target image and a preset mapping relationship; wherein the first head size and the second head size have the same measurement direction. A fused head size determination module 140 for fusing the first head size and the second head size to obtain a fused head size corresponding to the target object. A head region determination module 150 for determining the head region corresponding to the target object in the target image based on the fused head size and the position of the target object in the target image, and using it as the detection result.

[0077] In one possible implementation, determining the first head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected includes: determining the head center point corresponding to the target object in the image to be detected; determining the first head size corresponding to the target object in the image to be detected based on the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device; wherein, the pixel coordinates corresponding to the head center point are used to represent the position of the head center point in the image to be detected.

[0078] In one possible implementation, determining the second head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected and a preset mapping relationship includes: determining the head center point corresponding to the target image in the image to be detected; and determining the second head size corresponding to the target object in the image to be detected based on the pixel coordinates corresponding to the head center point and a preset mapping relationship.

[0079] In a possible implementation, the detection apparatus further includes a mapping relationship determination module configured to perform the following steps: obtaining a training image with a training object; wherein the training object is of the same object category as the target object; determining a body size corresponding to the training object according to the training image; determining a head size corresponding to the training object according to the body size; and determining the preset mapping relationship according to the head size corresponding to the training object and pixel coordinates corresponding to the training object; wherein the pixel coordinates corresponding to the training object are used to represent a position of the training object in the training image.

[0080] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image according to the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device includes: determining a conversion relationship between pixel coordinates in the to-be-detected image and world coordinates in a world coordinate system according to the calibration parameters corresponding to the image acquisition device; and determining the first head size corresponding to the target object in the to-be-detected image according to the conversion relationship, preset human body parameters, and the pixel coordinates corresponding to the head center point; wherein the preset human body parameters are measured based on a measurement standard of the world coordinate system.

[0081] In a possible implementation, the determining the first head size corresponding to the target object in the to-be-detected image according to the conversion relationship, the preset human body parameters, and the pixel coordinates corresponding to the head center point includes: determining a first world sub-coordinate of the head center point in the world coordinate system according to an installation height of the image acquisition device and the preset human body parameters; wherein the first world sub-coordinate is of the same measurement direction as the first head size; determining a second world sub-coordinate of the head center point in the world coordinate system according to the conversion relationship, the pixel coordinates corresponding to the head center point, and the first world sub-coordinate; wherein the second world sub-coordinate is of a different measurement direction from the first world sub-coordinate; determining world coordinates of two vertices of a head region corresponding to the head center point in the world coordinate system according to the first world sub-coordinate and the preset human body parameters; wherein the world coordinates of the two vertices of the head region are different in a value in the measurement direction of the first world sub-coordinate; determining pixel coordinates corresponding to the two vertices of the head region according to the world coordinates of the two vertices of the head region, the second world sub-coordinate, and the conversion relationship; and determining the first head size corresponding to the target object in the to-be-detected image based on an offset of the two vertices of the head region in the measurement direction of the first world sub-coordinate; wherein the offset is determined based on the pixel coordinates corresponding to the two vertices of the head region.

[0082] In a possible implementation, the first head size of the target object in the to-be-detected image is determined based on the offset of the two vertexes of the head region in the first world sub-coordinate measurement direction, and the method comprises the following steps: determining the pixel coordinates corresponding to the two vertexes of the head region, and the offset in the first world sub-coordinate measurement direction; in a case where the offset is greater than a preset offset, taking the offset as the first head size of the target object in the to-be-detected image.

[0083] In a possible implementation, the head region corresponding to the target object in the to-be-detected image is determined according to the fused head size and the position of the target object in the to-be-detected image, and the method comprises the following steps: determining a third head size of the target object in the to-be-detected image according to the fused head size and a preset proportional relationship between the fused head size and the third head size; wherein the fused head size and the third head size are one of the width and the height of the head region; and determining the head region corresponding to the target object in the to-be-detected image according to the position of the target object in the to-be-detected image, the fused head size and the third head size.

[0084] In a possible implementation, the first head size and the second head size are fused to obtain the fused head size corresponding to the target object, and the method comprises the following steps: fusing the first head size, a weight corresponding to the first head size, the second head size and a weight corresponding to the second head size to obtain the fused head size corresponding to the target object; wherein the weights are related to the position of the target object in the to-be-detected image and the difference between the first head size and the second head size.

[0085] In a possible implementation, the to-be-detected image is obtained by at least two image acquisition devices, and the head region corresponding to the target object in the to-be-detected image is determined according to the fused head size and the position of the target object in the to-be-detected image, and the method comprises the following steps: fusing the fused head size obtained by each image acquisition device to obtain a multi-device fusion result; and determining the head region corresponding to the target object in each to-be-detected image according to the multi-device fusion result and the position of the target object in the to-be-detected image.

[0086] The method has specific technical correlation with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing data storage, reducing data transmission, improving hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in accordance with the natural law.

[0087] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and specific implementations can refer to the descriptions of the above method embodiments. For brevity, they will not be repeated here.

[0088] The embodiments of the present disclosure also propose a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the above method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0089] The embodiments of the present disclosure also propose an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to execute the above method.

[0090] The embodiments of the present disclosure also provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in the processor of the electronic device, the processor in the electronic device executes the above method.

[0091] The electronic device can be provided as a server or other form of device.

[0092] Figure 4 A block diagram of an electronic device 1900 is shown according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 4 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932, for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.

[0093] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Microsoft Windows Server TM ), Apple's graphical user interface-based operating system (Mac OSX TM ), a multi-user multi-processing computer operating system (Unix TM ), a free and open-source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ), or the like.

[0094] In an exemplary embodiment, there is also provided a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above-described method.

[0095] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0096] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0097] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0098] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0099] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0100] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer readable storage medium having no data signals on it. The instructions can be executed by one or more processors of a computer or other programmable data processing apparatus to produce a computer implemented process such that the instructions, which execute via the one or more processors of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0101] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer readable storage medium having no data signals on it. The instructions can be executed by one or more processors of a computer or other programmable data processing apparatus to produce a computer implemented process such that the instructions, which execute via the one or more processors of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0102] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0103] The computer program product can be embodied by a hardware, a software or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) or the like.

[0104] The above description of the various embodiments is intended to be illustrative and not restrictive. Many other embodiments will be readily apparent to those skilled in the art upon reviewing the above description, and variations from the specific embodiments described herein will be readily apparent to those with ordinary skill in the art. By introducing elements expressed in the dependent claims, alternative embodiments can be implemented.

[0105] Those skilled in the art can understand that, in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0106] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and the personal information will be collected. If the person voluntarily enters the collection range, it is considered to agree to collect the personal information. Or, on the device for processing personal information, the personal information processing rules are informed by using obvious signs / information, and the personal authorization is obtained by means of pop-up information or asking the person to upload his / her personal information. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type, etc.

[0107] The above has described various embodiments of the present disclosure, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical application, or improvement of technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A detection method, characterized in that, The detection method includes: Acquire the image to be detected; wherein the image to be detected is acquired by an image acquisition device; Based on the position of the target object in the image to be detected and the calibration parameters of the image acquisition device, the first head size corresponding to the target object in the image to be detected is determined; Based on the position of the target object in the image to be detected and a preset mapping relationship, the second head size corresponding to the target object in the image to be detected is determined; wherein, the measurement direction of the first head size and the second head size is the same, and the mapping relationship is a representation of the relationship between the head size corresponding to the object and the position of the object in the image; The first head size and the second head size are merged to obtain the merged head size corresponding to the target object; Based on the fused head size and the position of the target object in the image to be detected, the head region corresponding to the target object in the image to be detected is determined and used as the detection result; The step of determining the head region corresponding to the target object in the image to be detected based on the fused head size and the position of the target object in the image to be detected includes: Based on the fused head size and the preset proportional relationship between the fused head size and the third head size, the third head size corresponding to the target object in the image to be detected is determined; wherein, the fused head size and the third head size are respectively one of the width and height of the head region; Based on the position of the target object in the image to be detected, the fused head size, and the third head size, the head region corresponding to the target object in the image to be detected is determined.

2. The detection method as described in claim 1, characterized in that, The step of determining the first head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected includes: Determine the center point of the head of the target object in the image to be detected; Based on the pixel coordinates corresponding to the head center point and the calibration parameters of the image acquisition device, the first head size corresponding to the target object in the image to be detected is determined; wherein, the pixel coordinates corresponding to the head center point are used to represent the position of the head center point in the image to be detected.

3. The detection method as described in claim 1 or 2, characterized in that, The step of determining the second head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected and a preset mapping relationship includes: Determine the head center point corresponding to the target image in the image to be detected; Based on the pixel coordinates corresponding to the center point of the head and a preset mapping relationship, the second head size corresponding to the target object in the image to be detected is determined.

4. The detection method as described in claim 3, characterized in that, The detection method further includes: Acquire training images containing training objects; wherein the training objects and the target objects belong to the same object category; Based on the training images, determine the body dimensions corresponding to the training object; Based on the body dimensions, determine the head dimensions corresponding to the training subject; The preset mapping relationship is determined based on the head size and pixel coordinates of the training object; wherein the pixel coordinates of the training object are used to represent the position of the training object in the training image.

5. The detection method as described in claim 2, characterized in that, The step of determining the first head size corresponding to the target object in the image to be detected based on the pixel coordinates corresponding to the center point of the head and the calibration parameters of the image acquisition device includes: Based on the calibration parameters corresponding to the image acquisition device, determine the transformation relationship between the pixel coordinates in the image to be detected and the world coordinates in the world coordinate system; Based on the transformation relationship, preset human body parameters, and the pixel coordinates corresponding to the head center point, the first head size corresponding to the target object in the image to be detected is determined; wherein, the preset human body parameters are measured based on the measurement standard of the world coordinate system.

6. The detection method as described in claim 5, characterized in that, The step of determining the first head size corresponding to the target object in the image to be detected based on the transformation relationship, preset human body parameters, and the pixel coordinates corresponding to the head center point includes: Based on the installation height of the image acquisition device and preset human body parameters, the first world sub-coordinates of the head center point in the world coordinate system are determined; wherein, the first world sub-coordinates are in the same direction as the measurement of the first head size; Based on the transformation relationship, the pixel coordinates corresponding to the head center point, and the first world sub-coordinates, the second world sub-coordinates of the head center point in the world coordinate system are determined; wherein, the measurement direction of the second world sub-coordinates is different from that of the first world sub-coordinates; Based on the first world sub-coordinates and the preset human body parameters, determine the world coordinates of the two vertices of the head region corresponding to the head center point in the world coordinate system; wherein, the world coordinates of the two vertices of the head region have different values ​​in the measurement direction of the first world sub-coordinates. The pixel coordinates corresponding to the two vertices of the head region are determined based on the world coordinates of the two vertices of the head region, the second world sub-coordinates, and the transformation relationship. The first head size of the target object in the image to be detected is determined based on the offset of the two vertices of the head region in the first world sub-coordinate measurement direction; wherein the offset is determined based on the pixel coordinates of the two vertices of the head region.

7. The detection method as described in claim 6, characterized in that, Determining the first head size corresponding to the target object in the image to be detected based on the offset of the two vertices of the head region in the first world sub-coordinate metric direction includes: Determine the pixel coordinates corresponding to the two vertices of the head region, and the offset in the first world sub-coordinate measurement direction; If the offset is greater than a preset offset, the offset is used as the first head size of the target object in the image to be detected.

8. The detection method according to any one of claims 1 or 2, characterized in that, The step of fusing the first head size and the second head size to obtain the fused head size corresponding to the target object includes: Based on the first head size, the weight corresponding to the first head size, the second head size, and the weight corresponding to the second head size, the first head size and the second head size are fused to obtain the fused head size corresponding to the target object; wherein, the weight is related to the position of the target object in the image to be detected and the difference between the first head size and the second head size.

9. The detection method according to any one of claims 1 or 2, characterized in that, The acquisition of the image to be detected includes: acquiring at least two images to be detected containing the target object through at least two image acquisition devices; The step of determining the head region corresponding to the target object in the image to be detected based on the fused head size and the position of the target object in the image to be detected includes: The fusion head dimensions obtained from each image acquisition device are fused together to obtain the multi-device fusion result; Based on the multi-device fusion result and the position of the target object in the image to be detected, the head region corresponding to the target object in each image to be detected is determined.

10. A detection device, characterized in that, The detection device includes: The image to be detected module is used to acquire the image to be detected; wherein, the image to be detected is acquired by an image acquisition device; The first head size generation module is used to determine the first head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected and the calibration parameters of the image acquisition device; wherein, different first head sizes have the same measurement direction but different generation methods; The second head size generation module is used to determine the second head size corresponding to the target object in the image to be detected based on the position of the target object in the image to be detected and a preset mapping relationship; wherein, the first head size and the second head size are measured in the same direction, and the mapping relationship is a representation of the relationship between the object position and the object head size; A fused head size determination module is used to fuse the first head size and the second head size to obtain the fused head size corresponding to the target object; The head region determination module is used to determine the head region corresponding to the target object in the image to be detected based on the fused head size and the position of the target object in the image to be detected, and use it as the detection result; The head region determination module is further used for: Based on the fused head size and the preset ratio between the fused head size and the third head size, the third head size corresponding to the target object in the image to be detected is determined; wherein, the fused head size and the third head size are respectively one of the width and height of the head region; based on the position of the target object in the image to be detected, the fused head size, and the third head size, the head region corresponding to the target object in the image to be detected is determined.

11. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the detection method according to any one of claims 1 to 9.

12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Key point detection method and device, electronic equipment and storage medium

    CN111243011A

  • Sample image generation method of pedestrian head image classifier and corresponding training method

    CN112257797A