Target positioning method, track matching method, electronic equipment and storage medium

By acquiring and correcting the coordinate information of the detection box, the height of the target object is predicted and corrected, which solves the problem of low detection box accuracy when the target object is occluded, and achieves higher accuracy in target localization and trajectory matching.

CN121582328APending Publication Date: 2026-02-27ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511537076.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

When the target object in an image is occluded, the detection box obtained by traditional object detection methods has low accuracy, which affects the downstream processing results. This is especially true in traffic scenes where vehicle targets are occluded, leading to a decrease in the accuracy of the detection box.

Method used

By acquiring the coordinates of the initial detection box, the target object's contact point coordinates, and the image acquisition device coordinates, the initial height of the target object is predicted. Based on the initial height, the detection box is corrected to generate a corrected detection box, thereby improving the accuracy of the detection box.

Benefits of technology

When the target object is occluded, the detection bounding box generated through the correction process has high accuracy and can more accurately represent the 2D or 3D localization results of the target object, thus improving the accuracy of downstream processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582328A_ABST
    Figure CN121582328A_ABST
Patent Text Reader

Abstract

The invention discloses a target positioning method, a track matching method, electronic equipment and a storage medium. The target positioning method comprises the following steps: in response to shielding of a target object in an initial detection frame of a current frame image, obtaining coordinate information of the initial detection frame, coordinate information of a contact point of the target object in the initial detection frame and coordinate information of image acquisition equipment; predicting the initial height of the target object according to the coordinate information of the initial detection frame, the coordinate information of the attachment point of the target object and the coordinate information of the image acquisition equipment; the initial detection frame is corrected according to the initial height, a corrected detection frame is obtained, the detection frame represents a 2D positioning result and / or a 3D positioning result of the target object in the current frame image, and the detection frame comprises the initial detection frame and the corrected detection frame. According to the scheme, the accuracy of the determined detection frame can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, and in particular, to a target positioning method, a trajectory matching method, an electronic device, and a storage medium. BACKGROUND

[0002] At present, the detection box obtained by target detection on the image collected by the target object can provide the position information of each target in the image. The detection box in the image is widely applied in downstream processing, such as target tracking algorithm, collision prediction between different target objects, and the like. However, the target detection depends on the completeness of the target object in the image. If the target object is partially occluded in the image, the accuracy of the detection box obtained by the target detection is low. Specifically, in a traffic scene, the traffic environment is complex and changeable. The vehicle target is often occluded due to dense vehicles, interlaced pedestrians, or weather factors, thereby causing the accuracy of the detection box obtained by the traditional target detection method to be low, and further affecting the result obtained by the detection box in the image in downstream processing.

[0003] Therefore, there is an urgent need for a target positioning method. SUMMARY

[0004] The present application provides at least a target positioning method, a trajectory matching method, an electronic device, and a storage medium.

[0005] The present application provides a target positioning method. The target positioning method comprises: in response to the existence of occlusion of a target object in an initial detection box of a current frame image, obtaining coordinate information of the initial detection box, patch coordinate information of the target object in the initial detection box, and coordinate information of an image acquisition device; predicting an initial height of the target object according to the coordinate information of the initial detection box, the patch coordinate information of the target object, and the coordinate information of the image acquisition device; and performing correction processing on the initial detection box according to the initial height to obtain a corrected detection box, wherein the detection box represents a 2D positioning result and / or a 3D positioning result of the target object in the current frame image, and the detection box comprises the initial detection box and the corrected detection box.

[0006] The present application provides a trajectory matching method. The trajectory matching method comprises: in response to the existence of occlusion of a target object in an initial detection box of a current frame image, obtaining a corrected detection box corresponding to the initial detection box, the corrected detection box being obtained by the above-mentioned target positioning method; and performing tracking matching based on the corrected detection box to obtain trajectory information of the target object.

[0007] The application provides a target positioning device, comprising: an acquisition module, a prediction module and a correction module; the acquisition module is used for acquiring coordinate information of an initial detection frame, sticking point coordinate information of a target object in the initial detection frame and coordinate information of an image acquisition device in response to the existence of occlusion of the target object in the initial detection frame of a current frame image; the prediction module is used for predicting an initial height of the target object according to the coordinate information of the initial detection frame, the sticking point coordinate information of the target object and the coordinate information of the image acquisition device; and the correction module is used for performing correction processing on the initial detection frame according to the initial height to obtain a corrected detection frame, wherein the detection frame represents a 2D positioning result and / or a 3D positioning result of the target object in the current frame image, and the detection frame comprises the initial detection frame and the corrected detection frame.

[0008] The application provides a trajectory matching device, comprising: a preset acquisition module and a tracking matching module; the preset acquisition module is used for acquiring a corrected detection frame corresponding to an initial detection frame of a current frame image in response to the existence of occlusion of a target object in the initial detection frame, the corrected detection frame being obtained by the above target positioning method; and the tracking matching module is used for performing tracking matching based on the corrected detection frame to obtain trajectory information of the target object.

[0009] The application provides an electronic device, comprising a memory and a processor, the processor being used for executing program instructions stored in the memory to realize the above target positioning method and / or the above trajectory matching method.

[0010] The application provides a computer readable storage medium, the program instructions being stored on the computer readable storage medium, the program instructions being executed by a processor to realize the above target positioning method and / or the above trajectory matching method.

[0011] The above scheme is used for predicting an initial height of a target object according to acquired coordinate information of an initial detection frame, sticking point coordinate information of the target object in the initial detection frame and coordinate information of an image acquisition device in the case that the target object exists occlusion in the initial detection frame of a current frame image, so that the predicted initial height has high accuracy, and then performing correction processing on the initial detection frame according to the initial height to obtain a corrected detection frame with high accuracy.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the application. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the application and, together with the description, serve to explain the principles of the application.

[0014] Figure 1is a flowchart of an exemplary embodiment of the target positioning method of the present application; Figure 2 is Figure 1 is a sub-flowchart of step S12 in the method of Figure 3 is Figure 1 is a sub-flowchart of step S13 in the method of Figure 4 is Figure 3 is a sub-flowchart of step S31 in the method of Figure 5 is Figure 3 is a sub-flowchart of step S32 in the method of Figure 6 is a flowchart of an exemplary embodiment of the trajectory matching method of the present application; Figure 7a is a framework diagram of an exemplary embodiment of the target positioning method of the present application; Figure 7b is a geometric relationship diagram in an exemplary embodiment of the target positioning method of the present application; Figure 7c is an initial detection frame diagram in an exemplary embodiment of the target positioning method of the present application; Figure 7d is a corrected detection frame diagram in an exemplary embodiment of the target positioning method of the present application; Figure 8 is a structural diagram of an embodiment of the target positioning apparatus of the present application; Figure 9 is a structural diagram of an embodiment of the trajectory matching apparatus of the present application; Figure 10 is a structural diagram of an embodiment of the electronic device of the present application; Figure 11 is a structural diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0016] In the following description, specific details are set forth in order to provide a thorough understanding of the present application. The present application may, however, be practiced without these details. In other instances, well-known methods, structures and techniques have not been described in detail in order to avoid obscuring the present application.

[0017] The term "and / or", as used herein, merely describes association between associated objects, and can indicate that three cases can exist, for example, A and / or B can indicate that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally indicates that the front and rear associated objects are in an "or" relationship. In addition, "multiple" herein indicates two or more than two. In addition, the term "at least one" herein indicates any one of multiple or any combination of at least two of multiple, for example, at least one of A, B and C can indicate any one or more elements selected from the set consisting of A, B and C.

[0018] The present application provides some target positioning methods and target positioning devices. The application scenarios of the target positioning method include but are not limited to target positioning scenarios and target tracking matching scenarios. The execution subject of the target positioning method can be a target positioning device, for example, the target positioning device can be arranged in a terminal device or a server or other processing device, wherein the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementation manners, the target positioning method can be realized by calling computer readable instructions stored in a memory by a processor.

[0019] Please refer to Figure 1 , Figure 1 is a flowchart of an exemplary embodiment of the target positioning method of the present application. Specifically, the target positioning method can include the following steps: Step S11: In response to the existence of occlusion in the target object in the initial detection frame of the current frame image, obtaining coordinate information of the initial detection frame, patch coordinate information of the target object in the initial detection frame and coordinate information of the image acquisition device.

[0020] The current frame image can be an image obtained by image acquisition on the target scene at the current time. The target scene can be a traffic scene, a logistics transportation scene, etc. The present application takes the target scene as a traffic scene as an example. In some application scenarios, the current frame image can be an image frame at the current time in the video data obtained by collecting the target scene. The initial detection box represents a detection box obtained by performing a preset target detection processing on the current frame image. The preset target detection processing can be any target detection algorithm. Specifically, the preset target detection processing can be 2D detection processing and / or 3D detection processing. The initial detection box represents a 2D initial positioning result and / or a 3D initial positioning result of a target object in the current frame image. The target object can be an object in the initial detection box. For example, in the traffic scene, the target object can be a vehicle, and in the logistics transportation scene, the target object can be a delivery box waiting for transportation. The present application takes the target object as a vehicle as an example, and the subsequent description will not be repeated. The target object is related to the target scene, and the target object is an object in the target scene that needs to be detected and / or tracked and matched. The coordinate information of the initial detection box represents the coordinate information of at least one point in the initial detection box. Specifically, the coordinate information of the initial detection box includes first coordinate information of a first preset position and third coordinate information of a third preset position on the top edge of the initial detection box, and second coordinate information of a second preset position on the bottom edge of the initial detection box. Wherein, in the case that the first preset position and the third preset position are the same position, the first coordinate information and the third coordinate information are the coordinate information of the same position on the top edge of the initial detection box. The contact point coordinate information represents the coordinate information of the contact point between the target object and the road in the target scene. For example, in the case that the target object is a vehicle, the contact point coordinate information represents the coordinate information of the contact point between each tire corresponding to the target object and the road in the target scene. The image acquisition device is a device for collecting the target scene to obtain the current frame image. The coordinate information of the image acquisition device represents the coordinate information of any one preset position in the image acquisition device. For example, the coordinate information of the image acquisition device represents the coordinate information of the optical center of the image acquisition device.

[0021] It can be understood that the number of target objects can be one or more. The number of initial detection boxes can be one or more. In the case that the number of target objects and initial detection boxes is multiple, the present application performs the same processing on different target objects and different initial detection boxes to obtain the corrected detection box corresponding to each initial detection box. For the sake of brevity, the number of target objects and initial detection boxes is taken as one as an example, and the subsequent description will not be repeated.

[0022] In some application scenarios, before the step S11, in response to the presence of at least one target object in the current frame image, it is determined whether each target object in the current frame image is occluded. Specifically, for each target object, the occlusion degree of the target object in the current frame image is determined, and in response to the occlusion degree of the target object being greater than or equal to a threshold, it is determined that the target object is occluded.

[0023] In some application scenarios, the step S11 can be, in response to the presence of occlusion in the target object in the initial detection box of the current frame image, performing a preset target detection processing on the current frame image to obtain the initial detection box and the image coordinate information of the initial detection box. Wherein, the image coordinate information of the initial detection box is directly taken as the coordinate information of the initial detection box. Or, the image coordinate information of the initial detection box in the image coordinate system of the current frame image is converted to the world coordinate system to obtain the coordinate information of the initial detection box. The present application takes the coordinate information of the initial detection box as an example of the coordinate information in the world coordinate system.

[0024] In other application scenarios, the step S11 can be, in response to the presence of occlusion in the target object in the initial detection box of the current frame image, performing a preset attachment point detection processing on the current frame image to obtain the attachment point coordinate information of the target object. The preset attachment point detection processing can be to perform key point detection on the target object in the current frame image to obtain a key point detection result. The key point detection result includes the key points at the bottom of the target object in the current frame image, and the key points in the key point detection result are taken as the attachment points. For example, the attachment point coordinate information includes the coordinate information of the four attachment points at the bottom of the target object. In response to the number of key points in the key point detection result not reaching a preset number, the orientation of the target object in the current frame image and the experience size of the target object are obtained; and the attachment point coordinate information is determined based on the orientation of the target object and the experience size. Wherein, the experience size represents the shape and size formed by each attachment point at a historical time. For example, the preset number is four. Specifically, the attachment point coordinate information is determined based on the orientation of the target object and the experience size, including: in response to the number of key points in the key point detection result being equal to a first preset number, the experience size is fine-tuned according to the orientation of the target object to obtain a target attachment region, and each corner point in the target attachment region is taken as an attachment point, and the coordinate information of each corner point is taken as the attachment point coordinate information, wherein the first preset number is 0. Or, in response to the number of key points in the key point detection result being greater than the first preset number but less than a second preset number, taking each key point in the key point detection result as a reference point, and fitting the missing key points according to the orientation of the target object and the experience size. Each key point in the key point detection result and the fitted missing key points are taken as each attachment point.

[0025] In some application scenarios, the step S11 can be, in response to the target object in the initial detection frame of the current frame image being occluded, taking the coordinate information of a preset part of the image acquisition device as the coordinate information of the image acquisition device. The preset part can be the optical center part.

[0026] In some application scenarios, the coordinate information of the initial detection frame and the patch coordinate information of the target object in the initial detection frame of the current frame image are obtained by performing preset processing on the current frame image through a preset algorithm before the step S11, and are stored in a preset database; the coordinate information of the image acquisition device is also stored in the preset database. In the step S11, the coordinate information of the initial detection frame and the patch coordinate information of the target object associated with the current frame image are called from the preset database, and the coordinate information of the image acquisition device is also called.

[0027] Step S12: predicting the initial height of the target object according to the coordinate information of the initial detection frame, the patch coordinate information of the target object, and the coordinate information of the image acquisition device.

[0028] The initial height of the target object represents the estimated image height of the target object in the current frame image.

[0029] In some application scenarios, the step S12 can be inputting the coordinate information of the initial detection frame, the patch coordinate information of the target object, and the coordinate information of the image acquisition device into a height estimation model to obtain the initial height output by the height estimation model.

[0030] Step S13: performing correction processing on the initial detection frame according to the initial height to obtain a corrected detection frame.

[0031] The detection frame represents the 2D positioning result and / or the 3D positioning result of the target object in the current frame image, and the detection frame includes the initial detection frame and the corrected detection frame. The corrected detection frame represents the initial detection frame after the correction processing. The corrected detection frame represents the 2D target positioning result and / or the 3D target positioning result of the target object in the current frame image. The correction processing can be inputting the initial detection frame and the initial height into a preset correction processing module to obtain the detection frame output by the preset correction processing module, and taking the detection frame output by the preset correction processing module as the corrected detection frame.

[0032] In some application scenarios, the step S13 can be: determining the to-be-adjusted coordinate information of the target position on the bottom side of the initial detection frame according to the height of the image acquisition device from the ground, the first horizontal distance between the image acquisition device and the first preset position in the horizontal direction, the horizontal actual size of the target side of the target object on the horizontal plane, and the initial height; taking the intersection between the extension line of each frame side in the initial detection frame and the horizontal line where the to-be-adjusted coordinate information is located as the target intersection point; and generating the corrected detection frame with each target intersection point and each corner point in the top side of the frame as the reference corner point.

[0033] In some application scenarios, the step S13 can be: fine-tuning the initial height according to the occlusion degree of the target object and / or the fact that the target object is in the road region in the target scene at the current moment; determining the to-be-adjusted coordinate information of the target position on the bottom side of the initial detection frame according to the height of the image acquisition device from the ground, the first horizontal distance between the image acquisition device and the first preset position in the horizontal direction, the horizontal actual size of the target side of the target object on the horizontal plane, and the target height; taking the intersection between the extension line of each frame side in the initial detection frame and the horizontal line where the to-be-adjusted coordinate information is located as the target intersection point; and generating the corrected detection frame with each target intersection point and each corner point in the top side of the frame as the reference corner point.

[0034] The above scheme, in the case that the target object is occluded in the initial detection frame of the current frame image, predicts the initial height of the target object according to the obtained coordinate information of the initial detection frame, the sticking point coordinate information of the target object in the initial detection frame, and the coordinate information of the image acquisition device, so that the predicted initial height has high accuracy, and the corrected detection frame obtained by correcting the initial detection frame according to the initial height has high accuracy.

[0035] In some embodiments, the target positioning method further includes the following steps: in response to the fact that the target object is not occluded in the initial detection frame of the current frame image, confirming that the initial detection frame does not need to be corrected and outputting the initial detection frame as the corrected detection frame of the current frame image; or, in response to the fact that the target object is not occluded in the initial detection frame of the current frame image, obtaining the coordinate information of the initial detection frame, the sticking point coordinate information of the target object in the initial detection frame, and the coordinate information of the image acquisition device; predicting the initial height of the target object according to the coordinate information of the initial detection frame, the sticking point coordinate information of the target object, and the coordinate information of the image acquisition device, and storing the initial height for correcting the initial detection frame belonging to the same target object in the next frame image based on the initial height to obtain the corrected detection frame of the target object in the next frame image.

[0036] It can be understood that the application can record the initial height of the target object predicted according to the coordinate information of the initial detection box, the coordinate information of the sticking point of the target object and the coordinate information of the image acquisition device when the occlusion degree of the target object is small, and directly retain it as the historical height of the next frame of image, thereby improving the accuracy of the historical height corresponding to the same target object.

[0037] Please refer to Figure 2 , Figure 2 is Figure 1 a subflowchart diagram of step S12 in

[0038] In some embodiments, the coordinate information of the initial detection box includes second coordinate information of a second preset position on the bottom edge of the initial detection box of the target object and third coordinate information of a third preset position on the top edge of the initial detection box. The above step S12 can include the following steps: step S21: determining the horizontal actual size of the target side of the target object on the horizontal plane according to the coordinate information of the sticking point. Step S22: determining the initial height of the target object according to the second coordinate information, the third coordinate information, the coordinate information of the image acquisition device and the horizontal actual size.

[0039] The bottom edge of the initial detection box is the edge of the initial detection box closest to the road surface of the current frame of image. The top edge of the initial detection box is the edge of the initial detection box farthest from the road surface of the current frame of image. The second preset position can be the position of the bottom center point at the center of the bottom edge. The third preset position can be the position of the bottom center point at the center of the top edge.

[0040] The target side of the target object represents any one side of the outside of the target object. Specifically, the target side of the target object can be the side of the target object facing the image acquisition device, i.e. the side opposite to the image acquisition device. The target side of the target object can also be the side perpendicular to the opposite side of the target object and the image acquisition device. For example, in the case of a vehicle as the target object, the target side can be the side corresponding to the front of the vehicle or the side corresponding to the side of the vehicle. The horizontal actual size represents the actual size of the target side of the target object in the horizontal direction.

[0041] In some application scenarios, the step S21 can be: taking the sticking points on the target side as target sticking points, and taking the distance between the target sticking points as the horizontal actual size. In other application scenarios, the step S21 can be: determining whether each sticking point is on the target side; in response to any two sticking points being on the target side, taking the sticking points on the target side as target sticking points, and taking the distance between the target sticking points as the horizontal actual size. In response to one sticking point being on the target side, taking the sticking point on the target side as a target sticking point. In response to the angle between the line connecting the target sticking point and any other sticking point in the sticking points except the target sticking point and the target side being less than or equal to a preset angle, taking the other sticking point as a candidate sticking point. Taking the distance between the target sticking point and the candidate sticking point as the horizontal actual size.

[0042] In some application scenarios, the step S22 can include: obtaining a preset coordinate association relationship, the preset coordinate association relationship representing an association relationship between the coordinate information of the image acquisition device and the second preset position, the third preset position, the horizontal actual size, and the initial height in the initial detection frame. Inputting the second coordinate information, the third coordinate information, the coordinate information of the image acquisition device, and the horizontal actual size into the initial height estimation model to obtain the initial height output by the initial height estimation model. The initial height estimation model has the preset coordinate association relationship built therein.

[0043] In some embodiments, the step S22 can include the following steps: determining a second horizontal distance between the image acquisition device and the second preset position in the horizontal direction based on the second coordinate information and the coordinate information of the image acquisition device; determining a third horizontal distance between the image acquisition device and the third preset position in the horizontal direction based on the third coordinate information and the coordinate information of the image acquisition device; and determining the initial height of the target object according to a geometric relationship between the height of the image acquisition device from the ground, the second horizontal distance, the third horizontal distance, and the horizontal actual size.

[0044] The vertical line between the image acquisition device and the road ground is taken as a target vertical line. The second horizontal distance represents the vertical distance between the second preset position and the target vertical line. The third horizontal distance represents the vertical distance between the third preset position and the target vertical line. The height of the image acquisition device from the ground represents the length of the target vertical line. The geometric relationship represents that the height of the image acquisition device from the ground and the initial height satisfy a preset geometric proportion. The preset geometric proportion can represent the proportion of a preset similar triangle. In some application scenarios, the proportion of the preset similar triangle can be determined according to the second horizontal distance, the third horizontal distance, and the horizontal actual size. The initial height of the target object is determined based on the proportion of the preset similar triangle and the height of the image acquisition device from the ground.

[0045] In some application scenarios, the initial height of the target object is determined according to the geometric relationship between the height of the image acquisition device from the ground, the second horizontal distance, the third horizontal distance, and the horizontal actual size, including: taking the difference between the third horizontal distance and the second horizontal distance as an initial horizontal difference. Taking the difference between the initial horizontal difference and the horizontal actual size as a target horizontal difference. Taking the ratio between the target horizontal difference and the third horizontal distance as a target horizontal ratio. Taking the product between the height of the image acquisition device from the ground and the target horizontal ratio as the initial height. Specifically, the process of determining the initial height of the target object can refer to the following formula (1): Formula (1); Wherein, h represents the initial height of the target object. represents the height of the image acquisition device from the ground. represents the third horizontal distance. represents the second horizontal distance. represents the horizontal actual size. represents multiplication. In some application scenarios, the horizontal actual size can represent the length or width of the target object. In this application, the initial detection frame is the detection frame obtained by target detection of the front face of the target object facing the image acquisition device, the target side is the side of any one side face of the target object, and the horizontal actual size represents the length of the target object, i.e. the horizontal actual size represents the length of the vehicle, which will not be described again.

[0046] It can be considered that the initial height of the target object determined by the second coordinate information, the third coordinate information, the coordinate information of the image acquisition device, and the horizontal actual size is the image estimation height that meets the image frames at different times, so that the accuracy of the initial height is higher.

[0047] Please refer to Figure 3 , Figure 3 is Figure 1 a sub-process diagram of step S13 in

[0048] In some embodiments, the above step S13 can include the following steps: step S31: adjusting the initial height to obtain a target height. Step S32: adjusting the initial detection frame according to the target height to obtain a corrected detection frame.

[0049] In some application scenarios, the step S31 can be adjusting the initial height according to the height adjustment ratio corresponding to the occlusion degree of the target object to obtain the target height. In other application scenarios, the step S31 can be adjusting the initial height of the initial detection frame in the current frame image by using the historical initial heights of the historical initial detection frames in the historical frame images to obtain the target height. Specifically, the historical frame images can be each frame historical frame image containing the same target object, or each frame historical frame image containing the same target object and the target object not being occluded.

[0050] In some application scenarios, the step S32 can be taking the target height as the height of the initial detection frame. The side to be adjusted in the initial detection frame is obtained. The opposite side of the side to be adjusted is taken as the reference side, and the height of the initial detection frame is taken as the reference height to determine the adjusted side corresponding to the side to be adjusted, which is taken as the first adjusted side. Each side of the initial detection frame except the side to be adjusted and the opposite side of the side to be adjusted is taken as each candidate adjusted side, and each candidate adjusted side is extended or shortened to the adjusted side to obtain the corresponding second adjusted side. The frame formed between the first adjusted side, each second adjusted side, and the opposite side of the side to be adjusted is taken as the modified detection frame. When the occluded side in the initial detection frame is the bottom side of the frame, the bottom side of the frame in the initial detection frame is taken as the side to be adjusted. When the occluded side in the initial detection frame is the top side of the frame, the top side of the frame in the initial detection frame is taken as the side to be adjusted.

[0051] It can be considered that the target height obtained by adjusting the initial height is an estimated height combined with the image feature information of the target object, and the accuracy of estimating the height of the target object in the current frame image by using the target height is higher, so that the accuracy of the modified detection frame obtained by modifying the initial detection frame by using the target height is higher.

[0052] Please refer to Figure 4 , Figure 4 is Figure 3 the sub-process diagram of step S31 in FIG. 4.

[0053] In some embodiments, the step S31 can include the following steps: step S41: obtaining the historical height of the target object determined by at least one historical frame image. Step S42: determining the target height based on each historical height and the initial height.

[0054] The collection time of the historical frame image is earlier than the collection time of the current frame image. The historical frame image is an image collected by the image collection device at a historical time on the same target object. In this application, the historical frame image is taken as an example of an image collected by the image collection device at a historical time on the same target object and the target object in the historical frame image not being occluded, and the subsequent description will not be repeated.

[0055] In some application scenarios, the step S41 can be: for each historical frame image, in response to the target object in the historical initial detection box of the historical frame image not being occluded, obtaining coordinate information of the historical initial detection box, historical pasting coordinate information of the target object in the historical initial detection box, and coordinate information of the image acquisition device. The historical initial height of the target object in the historical frame image is predicted according to the coordinate information of the historical initial detection box, the historical pasting coordinate information of the target object, and the coordinate information of the image acquisition device. This process is the same as the process performed in the step S12 described above, and will not be repeated here. In some application scenarios, the historical initial height in the historical frame image is taken as the historical height of the target object in the historical frame image. In other application scenarios, the historical initial height is adjusted to obtain a target historical height, and the target historical height is taken as the historical height of the target object in the historical frame image. This process is the same as the process performed in the step S31 described above, and will not be repeated here. In other application scenarios, the step S41 can be: calling a historical height corresponding to at least one historical frame image from a preset database.

[0056] In some application scenarios, the step S42 can be: performing weighted fusion on each historical height and the initial height to obtain a target height corresponding to the initial height. In the weighted fusion process, the weights of each historical height and the initial height can be the same preset weight, that is, the result obtained by performing mean value processing on each historical height and the initial height is taken as the target height. In the weighted fusion process, the weights of each historical height and the initial height can be different weights.

[0057] In some embodiments, the step S42 can include the following steps: obtaining a first weight of each historical height and a second weight of the initial height. The target height is obtained according to each historical height, the first weight of each historical height, the initial height, and the second weight.

[0058] The first weight represents the accuracy when the corresponding historical height is taken as the estimated height in the historical frame image. The second weight represents the accuracy when the initial height is taken as the estimated height in the current frame image.

[0059] In some application scenarios, the estimated height of the target object in each image changes according to a preset rule according to different road regions where each image is located, for example, the closer the distance between each image and the image capture device, the higher the estimated height of the target object in each image. Wherein, the road region where each image is located can be a specific road position of the target object appearing at the image capture moment in the road within the capture range of the image capture device. For example, the empirical height of the target object is determined according to the vehicle type to which the target object belongs. The road in the image captured by the image capture device is divided into a plurality of preset road regions, and the empirical sub-height corresponding to each preset road region is adjusted and set according to the empirical height. Specifically, the way to obtain the first weight of each historical height can be as follows: for each historical frame image, the following steps are performed: matching the road region where the target object is located in the historical frame image with a plurality of preset road regions to obtain a target preset road region. Based on the difference between the empirical sub-height corresponding to the target preset road region and the historical height corresponding to the historical frame image, the first weight of the historical height corresponding to the historical frame image is determined. Specifically, the way to obtain the second weight of the initial height can be as follows: matching the road region where the target object is located in the current frame image with a plurality of preset road regions to obtain a target preset road region. Based on the difference between the empirical sub-height corresponding to the target preset road region and the initial height corresponding to the current frame image, the first weight of the initial height corresponding to the current frame image is determined.

[0060] In some embodiments, the above step of obtaining the first weight of each historical height can include the following steps: determining the confidence of the historical frame image corresponding to each historical height according to a preset rule; and taking the confidence corresponding to each historical frame image as the first weight of the corresponding historical height.

[0061] In some embodiments, the above step of obtaining the second weight of the initial height can include the following steps: determining the confidence of the current frame image according to a preset rule; and taking the confidence corresponding to the current frame image as the second weight of the corresponding initial height.

[0062] The preset rule can be inputting each frame image into a confidence determination module, and determining the confidence output by the confidence determination module as the confidence of the corresponding image. Each frame image can be each historical frame image and / or the current frame image, which will not be described in detail hereinafter. Specifically, the processing of the image according to the preset rule includes but is not limited to at least one of the following: detection confidence, occlusion degree, perspective information, position distance, initial detection box size, timestamp, etc. The detection confidence can be the confidence of the initial detection box output when the image is processed by a preset target detection. The occlusion degree represents the degree of occlusion of the target object in the initial detection box. The perspective information represents the heading angle of the target object in the image or the driving direction of the target object, etc. The position distance represents the distance between the target object in the image and the image capture device. The timestamp represents the image capture moment of the image.

[0063] In some embodiments, the step S42 can include the following steps: constructing a compensation height of the initial height according to each historical height. Based on the initial height and the compensation height, the target height is determined.

[0064] The compensation height of the initial height represents an adjustment value of the initial height.

[0065] In some embodiments, the step of constructing the compensation height of the initial height according to each historical height can include the following steps: taking a difference between an average value of each historical height and a preset average value as an initial compensation height of the initial height; updating the initial compensation height to obtain the compensation height according to a difference between a first historical height and the initial height, the first historical height representing a historical height determined by the vehicle at a capture time of the target historical frame image; or, adjusting the initial compensation height to obtain the compensation height according to a confidence degree of the current frame image; or, adjusting the initial compensation height to obtain the compensation height according to a road region to which the vehicle belongs at the capture time of the current frame image.

[0066] Exemplarily, the step of constructing the compensation height according to the initial height of each historical height comprises the following: dividing the road in the image captured by the image acquisition device into a plurality of preset road sections. The division manner can be equal interval division or other road section division manner, which is not limited herein. The historical height corresponding to the case that the target object in the historical frame image is not occluded is determined as the candidate historical height. In some application scenarios, in the case that the target object in the current frame image is in the first road section, the first average value obtained by performing mean value processing on the initial heights of all target objects in all images collected in the first road section is taken as the compensation height corresponding to the initial height of the target object in the current frame image. The all images include the current frame image and the historical frame image corresponding to the current frame image, and the target object in each image is not occluded and / or other target objects are not occluded. The other target objects in each frame image include target objects other than the target object, and the other target objects and the target object belong to the same object type, which are vehicles or express boxes, etc. In other application scenarios, in the case that the target object in the current frame image is not in the first road section, the road section where the target object in the current frame image is located is taken as the current section. The first average value is obtained by performing mean value processing on the initial heights of all target objects in all images collected in all road sections. The road section adjacent to the current section is taken as the candidate section, and the second average value is obtained by performing mean value processing on the candidate historical heights of the same target object in the candidate section in the historical frame image. The first average value and the second average value are processed by mean value processing to obtain the target average value. The target average value is taken as the offset of the current section. The offsets of all road sections are accumulated to obtain the compensation height corresponding to the current section, and the compensation height corresponding to the current section is taken as the compensation height corresponding to the initial height of the current frame image.

[0067] In some application scenarios, the step of determining the target height based on the initial height and the compensation height comprises: directly taking the difference between the initial height and the compensation height as the target height corresponding to the initial height. In other application scenarios, the step of determining the target height based on the initial height and the compensation height comprises: obtaining each historical height and a historical compensation height corresponding to each historical height. It can be understood that, for the determination of each historical compensation height, an image collected earlier than the historical frame image can be taken as a target historical frame image. The historical compensation height of the historical frame image is constructed according to the historical height corresponding to each target historical frame image, and this process can refer to the step of constructing the compensation height of the initial height according to each historical height, which will not be described here. Then, the difference between the historical height and the historical compensation height of the historical height is taken as an adjusted historical height of the historical height. Based on each adjusted historical height and the initial height, the target height is determined, specifically comprising: obtaining a third weight of each adjusted historical height and a second weight of the initial height. The target height is obtained according to each adjusted historical height, the third weight of each adjusted historical height, the initial height and the second weight. The determination manner of the third weight of each adjusted historical height is the same as that of the first weight of the corresponding historical height. In some application scenarios, the third weight of each adjusted historical height is the same as the first weight of the corresponding historical height.

[0068] It can be considered that the target height obtained by adjusting the compensation height of the initial height can represent the accurate estimated height of the target object in the current frame image, and the modified detection frame obtained by modifying the initial detection frame with the target height is more accurate.

[0069] Please refer to Figure 5 , Figure 5 is Figure 3 the sub-flowchart of step S32 in

[0070] In some embodiments, the coordinate information of the initial detection frame comprises first coordinate information of a first preset position on the top side of the initial detection frame. The step S32 can comprise the following steps: step S51: determining the to-be-adjusted coordinate information of the target position on the bottom side of the initial detection frame according to the height of the image collection device from the ground, the first horizontal distance of the image collection device and the first preset position in the horizontal direction, the horizontal actual size of the target side of the target object on the horizontal plane and the target height. Step S52: taking the intersection point between the extension line of each side of the initial detection frame and the horizontal line where the to-be-adjusted coordinate information is located as a target intersection point. Step S53: generating the modified detection frame with each target intersection point and each corner point on the top side as a reference corner point.

[0071] The first preset position and the third preset position can be the same or different positions on the top edge of the initial detection frame. The preset positions represent unoccluded positions on the top edge of the initial detection frame. For example, the first preset position can be a position of a bottom center point of the frame that is at the center of the top edge of the frame. The target position can be a position to be adjusted on the bottom edge of the initial detection frame. For example, the target position can be a position of a bottom center point of the frame that is at the center of the bottom edge of the frame. The adjusted coordinate information represents coordinate information to be adjusted corresponding to the target position. The horizontal line on which the coordinate information to be adjusted is located is parallel to the top edge of the initial detection frame. The step S53 can be to use the line connecting the adjacent reference corner points as each edge of the corrected detection frame.

[0072] In some application scenarios, the scale of the preset similar triangle is determined based on the height of the image acquisition device from the ground and the target height. The coordinate information to be adjusted of the target position on the bottom edge of the initial detection frame is determined based on the scale of the preset similar triangle, the first horizontal distance between the image acquisition device and the first preset position in the horizontal direction, and the horizontal actual size of the target side of the target object on the horizontal plane. In other application scenarios, the height of the image acquisition device from the ground, the first horizontal distance between the image acquisition device and the first preset position in the horizontal direction, the horizontal actual size of the target side of the target object on the horizontal plane, and the target height are input into a coordinate information to be adjusted estimation model to obtain the coordinate information to be adjusted output by the coordinate information to be adjusted estimation model.

[0073] It can be considered that the corrected detection frame obtained by adjusting the initial detection frame based on the target height is a detection frame that is more consistent with the estimated height of the target object in the current frame image, and thus the accuracy of the corrected detection frame is higher.

[0074] Some trajectory matching methods and trajectory matching devices are provided. The application scenarios of the trajectory matching method include but are not limited to target positioning and trajectory matching scenarios, target tracking and trajectory matching scenarios. The execution subject of the trajectory matching method can be a trajectory matching device or the above-mentioned target positioning device. For example, the trajectory matching device can be arranged in a terminal device or a server or other processing device. The terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementation manners, the trajectory matching method can be realized by a processor calling computer readable instructions stored in a memory. In some possible implementation manners, the trajectory matching method can be realized by the processor calling the corrected detection frame output by the above-mentioned target positioning device.

[0075] Please refer to Figure 6 , Figure 6is a flowchart of an exemplary embodiment of the trajectory matching method of the present application. Specifically, the trajectory matching method can include the following steps: step S61: in response to the existence of occlusion in the initial detection box of the target object of the current frame image, obtaining a corrected detection box corresponding to the initial detection box. Step S62: based on the corrected detection box, tracking matching is performed to obtain the trajectory information of the target object.

[0076] wherein the corrected detection box is obtained by the above-mentioned target positioning method. The above-mentioned step S61 can refer to the above-mentioned steps S11 to S13, which will not be described here. The tracking matching can be a preset tracking matching algorithm. Specifically, the preset tracking matching algorithm can be provided with Kalman filtering and Hungarian algorithm, or only the intersection over union of the detection box is used for trajectory matching.

[0077] Exemplarily, please refer to Figure 7a , Figure 7a is a framework diagram of an exemplary embodiment of the target positioning method of the present application.

[0078] As Figure 7a shown, the execution process of the above-mentioned steps S11 to S13 can be inputting the current frame image into the target positioning model to obtain the corrected detection box output by the target positioning model. The current frame image sequentially passes through the initial height prediction module, the compensation height construction module, the target height prediction module and the initial target positioning module in the target positioning model to obtain the output corrected detection box.

[0079] wherein the above-mentioned steps S11 and S12 can be inputting the current frame image into the initial height prediction module to obtain the initial height output by the initial height prediction module. Specifically, the implementation process of the initial height prediction module can include the following contents: using a 2D detection algorithm (such as YOLO series, RCNN series, CenterNet series), obtaining the 2D bounding box (i.e. initial detection block) of all target objects in the current frame image, for a certain target object , assuming that the left upper corner point and the right lower corner point of the initial detection box of the target object are and , respectively represented as and . According to the left upper corner point and the right lower corner point of the initial detection box, the center point of the top edge and the center point of the bottom edge of the initial detection box of the target object on the current frame image can be obtained, respectively represented as and , respectively represented as and . Assuming that the optical center point of the image acquisition device of the current frame image is , represented as .

[0080] The present application takes a vehicle as an example of a target object, and uses a vehicle projection point estimation algorithm, such as a key point detection algorithm (Simple Baselin, SBL), to obtain first patch point coordinate information of four patch points of the vehicle, wherein the first patch point coordinate information is in an image coordinate system. The four patch points include a left front projection point, a right front projection point, a left rear projection point, and a right rear projection point of the vehicle. Different patch points represent contact points between tires of the vehicle and a road surface. The left front projection point, the right front projection point, the left rear projection point, and the right rear projection point are respectively represented as , , and , and are respectively represented as , , and . Then, the four patch points are converted from the image coordinate system to a world coordinate system to obtain corresponding second patch point coordinate information. The second patch point coordinate information is in the world coordinate system. Each second patch point coordinate information is taken as patch point coordinate information of each patch point. Specifically, the left front projection point, the right front projection point, the left rear projection point, and the right rear projection point in the world coordinate system are respectively represented as , , and , and are respectively represented as , , and , wherein represents a projection transformation coefficient from an image plane to a ground plane. In some application scenarios, the horizontal actual size determined in the above step S21 can be a length and / or a width of the vehicle. In some application scenarios, the horizontal actual size is determined based on patch point coordinate information of two patch points on a target side. In other application scenarios, the horizontal actual size is determined based on patch point coordinate information of all patch points. Specifically, a process of determining the length of the target object in the horizontal actual size is described in the following formula (2), and a process of determining the width of the target object in the horizontal actual size is described in the following formula (3).

[0081] Formula (2); Formula (3); wherein represents the length of the target object. represents the width of the target object. The remaining parameters in the formula (2) and the formula (3) refer to the coordinate information of the left front projection point, the right front projection point, the left rear projection point, and the right rear projection point in the world coordinate system described above. For example, represents a distance between a left front projection point and a left rear projection point in a world coordinate system. It can be understood that the length and / or width of the target object are used to determine the above horizontal actual size according to the relative relationship between each face in the target object and the image acquisition device.

[0082] In some application scenarios, the step S22 can be estimating the initial height of the target object using the extrinsic parameter of the image acquisition device, the homography matrix between the image plane in the current frame image and the ground plane and geometric principles, and specifically includes the following contents: using the center point of the top edge and the center point of the bottom edge in the initial detection frame are converted from the coordinates on the image to the coordinates on the ground plane, and specifically, the projection point of the center point of the bottom edge is represented as and the projection point of the center point of the top edge is represented as Specifically, the process of calculating the horizontal distance between the target object and the image acquisition device can refer to the following formulas (4) and (5), and the process of deducing the initial height of the target object using geometric relationships can refer to the above formula (1).

[0083] Formula (4); Formula (5); wherein the coordinate information of the optical center point of the image acquisition device is represented as , represents a third horizontal distance. represents a second horizontal distance. represents the coordinate information of the projection point of the center point of the bottom edge in the initial detection frame in the world coordinate system. represents the coordinate information of the projection point of the center point of the top edge in the initial detection frame in the world coordinate system. Specifically, and are respectively , and the distance between the camera and the position on the ground plane. As shown in the geometric relationship Figure 7b , the process of the initial height obtained by the above formula (1) can include the following contents, wherein P1 represents the coordinate information of the optical center point of the image acquisition device in the world coordinate system, P2 represents the coordinate information of the third preset position of the initial detection frame in the world coordinate system, and P3 represents the coordinate information of the second preset position of the initial detection frame in the world coordinate system. H represents the height of the image acquisition device from the ground. l represents the horizontal actual size of the target object. h represents the initial height of the target object. D1 represents the second horizontal distance, that is . D2 represents the third horizontal distance, that is .

[0084] In some application scenarios, considering that the inverse perspective transformation relies on the flat ground plane assumption, when the assumption does not hold, the height of the same target estimated at different positions is inconsistent, to alleviate the problem, the application introduces a compensation height construction module to construct a compensation height corresponding to the initial height, for details, please refer to the above content that the initial height corresponding to the compensation height is constructed in the case that the target object in the current frame image is in the first road section or not in the first road section, which will not be repeated here.

[0085] Further, when the target object in the current frame image is not positioned accurately due to being blocked and needs to be corrected using the estimated height, the TopN weighted historical heights in the historical record associated with the target object are taken to adjust the initial height to obtain the target height, for details, please refer to the above step of determining the target height based on each adjusted historical height and the initial height, which will not be repeated here. The target height of the target object in the current frame image is obtained by weighted average of each adjusted historical height and the initial height.

[0086] Considering that the height of the same target object estimated in different video frames is theoretically consistent. However, due to the existence of 2D positioning or 3D positioning errors, system errors caused by non-flat ground, the height estimated in different frames may be different. In order to alleviate this problem, the application introduces a target height prediction module to perform the following weighted average algorithm: first, determine the weight of the initial height corresponding to each frame image according to the preset rule, for details, please refer to the above content which will not be repeated here. Initialize the minimum heap, create a minimum heap structure, and set the maximum capacity to n. When inserting data, for each frame image, the initial height determined for each frame image and the weight corresponding to the initial height are taken as the data to be inserted, and the size of the minimum heap is checked, which specifically includes: if the size of the heap is less than the maximum capacity, directly insert the data to be inserted into the heap. If the size of the heap is equal to the maximum capacity, compare the size of the weight in the data to be inserted with the top element (the smallest weight) of the heap, if the weight in the data to be inserted is larger, then pop the top element of the heap (delete the minimum weight and the historical height corresponding to the historical frame image associated with the minimum weight), and then insert the data to be inserted into the heap. If the weight in the data to be inserted is smaller than the top element of the heap, ignore or delete the data to be inserted.

[0087] In the target height prediction module, when the target height needs to be calculated, the current n elements are taken out from the minimum heap, the corresponding adjusted historical height is first determined using the historical compensation height corresponding to each historical height, and then weighted average is performed. The process of determining the target height is described in the following formula (6): Formula (6); Wherein, represents the target height. i represents the coordinate index corresponding to any one of the initial height and each historical height. For example, i represents any one of the current n elements and the element corresponding to the initial height. represents the compensation height of the initial height or the historical compensation height. represents the compensation height of the initial height. represents the first weight corresponding to the historical height or the third weight corresponding to the adjusted historical height, or the second weight corresponding to the initial height. represents the total number of the current n elements or the total number of the initial height and each historical height in the min-heap. represents multiplication.

[0088] Table 1: Example table of estimated height

[0089] Specifically, in combination with Table 1, if the current frame image is an image collected at coordinate 200 of a target object driving from far to near, and the target object is not occluded at coordinates 500 and 300, the estimated height 1.9 of the target object at coordinate 500 is a historical height, the height compensation 0.5 is a historical compensation height, and the weight 0.8 is the first weight of the historical height. The estimated height 1.6 of the target object at coordinate 300 is a historical height, the height compensation 0.2 is a historical compensation height, and the weight 0.9 is the first weight of the historical height. The estimated height 1.4 of the target object at coordinate 200 is an initial height, the height compensation 0.0 is a compensation height of the initial height, and the weight 0.95 is the second weight of the initial height. It can be understood that the target height corresponding to coordinate 200 is obtained by substituting the three groups of data associated with coordinates 200, 500 and 300 into the above formula (6).

[0090] In some application scenarios, considering that in the intelligent transportation system, when the target is densely occluded, directly performing 2D positioning and 3D positioning, the results obtained are likely to be biased from the real position, such as Figure 7c and Figure 7d The rear vehicle, i.e., target object 1 (i.e., A1), is severely occluded by the front vehicle, i.e., target object 2 (i.e., A2), so that the lower bottom edge of the 2D detection box (i.e., K1) is too high, and the 3D position (i.e., K2, representing the lower bottom frame in the 3D positioning projection box) is inaccurate. The target positioning module is introduced to correct the initial detection box to obtain a corrected detection box, i.e., a corrected 2D detection box (i.e., K3) and a corrected 3D position (i.e., K4). It can be understood that, Figure 7c K1 and K2 in represent the 2D detection box and the 3D positioning projection box of the target object 1 in the initial detection box, respectively. Figure 7dIn the diagram, K3 and K4 represent the corrected 2D detection box and the corrected 3D positioning projection box of target object 1 in the corrected detection box, respectively.

[0091] It is understandable that, when the initial detection box represents the initial 3D localization result, i.e., the 3D detection box, the initial detection box in the above steps can be the front frame and the bottom frame facing the image acquisition device in the 3D detection box. The initial detection box that needs to be corrected in step S13 can be the bottom area of ​​the front frame facing the image acquisition device and the area where the bottom frame is located in the 3D detection box.

[0092] In some applications, the initial detection box can represent the 2D initial localization result, and the corrected detection box can represent the 2D target localization result, enabling trajectory tracking and matching of target objects in the current frame image based on the 2D target localization result. In other applications, the initial detection box can represent the 3D initial localization result, and the corrected detection box can represent the 3D target localization result, enabling trajectory tracking and matching of target objects in the current frame image based on the 3D target localization result. In still other applications, the initial detection box can represent both 2D and 3D initial localization results, and the corrected detection box can represent both 2D and 3D target localization results, enabling trajectory tracking and matching of target objects in the current frame image based on both 2D and 3D target localization results. During the correction process of the initial detection box, the target height can be determined based on the 2D and / or 3D initial localization results, which will not be elaborated further here.

[0093] In some application scenarios, the initial target localization module executes steps S51 to S53. The specific implementation process of the initial target localization module is as follows: Obtain the coordinates of the center point of the top edge of the 2D detection box (i.e., the initial detection box) of the target object in the current frame image. Considering that the target object may be severely occluded, there may be cases where the 2D detection box cannot be detected. To obtain the coordinates of any point on the top edge of the initial detection box, methods such as roof component detection or roof vertex estimation can be introduced to assist in obtaining the top edge coordinates. Only the top edge of the box corresponding to the roof is detected. Specifically, the method for obtaining the first coordinate information of the first preset position can include the following: the coordinate information of the first preset position in the image coordinate system is represented as... Reuse Transform the coordinates on the image to coordinates on the ground plane. This yields the coordinate information of the first preset position in the world coordinate system, specifically represented as follows: Step S51 above can be used to estimate the position to be adjusted at the bottom edge of the initial detection frame of the target object in the world coordinate system by using the target height obtained through weighted averaging, combined with camera extrinsic parameters and relevant geometric relationships. The specific determination of the to-be-adjusted coordinate information of the target position on the bottom side of the initial detection frame will be described below with reference to formulas (7) and (8): Formula (7); Formula (8); wherein, represents the coordinate information of the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the world coordinate system, that is, the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the world coordinate system is represented as . represents the target height. represents the new second horizontal distance. In some application scenarios, the first horizontal distance is equal to the third horizontal distance. represents the horizontal actual size. In this application, the horizontal actual size is taken as an example of the length of the target object. The optical center point coordinate information of the image acquisition device is represented as . The remaining parameters in formulas (7) and (8) are described above and will not be repeated here.

[0094] In some application scenarios, after obtaining the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the world coordinate system, the ground plane to image projection transformation is used to convert the to-be-adjusted position of the target position on the bottom side from the coordinate on the ground plane back to the coordinate on the image to obtain the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the image coordinate system . The process of determining the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the image coordinate system can be described below with reference to formula (9): Formula (9); wherein, represents the coordinate information of the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the image coordinate system, that is, the to-be-adjusted position of the target object on the bottom side of the initial detection frame in the image coordinate system is represented as . The remaining parameters in formula (9) are described above and will not be repeated here. The to-be-adjusted position of the target object on the bottom side of the initial detection frame in the image coordinate system is taken as the to-be-adjusted coordinate information of the target position on the bottom side of the initial detection frame output by step S51, and then steps S52 to S53 are executed.

[0095] In some application scenarios, when the target is occluded, the 2D detection box and 3D position (i.e. 3D positioning projection box) in the initial detection box obtained by directly performing the preset target detection processing on the image may not be accurate. At this time, directly using the 2D detection box / 3D position for tracking matching may cause errors. To alleviate this problem, the application introduces the following tracking matching algorithm of the modified detection box depending on the output of the above target positioning model: first, for the occluded target object, calculate the intersection over union (i.e. IOU) of the 2D detection box and the 2D detection box of all target objects in the historical trajectory pool, calculate the distance between the 3D position and the 3D position of the trajectory, and exclude trajectories with too small IOU and too large distance. Traverse the remaining trajectories, and use the initial height or target height of the target object in the remaining trajectory to correct the position of the 2D detection box and the 3D position of the current target object, and record the corrected 2D box and 3D position, i.e. to obtain the above modified detection box. Considering the IOU of the modified 2D detection box and the 2D detection box of the trajectory, and the distance between the modified 3D position and the 3D position of the trajectory, the cost matrix of the Hungarian algorithm is obtained. Through the Hungarian algorithm, the optimal matching of all occluded targets after modification and historical trajectories is obtained, and it is used as the tracking matching to obtain the trajectory information of the target object.

[0096] The above scheme, in the case that the target object in the initial detection box of the current frame image is occluded, predicts the initial height of the target object according to the coordinate information of the initial detection box, the attachment point coordinate information of the target object in the initial detection box, and the coordinate information of the image acquisition device, so that the accuracy of the predicted initial height is relatively high, and then the modified detection box obtained by correcting the initial detection box according to the initial height has high accuracy.

[0097] Please refer to Figure 8 , Figure 8 is a structural schematic diagram of an embodiment of the target positioning device. The target positioning device 80 comprises an acquisition module 81, a prediction module 82, and a correction module 83. The acquisition module 81 is configured to, in response to the target object being occluded in the initial detection box of the current frame image, acquire the coordinate information of the initial detection box, the attachment point coordinate information of the target object in the initial detection box, and the coordinate information of the image acquisition device. The prediction module 82 is configured to predict the initial height of the target object according to the coordinate information of the initial detection box, the attachment point coordinate information of the target object, and the coordinate information of the image acquisition device. The correction module 83 is configured to correct the initial detection box according to the initial height to obtain a modified detection box. The detection box represents the 2D positioning result and / or the 3D positioning result of the target object in the current frame image, and the detection box includes the initial detection box and the modified detection box.

[0098] The functions performed by each module can be referred to the target positioning method, which will not be described here.

[0099] According to the scheme, in the case that the target object in the initial detection frame of the current frame image is occluded, the initial height of the target object is predicted according to the coordinate information of the obtained initial detection frame, the coordinate information of the sticking point of the target object in the initial detection frame and the coordinate information of the image acquisition device, so that the accuracy of the predicted initial height is relatively high, and then the initial detection frame is corrected according to the initial height to obtain a corrected detection frame with high accuracy.

[0100] Please refer to Figure 9 , Figure 9 is a structural schematic diagram of an embodiment of the trajectory matching device. The trajectory matching device 90 includes a preset acquisition module 91 and a tracking matching module 92. The preset acquisition module 91 is configured to, in response to the target object in the initial detection frame of the current frame image being occluded, acquire a corrected detection frame corresponding to the initial detection frame, the corrected detection frame being obtained by the trajectory matching method. The tracking matching module 92 is configured to perform tracking matching based on the corrected detection frame to obtain trajectory information of the target object.

[0101] The functions performed by each module are described with reference to the trajectory matching method, which will not be described here again. In some application scenarios, the preset acquisition module 91 in the trajectory matching device 90 can be the target positioning device, or the trajectory matching device can include the target positioning model or the target positioning device. In another application scenario, the trajectory matching device can directly call the corrected detection frame output by the target positioning device, which is not limited here.

[0102] According to the scheme, in the case that the target object in the initial detection frame of the current frame image is occluded, the initial height of the target object is predicted according to the coordinate information of the obtained initial detection frame, the coordinate information of the sticking point of the target object in the initial detection frame and the coordinate information of the image acquisition device, so that the accuracy of the predicted initial height is relatively high, and then the initial detection frame is corrected according to the initial height to obtain a corrected detection frame with high accuracy.

[0103] Please refer to Figure 10 , Figure 10 is a structural schematic diagram of an embodiment of the electronic device. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is configured to execute program instructions stored in the memory 101 to implement the steps in the target positioning method and / or the trajectory matching method. In a specific implementation scenario, the electronic device 100 can include but is not limited to a microcomputer, a server, and in addition, the electronic device 100 can also include a notebook computer, a tablet computer and other mobile devices, which are not limited here.

[0104] Specifically, the processor 102 is configured to control itself and the memory 101 to implement the steps in the above target positioning method and / or the above trajectory matching method embodiments. The processor 102 can also be referred to as a CPU (Central Processing Unit). The processor 102 can be an integrated circuit chip having a processing capability of signals. The processor 102 can also be a general processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. In addition, the processor 102 can be jointly implemented by integrated circuit chips.

[0105] The above scheme, in the case that the target object is occluded in the initial detection frame of the current frame image, predicts the initial height of the target object according to the obtained coordinate information of the initial detection frame, the pasting point coordinate information of the target object in the initial detection frame and the coordinate information of the image acquisition device, so that the predicted initial height has high accuracy, and then the initial detection frame is corrected according to the initial height to obtain a corrected detection frame with high accuracy.

[0106] Please refer to Figure 11 , Figure 11 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 110 has program instructions 1101 stored thereon, and the program instructions 1101 are executed by the processor to implement the steps in any of the above target positioning method and / or the above trajectory matching method embodiments.

[0107] The above scheme, in the case that the target object is occluded in the initial detection frame of the current frame image, predicts the initial height of the target object according to the obtained coordinate information of the initial detection frame, the pasting point coordinate information of the target object in the initial detection frame and the coordinate information of the image acquisition device, so that the predicted initial height has high accuracy, and then the initial detection frame is corrected according to the initial height to obtain a corrected detection frame with high accuracy.

[0108] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, it will not be repeated here.

[0109] The above description of the various embodiments tends to emphasize differences between the various embodiments, and the same or similar elements can be referred to each other for brevity, and will not be repeated here.

[0110] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the above-described device implementation is only schematic; for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed elements can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0111] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0112] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that makes a contribution or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0113] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has clearly informed the personal information processing rules before processing the personal information and obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his / her personal information, the individual's authorization is obtained under the condition of using obvious mark / information to inform the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.

Claims

1. A target localization method, characterized in that, The method includes: In response to the presence of occlusion of a target object in the initial detection box of the current frame image, the coordinate information of the initial detection box, the coordinate information of the contact point of the target object in the initial detection box, and the coordinate information of the image acquisition device are obtained; The initial height of the target object is predicted based on the coordinate information of the initial detection box, the coordinate information of the target object's contact point, and the coordinate information of the image acquisition device. The initial detection box is corrected based on the initial height to obtain a corrected detection box. The detection box represents the 2D and / or 3D positioning results of the target object in the current frame image. The detection box includes the initial detection box and the corrected detection box.

2. The method according to claim 1, characterized in that, The step of correcting the initial detection box based on the initial height to obtain the corrected detection box includes: The initial height is adjusted to obtain the target height; The initial detection box is adjusted according to the target height to obtain the corrected detection box.

3. The method according to claim 2, characterized in that, The step of adjusting the initial height to obtain the target height includes: The historical height of the target object is determined by acquiring at least one historical frame image, wherein the acquisition time of the historical frame image is earlier than the acquisition time of the current frame image; The target altitude is determined based on the historical altitudes and the initial altitude.

4. The method according to claim 3, characterized in that, The step of determining the target altitude based on historical altitudes and the initial altitude includes: Obtain the first weight of each historical height and the second weight of the initial height; The target height is obtained based on each historical altitude, the first weight of each historical altitude, the initial altitude, and the second weight.

5. The method according to claim 3, characterized in that, The step of determining the target altitude based on historical altitudes and the initial altitude includes: The compensation height is constructed based on the historical altitudes to establish the initial altitude. The target height is determined based on the initial height and the compensated height.

6. The method according to claim 2, characterized in that, The coordinate information of the initial detection frame includes the first coordinate information of the first preset position on the top edge of the initial detection frame; the step of adjusting the initial detection frame according to the target height to obtain the corrected detection frame includes: Based on the height of the image acquisition device above the ground, the first horizontal distance between the image acquisition device and the first preset position in the horizontal direction, the actual horizontal dimensions of the target side of the target object on the horizontal plane, and the target height, determine the coordinate information to be adjusted for the target position on the bottom edge of the initial detection frame; The intersection point between the extended lines of each side of the initial detection frame and the horizontal line where the coordinate information to be adjusted is located is taken as the target intersection point; The corrected detection box is generated using the intersection points of each target and the corner points of the top edge of the box as reference corner points.

7. The method according to claim 1, characterized in that, The coordinate information of the initial detection box includes the second coordinate information of a second preset position on the bottom edge of the initial detection box of the target object and the third coordinate information of a third preset position on the top edge of the box; the step of predicting the initial height of the target object based on the coordinate information of the initial detection box, the contact point coordinate information of the target object, and the coordinate information of the image acquisition device includes: Determine the actual horizontal dimensions of the target side of the target object on the horizontal plane based on the location coordinate information. The initial height of the target object is determined based on the second coordinate information, the third coordinate information, the coordinate information of the image acquisition device, and the actual horizontal dimensions.

8. A trajectory matching method, characterized in that, The method includes: In response to the presence of occlusion of the target object in the initial detection box of the current frame image, a corrected detection box corresponding to the initial detection box is obtained, wherein the corrected detection box is obtained by any one of the target localization methods in claims 1 to 7; The trajectory information of the target object is obtained by tracking and matching based on the corrected detection box.

9. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the method as claimed in any one of claims 1-7 and / or the method as claimed in claim 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they are used to implement the method as described in any one of claims 1-7 and / or the method as described in claim 8.