A monocular image target positioning method, device and computer program product
By constructing a virtual field of view and mapping the pixel coordinates of the target point, the problem of large positioning error in monocular positioning technology at large pitch angles is solved, achieving high-precision monocular image target positioning and improving the practicality of the positioning method.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LCFC HEFEI ELECTRONICS TECH
- Filing Date
- 2026-06-17
- Publication Date
- 2026-07-31
AI Technical Summary
Existing monocular positioning technology fails at large pitch angles, leading to a sharp increase in lateral and longitudinal positioning errors for distant targets, which cannot meet the requirements for accurate positioning.
By acquiring parameter information from a monocular image acquisition device, a virtual field of view is constructed, and the pixel coordinates of the target point are mapped to the virtual field of view. Geometric compensation is used to correct trapezoidal distortion, thereby realizing a secondary projection model based on the virtual field of view and correcting the positioning error caused by trapezoidal distortion.
It effectively overcomes the positioning error problem of traditional linear models in deep top-down scenarios, improves the accuracy and practicality of target positioning in monocular images, and reduces the complexity of data processing.
Smart Images

Figure CN122492804A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a monocular image target localization method, apparatus, and computer program product. Background Technology
[0002] In fields such as smart cities, intelligent transportation, and perimeter security, locating targets of interest (such as vehicles, people, and abandoned objects) using monocular surveillance cameras is a core requirement. Existing monocular positioning technologies typically utilize the pinhole imaging principle, calculating the target's distance based on the camera's installation height and tilt angle, combined with the assumption of a ground plane. However, this approach suffers from model failure at large tilt angles in practical applications. Existing algorithms usually assume that the projection of the field of view onto the ground is approximately rectangular or undergoes simple linear scaling. When the camera tilt angle is large (i.e., deep-view scenes), the ground field of view actually exhibits a significant trapezoidal shape. If a linear model is continued in this case, the lateral and longitudinal positioning errors of distant targets will increase dramatically and non-linearly, reaching tens of meters, failing to meet the requirements for accurate positioning. Summary of the Invention
[0003] To address the aforementioned technical problems in the current technology, this application provides a monocular image target localization method, apparatus, and computer program product.
[0004] This application provides a method for target localization in a monocular image, comprising: acquiring parameter information of an acquisition device for acquiring a monocular image; the parameter information including at least pitch angle information corresponding to the acquisition device; constructing a virtual field of view section when the parameter information is not less than a preset parameter; mapping the pixel coordinates of the target point of the object to be measured in the monocular image to the virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point; and determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates.
[0005] In some embodiments, mapping the pixel coordinates of the target point of the object under test in the monocular image to a virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point specifically includes: determining the vertical distance between the acquisition device and the virtual field of view section; and determining the mapped virtual field of view coordinates corresponding to the target point based on the vertical distance and the parameter information.
[0006] In some embodiments, determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates specifically includes: calculating the relative coordinate information of the target point based on the vertical distance, the parameter information, and the virtual field of view coordinates.
[0007] In some embodiments, after determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates, the method further includes: converting the relative coordinate information of the target point in the monocular image into the absolute latitude and longitude information of the target point; and displaying the absolute latitude and longitude information of the target point on the monocular image in real time.
[0008] In some embodiments, the method further includes: when the parameter information is less than a preset parameter, using a linear model to calculate the relative coordinate information of the target point in the monocular image.
[0009] In some embodiments, the step of calculating the relative coordinate information of the target point in the monocular image using a linear model specifically includes: determining the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin based on the pixel coordinates of the target point in the monocular image; determining the angular offset parameter of the object under test in the monocular image relative to the acquisition device based on the physical pixel coordinates of the target point and the parameter information; and determining the relative coordinate information of the target point based on the angular offset parameter.
[0010] In some embodiments, determining the mapped virtual field-of-view coordinates corresponding to the target point based on the vertical distance and the parameter information specifically includes: determining the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin based on the pixel coordinates of the target point in the monocular image; determining the angular offset parameter of the object under test in the monocular image relative to the acquisition device based on the physical pixel coordinates of the target point and the parameter information; and determining the mapped virtual field-of-view coordinates corresponding to the target point according to the vertical distance and the angular offset parameter.
[0011] In some embodiments, the parameter information may include at least one of the following: the installation height of the acquisition device, the vertical field of view of the acquisition device, the horizontal field of view of the acquisition device, and the azimuth angle of the acquisition device.
[0012] This application also provides a monocular image target localization device, including an acquisition module, a construction module, a mapping module, and a determination module. The acquisition module acquires parameter information of the acquisition device for acquiring monocular images; the parameter information includes at least elevation angle information corresponding to the acquisition device. The construction module constructs a virtual field of view section when the parameter information is not less than preset parameters. The mapping module maps the pixel coordinates of the target point of the object to be measured in the monocular image to the virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point. The determination module determines the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates.
[0013] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the monocular image target localization method according to any one of claims 1 to 8.
[0014] Compared with the current technology, the beneficial effects of the embodiments of this application are as follows: By acquiring the parameter information of the acquisition device for acquiring monocular images, this application can construct a virtual field of view section when the parameter information is not less than the preset parameters. The pixel coordinates of the target point of the object to be measured in the monocular image are mapped to the virtual field of view section, and then the relative coordinate information of the target point of the monocular image is obtained based on this. Thus, a secondary projection model based on the virtual field of view section is realized. By correcting the trapezoidal distortion through geometric compensation, the problem of far-end target positioning error caused by trapezoidal distortion in deep top-down scenes of traditional linear models is effectively overcome. It can also reduce the complexity of data processing, improve the practicality of the positioning method of the object to be measured in monocular images, and is conducive to market promotion and application. Attached Figure Description
[0015] In drawings that are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. The drawings generally illustrate various embodiments by way of example rather than limitation and are used, together with the description and claims, to illustrate the disclosed embodiments. Where appropriate, the same reference numerals are used in all drawings to refer to the same or similar parts. Such embodiments are illustrative and not intended to be exhaustive or exclusive embodiments of the apparatus or method.
[0016] Figure 1 This is a first flowchart of the monocular image target localization method according to an embodiment of this application.
[0017] Figure 2 This is a schematic diagram illustrating the principle of mapping a monocular image target localization method to a virtual field of view in an embodiment of this application.
[0018] Figure 3 This is a schematic diagram showing the relationship between the imaging area and the pitch angle of the acquisition device in the embodiments of this application.
[0019] Figure 4 This is a schematic diagram of the imaging area of the acquisition device in an embodiment of this application.
[0020] Figure 5 This is a second flowchart of the monocular image target localization method according to an embodiment of this application.
[0021] Figure 6 This is the third flowchart of the monocular image target localization method according to an embodiment of this application.
[0022] Figure 7This is the fourth flowchart of the monocular image target localization method according to an embodiment of this application.
[0023] Figure 8 This is the fifth flowchart of the monocular image target localization method in this application embodiment.
[0024] Figure 9 This is a structural block diagram of a monocular image target localization device according to an embodiment of this application. Detailed Implementation
[0025] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0026] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0027] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0028] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0029] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0030] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the embodiments are merely examples of this application, which may be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0031] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0032] This application provides a monocular image target localization method, which can be applied in fields such as smart cities, intelligent transportation, and perimeter security. Specifically, it is applied in scenarios where a monocular acquisition device (such as a camera) is used to locate an object of interest (such as a vehicle, person, or object).
[0033] like Figure 1 As shown, the monocular image target localization method includes steps S101 to S104.
[0034] Step S101: Obtain parameter information of the acquisition device 1 for acquiring monocular images; the parameter information includes at least the pitch angle information corresponding to the acquisition device 1.
[0035] The monocular image mentioned above can be understood as image information obtained from a single camera sensor.
[0036] The aforementioned data acquisition device 1 can be installed on a pan-tilt unit. The geographical coordinates, azimuth angle, and elevation angle of the pan-tilt unit are the geographical coordinates, azimuth angle, and elevation angle of the data acquisition device 1.
[0037] The aforementioned parameter information may include external and internal parameters of the acquisition device 1. The external parameters include at least one of the following: pitch angle information of the acquisition device 1. The installation height H of data acquisition device 1, and the geographical coordinates of data acquisition device 1. To collect the longitude of the location of device 1 itself, (latitude of the location of data acquisition device 1) and azimuth angle of data acquisition device 1. The internal parameters include at least one of the following: the vertical field of view of acquisition device 1. and horizontal field of view .
[0038] Step S102: If the parameter information is not less than the preset parameters, construct a virtual field of view section.
[0039] The aforementioned preset parameters can be preset pitch distortion thresholds. By comparing the pitch angle information corresponding to the current acquisition device 1 with the pitch angle distortion threshold, it can be determined whether the current field of view distortion is not obvious or there is obvious trapezoidal distortion.
[0040] Specifically, in In such cases, the scene is determined to have trapezoidal distortion; In such cases, the scene is judged to have insignificant field distortion.
[0041] The above pitch distortion threshold The range can be from 45° to 60°.
[0042] likeFigure 2 As shown, the virtual field of view section can be perpendicular to the optical axis of the acquisition device 1, and the virtual field of view section and the imaging area of the acquisition device 1 share a side, which is the side of the imaging area closest to the acquisition device 1 (e.g., the side closest to the acquisition device 1). Figure 2 (As shown on the right), the line connecting the center point of the virtual field of view section and the optical center of the acquisition device 1 is perpendicular to the virtual field of view section, with a specific vertical length of L.
[0043] This can solve the trapezoidal distortion error of linear models in deep top-down scenarios, thereby improving the target positioning accuracy of the object under test.
[0044] Step S103: Map the pixel coordinates of the target point of the object to be tested in the monocular image to the virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point.
[0045] The pixel coordinates (PX, PY) of the target point of the object under test can be mapped to a virtual field of view section to obtain virtual field of view coordinates (MX, MY). Specifically, the principle of similar triangles can be used to map the pixel coordinates of the image plane to the coordinate system of the virtual field of view section, thereby generating virtual field of view coordinates (MX, MY).
[0046] The target point of the object to be tested can be the bottom center point of the target detection box of the object to be tested, or it can be the top center point of the target detection box of the object to be tested. This application does not make specific limitations on this, as long as it can characterize the position of the object to be tested in the monocular image.
[0047] Step S104: Determine the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates.
[0048] like Figure 2 As shown, the aforementioned relative coordinate information (RX, RY) can be the coordinates of the field of view on the physical ground plane, specifically the coordinates of the target point in the monocular image within the imaging area.
[0049] This application, by acquiring parameter information of the acquisition device 1 for acquiring monocular images, can construct a virtual field of view section when the parameter information is not less than preset parameters. The pixel coordinates of the target point of the object under test in the monocular image are mapped to the virtual field of view section, and then the relative coordinate information of the target point in the monocular image is obtained based on this. This realizes a secondary projection model based on the virtual field of view section. By correcting trapezoidal distortion through geometric compensation, it effectively overcomes the problem of far-end target positioning error caused by trapezoidal distortion in deep top-down scenes using traditional linear models. It also reduces the complexity of data processing, improves the practicality of the method for locating the object under test in monocular images, and is beneficial for market promotion and application.
[0050] In some embodiments, such as Figure 5 As shown, step S103, which maps the pixel coordinates of the target point of the object to be tested in the monocular image to a virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point, specifically includes steps S201 to S202.
[0051] Step S201: Determine the vertical distance between the acquisition device 1 and the virtual field of view section.
[0052] Step S202: Determine the virtual field of view coordinates corresponding to the target point after mapping based on the vertical distance and the parameter information.
[0053] Thus, by determining the vertical distance between the acquisition device 1 and the virtual field of view section, and based on the vertical distance and camera intrinsic and extrinsic parameters, the two-dimensional pixel coordinates of the target point in the monocular image can be accurately converted into virtual field of view coordinates on the virtual field of view section. This realizes that by utilizing the geometric constraints of monocular imaging, points in the monocular image can be located on a preset virtual reference plane, thus providing a key coordinate transformation basis for subsequent positioning in a unified spatial coordinate system.
[0054] The vertical distance L between acquisition device 1 and the virtual field of view section can be calculated using the following formula: ; Where H is the installation height of data acquisition device 1; The vertical field of view of acquisition device 1; , To collect the pitch angle information of device 1.
[0055] The principle of similar triangles can be used to map the pixel coordinates of the target point to the coordinate system of the virtual field of view section using the following formula: ; ; in, The vertical offset angle is the angular offset parameter of the object under test in the monocular image relative to the acquisition device 1. The horizontal offset angle is the angular offset parameter of the object under test in the monocular image relative to the acquisition device 1.
[0056] In some embodiments, the step S104 of determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates specifically includes: The relative coordinates of the target point are calculated based on the vertical distance, the parameter information, and the virtual field of view coordinates.
[0057] Thus, by combining the determined vertical distance, parameter information, and the mapped virtual field of view coordinates, the precise relative coordinates of the target point with respect to the physical ground plane can be accurately calculated.
[0058] Specifically, a perspective transformation matrix can be constructed from the virtual field of view section to the physical ground plane to obtain the corrected relative coordinate information (RX, RY): ; .
[0059] Where MX is the abscissa of the virtual field of view coordinates of the corresponding target point after mapping; MY is the ordinate of the virtual field of view coordinates of the corresponding target point after mapping.
[0060] In some embodiments, such as Figure 6 As shown, after determining the relative coordinate information of the target point of the monocular image based on the virtual field of view coordinates in step S104, the method further includes steps S301 to S302.
[0061] Step S301: Convert the relative coordinate information of the target point in the monocular image into the absolute latitude and longitude information of the target point.
[0062] Step S302: Present the absolute latitude and longitude information of the target point on the monocular image in real time.
[0063] In this way, by converting the relative coordinates of the target point into its absolute latitude and longitude information, that is, into absolute latitude and longitude information with a globally unified reference system, and presenting it in the monocular image in real time, high-precision geospatial semantics are injected into the object under test. This allows operators to intuitively and instantly correlate and visualize the target points of the object under test identified in the monocular image with their precise geographical locations on the real Earth's surface, thereby realizing intelligent geographic information perception function and greatly improving the spatial positioning accuracy of the object under test in various scenarios.
[0064] Specifically, by combining the geographic coordinates (such as GPS location) and azimuth of the data acquisition device 1 itself, the relative coordinate information can be converted into the absolute latitude and longitude information of the target point, so that the positioning results can be directly uploaded to the map.
[0065] The relative coordinate information (RX, RY) of the target point in the calculated monocular image can be converted into latitude and longitude coordinates (such as WGS84 standard latitude and longitude coordinates) using the following formula.
[0066] ; ; Where N is the distance component of the object to be measured relative to the camera in the due north direction; E is the distance component of the object to be measured relative to the camera in the due east direction; The azimuth angle of data acquisition device 1.
[0067] The absolute latitude and longitude information of the target point can be calculated using the following formula: ; ; Where N is the distance component of the object to be measured relative to the camera in the due north direction; E is the distance component of the object to be measured relative to the camera in the due east direction; The longitude of the camera's location; The latitude of the camera's location; The longitude of the calculated target location; The latitude of the calculated target location; It is the Earth's average radius (approximately 6,371,000 meters).
[0068] In some embodiments, the method further includes: when the parameter information is less than a preset parameter, using a linear model to calculate the relative coordinate information of the target point in the monocular image.
[0069] Thus, by introducing a pitch angle threshold determination mechanism, it is possible to automatically use a linear model to save computing power in small-angle (low-distortion) scenarios; and to automatically switch to a secondary projection model based on a virtual field of view cross section in large-angle (high-distortion) scenarios, correcting trapezoidal distortion through geometric compensation. This enables adaptive switching of different algorithms to calculate relative coordinate information based on parameter information, balancing real-time performance and saving computing power.
[0070] In some embodiments, such as Figure 7 As shown, the calculation of the relative coordinate information of the target point in the monocular image using a linear model specifically includes steps S401 to S403.
[0071] Step S401: Determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin, based on the pixel coordinates of the target point in the monocular image.
[0072] Step S402: Determine the angle offset parameter of the object under test in the monocular image relative to the acquisition device 1 based on the physical pixel coordinates of the target point and the parameter information.
[0073] Step S403: Determine the relative coordinate information of the target point based on the angle offset parameter.
[0074] In this way, pixel coordinates can be geometrically transformed and mapped to the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin. Furthermore, the angular offset parameters of the physical pixel coordinates relative to the acquisition device 1 can be derived to output the relative coordinate information of the target point on the physical ground plane.
[0075] Specifically, the target object in a monocular image can be identified using object detection algorithms (such as YOLO). A physical pixel coordinate system is established with the center of the monocular image as the origin (0, 0). Let the width of the monocular image be W and the height be h. The physical pixel coordinates of the target point are calculated using the following formula: ; ; Where h is the height of the imaging area of acquisition device 1 (e.g., Figure 3 and Figure 4 As shown, Figure 4 (As shown in the figure, 'a' represents the upper length of the imaging area of acquisition device 1, and 'b' represents the lower length of the imaging area of acquisition device 1). The x-coordinate of the physical pixel coordinates of the target point; The vertical coordinate of the physical pixel coordinates of the target point.
[0076] The angular offset parameter of acquisition device 1 is calculated using the following formula: ; .
[0077] After determining the angle offset parameters of the acquisition device 1, the relative coordinate information of the target point can be calculated in step S403 using the following formula: ; .
[0078] In some embodiments, such as Figure 8 As shown, step S202, which involves determining the virtual field of view coordinates of the corresponding target point after mapping based on the vertical distance and the parameter information, specifically includes steps S501 to S503.
[0079] Step S501: Determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin, based on the pixel coordinates of the target point in the monocular image.
[0080] Step S502: Determine the angle offset parameter of the object under test in the monocular image relative to the acquisition device 1 based on the physical pixel coordinates of the target point and the parameter information.
[0081] Step S503: Determine the virtual field of view coordinates corresponding to the target point after mapping based on the vertical distance and angle offset parameters.
[0082] In this way, by using the image pixel coordinates to calculate the angular offset of the target relative to the camera, and then combining it with the known vertical distance, the original image coordinates affected by the viewing angle tilt are "corrected" and projected onto a virtual field of view plane perpendicular to the camera optical axis, thereby outputting a virtual field of view coordinate that eliminates the camera pitch and yaw errors, providing a stable and directly usable planar position reference for the positioning of the object under test.
[0083] Understandably, object detection algorithms (such as YOLO) can be used to identify targets in monocular images. A physical pixel coordinate system is established with the center of the monocular image as the origin (0, 0), and the physical pixel abscissa of the target point is calculated using the formula described above. and ordinate And calculate the angle offset parameter using the above formula. and .
[0084] After determining the angle offset parameters of the acquisition device 1, the relative coordinate information of the target point can be calculated in step S503 using the following formula: ; .
[0085] In some embodiments, the parameter information may include at least one of the following: the installation height of the acquisition device 1, the vertical field of view of the acquisition device 1, the horizontal field of view of the acquisition device 1, and the azimuth angle of the acquisition device 1.
[0086] This application also provides a monocular image target localization device 100. For example... Figure 9 As shown, the monocular image target localization device 100 includes an acquisition module 101, a construction module 102, a mapping module 103, and a determination module 104. The acquisition module 101 acquires parameter information of the acquisition device 1 that acquires monocular images; the parameter information includes at least the pitch angle information corresponding to the acquisition device 1. The construction module 102 constructs a virtual field of view section when the parameter information is not less than a preset parameter. The mapping module 103 maps the pixel coordinates of the target point of the object to be measured in the monocular image to the virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point. The determination module 104 determines the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates.
[0087] This application, by acquiring parameter information of the acquisition device 1 for acquiring monocular images, can construct a virtual field of view section when the parameter information is not less than preset parameters. The pixel coordinates of the target point of the object under test in the monocular image are mapped to the virtual field of view section, and then the relative coordinate information of the target point in the monocular image is obtained based on this. This realizes a secondary projection model based on the virtual field of view section. By correcting trapezoidal distortion through geometric compensation, it effectively overcomes the problem of far-end target positioning error caused by trapezoidal distortion in deep top-down scenes using traditional linear models. It also reduces the complexity of data processing, improves the practicality of the method for locating the object under test in monocular images, and is beneficial for market promotion and application.
[0088] In some embodiments, the mapping module 103 is further configured to: determine the vertical distance between the acquisition device 1 and the virtual field of view section; and determine the mapped virtual field of view coordinates corresponding to the target point based on the vertical distance and the parameter information.
[0089] In some embodiments, the determining module 104 is further configured to: calculate the relative coordinate information of the target point based on the vertical distance, the parameter information, and the virtual field of view coordinates.
[0090] In some embodiments, the monocular image target localization device 100 further includes a processing module, which is configured to: after determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates, convert the relative coordinate information of the target point in the monocular image into the absolute latitude and longitude information of the target point; and display the absolute latitude and longitude information of the target point on the monocular image in real time.
[0091] In some embodiments, the processing module is further configured to: calculate the relative coordinate information of the target point in the monocular image using a linear model when the parameter information is less than a preset parameter.
[0092] In some embodiments, the processing module is further configured to: determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin based on the pixel coordinates of the target point in the monocular image; determine the angular offset parameter of the object to be measured in the monocular image relative to the acquisition device 1 based on the physical pixel coordinates of the target point and the parameter information; and determine the relative coordinate information of the target point based on the angular offset parameter.
[0093] In some embodiments, the mapping module 103 is further configured to: determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin based on the pixel coordinates of the target point in the monocular image; determine the angular offset parameter of the object to be measured in the monocular image relative to the acquisition device 1 based on the physical pixel coordinates of the target point and the parameter information; and determine the virtual field of view coordinates of the corresponding target point after mapping based on the vertical distance and the angular offset parameter.
[0094] In some embodiments, the parameter information may include at least one of the following: the installation height of the acquisition device 1, the vertical field of view of the acquisition device 1, the horizontal field of view of the acquisition device 1, and the azimuth angle of the acquisition device 1.
[0095] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the monocular image target localization method according to any one of claims 1 to 8.
[0096] Note that the various units in the embodiments of this application can be implemented as computer-executable instructions stored in memory, which, when executed by a processor, can perform corresponding steps; they can also be implemented as hardware with corresponding logical computing capabilities; or they can be implemented as a combination of software and hardware (firmware). In some embodiments, the processor can be implemented as any of an FPGA, ASIC, DSP chip, SOC (System-on-a-Chip), MPU (e.g., but not limited to Cortex), etc. The processor can be communicatively coupled to the memory and configured to execute computer-executable instructions stored therein. The memory can include read-only memory (ROM), flash memory, random access memory (RAM), dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM, static memory (e.g., flash memory, static random access memory), etc., on which computer-executable instructions are stored in any format. The computer-executable instructions can be accessed by the processor, read from the ROM or any other suitable storage location, and loaded into the RAM for the processor to execute, to implement the wireless communication methods according to the embodiments of this application.
[0097] It should be noted that in the system of this application, the components are logically divided according to the functions they are to perform. However, this application is not limited to this and can re-divide or combine the components as needed. For example, some components can be combined into a single component, or some components can be further decomposed into more sub-components.
[0098] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the system according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form. Furthermore, this application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means can be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0099] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, and such examples will be interpreted as non-exclusive.
[0100] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments can be used by those skilled in the art when reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of the application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being able to be combined with each other in various combinations or arrangements. The scope of this application should be determined by reference to the appended claims and the full scope of their equivalents.
[0101] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for target localization in monocular images, characterized in that, include: Obtain parameter information of the acquisition device for acquiring monocular images; The parameter information includes at least the pitch angle information corresponding to the acquisition device; When the parameter information is not less than the preset parameters, a virtual field of view section is constructed; The pixel coordinates of the target point of the object under test in the monocular image are mapped to the virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point; The relative coordinate information of the target point in the monocular image is determined based on the virtual field of view coordinates.
2. The monocular image target localization method according to claim 1, characterized in that, The step of mapping the pixel coordinates of the target point of the object under test in the monocular image to a virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point specifically includes: Determine the vertical distance between the acquisition device and the virtual field of view section; The virtual field of view coordinates corresponding to the target point after mapping are determined based on the vertical distance and the parameter information.
3. The monocular image target localization method according to claim 2, characterized in that, The determination of the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates specifically includes: The relative coordinates of the target point are calculated based on the vertical distance, the parameter information, and the virtual field of view coordinates.
4. The monocular image target localization method according to claim 1, characterized in that, After determining the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates, the method further includes: The relative coordinate information of the target point in the monocular image is converted into the absolute latitude and longitude information of the target point; The absolute latitude and longitude information of the target point is displayed on the monocular image in real time.
5. The monocular image target localization method according to claim 1, characterized in that, The method further includes: If the parameter information is less than the preset parameter, the relative coordinate information of the target point in the monocular image is calculated using a linear model.
6. The monocular image target localization method according to claim 5, characterized in that, The calculation of the relative coordinate information of the target point in the monocular image using a linear model specifically includes: Based on the pixel coordinates of the target point in the monocular image, determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin; Based on the physical pixel coordinates of the target point and the parameter information, the angular offset parameter of the object under test in the monocular image relative to the acquisition device is determined; The relative coordinates of the target point are determined based on the angle offset parameter.
7. The monocular image target localization method according to claim 2, characterized in that, The process of determining the mapped virtual field-of-view coordinates corresponding to the target point based on the vertical distance and the parameter information specifically includes: Based on the pixel coordinates of the target point in the monocular image, determine the physical pixel coordinates of the target point in a coordinate system with the center of the monocular image as the origin; Based on the physical pixel coordinates of the target point and the parameter information, the angular offset parameter of the object under test in the monocular image relative to the acquisition device is determined; The virtual field of view coordinates corresponding to the target point after mapping are determined based on the vertical distance and angular offset parameters.
8. The monocular image target localization method according to claim 1 or 3, characterized in that, The parameter information shall include at least one of the following: the installation height of the acquisition device, the vertical field of view of the acquisition device, the horizontal field of view of the acquisition device, and the azimuth angle of the acquisition device.
9. A monocular image target localization device, characterized in that, include: An acquisition module is used to acquire parameter information of an acquisition device for acquiring monocular images; the parameter information includes at least the pitch angle information corresponding to the acquisition device. A construction module is used to construct a virtual field of view section when the parameter information is not less than a preset parameter; The mapping module is used to map the pixel coordinates of the target point of the object to be tested in the monocular image to a virtual field of view section to obtain the mapped virtual field of view coordinates corresponding to the target point. The determination module is used to determine the relative coordinate information of the target point in the monocular image based on the virtual field of view coordinates.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the monocular image target localization method according to any one of claims 1 to 8.