A positioning method, apparatus and system
By determining the pixel coordinates of the target object and the mapping relationship between the image grid and the physical space grid, the problems of high resource consumption and computational load in camera positioning methods are solved, and efficient target object positioning is achieved.
Patent Information
- Application Number
- CN202010821444.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-14
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-08-14
AI Technical Summary
In existing technologies, camera-based positioning methods consume a lot of resources for real-time positioning, involve a large amount of computation, and are difficult to obtain sample data.
By determining the pixel coordinates of the first localization pixel of the target object and utilizing the mapping relationship between the image grid and the physical space grid, real-time localization of the target object can be achieved, reducing the amount of computation and improving localization efficiency.
It achieves reduced resource consumption, improved positioning efficiency, and reduced computational load during real-time positioning.
Smart Images

Figure CN114076970B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of communication, and in particular to a positioning method, device and system. BACKGROUND
[0002] Cameras play an important role in the acquisition of dynamic information in cities, and through visual positioning technology, the position information of targets such as pedestrians and vehicles is extracted from video images. The position information of targets in cities is the basis for constructing a dynamic information library of cities. Cameras can be widely used in intelligent transportation, safe city, smart park and other scenarios. At present, the positioning methods for positioning targets in cities using cameras can be divided into the following two categories:
[0003] The first category is dynamic video positioning, that is, a monocular, binocular or multi-camera mounted on a mobile platform is used to estimate the relative position of the target in the city to the mobile platform by using stereo measurement, and then the global navigation satellite system (GNSS) and inertial navigation system mounted on the mobile platform are used to orient the target absolutely, so as to realize the spatial positioning of the target. For example, by using the Structure from Motion (SfM) or simultaneous localization and mapping (SLAM) technology, a monocular camera mounted on a micro four-rotor unmanned aerial vehicle is used to obtain video in real time, and the visual odometry construction, 2D pedestrian detection and tracking are completed on the unmanned aerial vehicle platform to recover the camera position and the 3D trajectory of the pedestrian in real time. In combination with the SfM algorithm and the vehicle target detection method, the boundary line of the vehicle in the 2D target detection is determined by tracking the sparse features of the SfM algorithm, and the monocular video sensor attitude is recovered to determine the three-dimensional vehicle position information in the scene in real time.
[0004] The second category is static video positioning, that is, a static camera is fixed in the city or community, a large amount of sample data is trained to estimate the distance depth from the image point coordinates of the target in the video to the camera, and then the spatial position of the target is inferred. For example, by relying on the depth learning method to recover the monocular image depth, pedestrians and vehicles in the road scene are detected and tracked to determine the position information of pedestrians and vehicles in the scene in real time.
[0005] However, in the first positioning method, the features of the image need to be extracted and matched in real time, and the amount of calculation is large; in the second positioning method, a large amount of sample data needs to be obtained in advance, and the sample data is difficult to obtain, and the calculation amount is large by estimating pixel by pixel, so the existing positioning method consumes more resources for real-time positioning. SUMMARY
[0006] The embodiment of the present application provides a positioning method, device and system, so that the operation amount is small when a target object is positioned in real time, resource consumption is reduced, and positioning efficiency is improved.
[0007] To achieve the above object, the embodiment of the present application adopts the following technical scheme.
[0008] In a first aspect, the embodiment of the present application provides a positioning method, which comprises the following steps: determining a first positioning pixel of a target object in a first image; determining a pixel coordinate of the first positioning pixel; determining a physical space grid corresponding to the first positioning pixel according to an image grid corresponding to the first positioning pixel and a mapping relationship between the image grid and a grid of a physical space, wherein the image grid of the first image is determined according to the grid of the physical space and parameter information of a camera collecting the first image; and determining geographical position information of the target object according to the physical space grid corresponding to the first positioning pixel, wherein the grid of the physical space has corresponding geographical position information. In the embodiment of the present application, the first positioning pixel of the target object is determined in the first image; the pixel coordinate of the first positioning pixel of the target object in the first image is determined; the image grid corresponding to the first positioning pixel of the target object is determined; and then the physical space grid corresponding to the first positioning pixel is found according to the image grid corresponding to the first positioning pixel and the mapping relationship between the image grid and the grid of the physical space, and the geographical position information represented by the grid is the geographical position information of the first positioning pixel of the target object, so that the positioning of the target object is realized, the operation amount is small when the target object is positioned in real time, resource consumption is reduced, and positioning efficiency is improved.
[0009] In the method of the first aspect, the image grid is determined according to the grid of the physical space and the parameter information of the camera, specifically: the pixel coordinate of a second positioning pixel of the image grid is determined according to the geographical position coordinate of the grid of the physical space, a projection matrix and the geographical position coordinate of the viewpoint center of the camera, and the projection matrix is used to represent the conversion relationship between the pixel coordinate of the pixel in the image and the geographical position coordinate of the grid in the physical space.
[0010] In a possible design, the parameter information of the camera comprises a focal length, a principal point position, a video CCD size and a pose parameter, and the pose parameter comprises the geographical position coordinate of the camera, a pitch angle, a roll angle and a side view angle.
[0011] In a possible design, the second positioning pixel is a pixel corresponding to the geometric center of the image grid.
[0012] In a possible design, the geographical position information comprises a global position code or a geographical position coordinate.
[0013] According to the method in the first aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined, specifically: according to the type of the target object, the pixel coordinates of the reference pixel of the target object in the first image are determined, and the reference pixel is used to determine the pixel coordinates of the first positioning pixel of the target object in the first image. It should be noted that the determination method of the first positioning pixel of the target object in the first image is different based on the type of the target object.
[0014] According to the method in the first aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined, specifically: if the type of the target object is a vehicle, the pixel coordinates, the height value and the width value of the reference pixel of the vehicle in the first image are determined. According to the pixel coordinates, the height value and the width value of the reference pixel, the pixel coordinates of the first positioning pixel of the vehicle in the first image are determined.
[0015] According to the method in the first aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined, specifically: if the type of the target object is a vehicle, the pixel coordinates, the height value and the width value of the reference pixel of the vehicle in the first image are determined. According to the pixel coordinates, the height value and the width value of the reference pixel, the pixel coordinates of the first positioning pixel of the vehicle in the first image are determined.
[0016] According to the method in the first aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined, specifically: if the type of the target object is a vehicle, the pixel coordinates, the height value and the width value of the reference pixel of the vehicle in the first image are determined. According to the pixel coordinates, the height value and the width value of the reference pixel, the pixel coordinates of the first positioning pixel of the vehicle in the first image are determined.
[0017] In a second aspect, the present application provides a positioning device, comprising: a first positioning pixel determination unit configured to determine a first positioning pixel of a target object in a first image; a first positioning pixel coordinate determination unit configured to determine a pixel coordinate of the first positioning pixel; a third determination unit configured to determine a physical space grid corresponding to the first positioning pixel according to an image grid corresponding to the first positioning pixel and a mapping relationship between the image grid and the physical space grid, wherein the image grid of the first image is determined according to the physical space grid and parameter information of a camera used to capture the first image; and a geographic location information determination unit configured to determine geographic location information of the target object according to the physical space grid corresponding to the first positioning pixel, wherein the physical space grid has corresponding geographic location information. In the present application, the first positioning pixel of the target object is determined in the first image, the pixel coordinate of the first positioning pixel of the target object in the first image is determined, the image grid corresponding to the first positioning pixel is determined, and then the physical space grid corresponding to the first positioning pixel is found according to the image grid corresponding to the first positioning pixel and the mapping relationship between the image grid and the physical space grid, and the geographic location information represented by the grid is the geographic location information of the first positioning pixel of the target object, so that the target object can be positioned, and the real-time positioning of the target object can be realized with less calculation amount, reduced resource consumption and improved positioning efficiency.
[0018] In a second aspect, the present application provides a positioning device, comprising: a first positioning pixel determination unit configured to determine a first positioning pixel of a target object in a first image; a first positioning pixel coordinate determination unit configured to determine a pixel coordinate of the first positioning pixel; a third determination unit configured to determine a physical space grid corresponding to the first positioning pixel according to an image grid corresponding to the first positioning pixel and a mapping relationship between the image grid and the physical space grid, wherein the image grid of the first image is determined according to the physical space grid and parameter information of a camera used to capture the first image; and a geographic location information determination unit configured to determine geographic location information of the target object according to the physical space grid corresponding to the first positioning pixel, wherein the physical space grid has corresponding geographic location information. In the present application, the first positioning pixel of the target object is determined in the first image, the pixel coordinate of the first positioning pixel of the target object in the first image is determined, the image grid corresponding to the first positioning pixel is determined, and then the physical space grid corresponding to the first positioning pixel is found according to the image grid corresponding to the first positioning pixel and the mapping relationship between the image grid and the physical space grid, and the geographic location information represented by the grid is the geographic location information of the first positioning pixel of the target object, so that the target object can be positioned, and the real-time positioning of the target object can be realized with less calculation amount, reduced resource consumption and improved positioning efficiency.
[0019] In a possible design, the parameter information of the camera includes focal length, principal point position, video CCD size and pose parameters, and the pose parameters include geographic location coordinates of the camera, pitch angle, roll angle and side view angle.
[0020] In a possible design, the second positioning pixel is a pixel corresponding to a geometric center of the image grid.
[0021] In a possible design, the geographic location information includes global position coding or geographic location coordinates.
[0022] The positioning device of the second aspect, the first positioning pixel determining unit comprises: a first determining subunit configured to determine the pixel coordinates, the height value and the width value of the reference pixel of the vehicle in the first image if the type of the target object is a vehicle; and a second determining subunit configured to determine the pixel coordinates of the first positioning pixel of the vehicle in the first image according to the pixel coordinates, the height value and the width value of the reference pixel.
[0023] The positioning device of the second aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined in the following manner: if the type of the target object is a pedestrian, the pixels corresponding to the feet of the pedestrian in the first image are determined as the reference pixels; and the pixel coordinates of the first positioning pixel of the pedestrian in the first image are determined according to the pixel coordinates of the two reference pixels.
[0024] The positioning device of the second aspect, the pixel coordinates of the first positioning pixel of the target object in the first image are determined in the following manner: if the type of the target object is a pedestrian, the pixel corresponding to the center of gravity of the pedestrian in the first image is determined as the first positioning pixel.
[0025] The third aspect of the embodiments of the present application provides a positioning system, which comprises at least one camera and a positioning device; wherein the camera is configured to collect images and send the images to the positioning device; and the positioning device is configured to receive the images collected by the camera and perform the positioning method of the first aspect or any possible design of the first aspect.
[0026] The fourth aspect of the embodiments of the present application provides an electronic device, which comprises a processor and a memory; the memory is coupled to the processor; the memory is configured to store computer program codes; the computer program codes comprise computer instructions; when the processor reads the computer instructions from the memory, the electronic device performs the positioning method of the first aspect or any possible design of the first aspect.
[0027] The fifth aspect of the embodiments of the present application provides a computer program product, which comprises computer instructions; when the computer instructions run on a computer, the computer performs the positioning method of the first aspect or any possible design of the first aspect.
[0028] The sixth aspect of the embodiments of the present application provides a computer readable storage medium, which comprises computer instructions; when the computer instructions run on a computer, the computer performs the positioning method of the first aspect or any possible design of the first aspect.
[0029] In a seventh aspect, the present application provides a chip system, which comprises one or more processors. When the one or more processors execute instructions, the one or more processors perform the positioning method in the first aspect or any possible design of the above aspect. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0031] Figure 1 An architecture schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1.
[0032] Figure 2 Another architecture schematic diagram of a positioning system provided by an embodiment of the present application is shown in FIG. 2.
[0033] Figure 3 Another architecture schematic diagram of a positioning system provided by an embodiment of the present application is shown in FIG. 3.
[0034] Figure 4 A flowchart of a positioning method provided by an embodiment of the present application is shown in FIG. 4.
[0035] Figure 5 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 5. Figure 1
[0036] An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 6. Figure 6 Figure 2 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 7.
[0037] Figure 7 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 8. Figure 3
[0038] An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 9. Figure 8 Figure 4 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 10.
[0039] An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 11. Figure 9 Figure 5 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 12.
[0040] An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 13. Figure 10 An application scenario schematic diagram of a positioning method provided by an embodiment of the present application is shown in FIG. 14.
[0041] Figure 11 An application scenario of a positioning method provided by an embodiment of the present application Figure 6 ;
[0042] Figure 12 An application scenario of a positioning method provided by an embodiment of the present application Figure 7 ;
[0043] Figure 13 A flowchart of another positioning method provided by an embodiment of the present application
[0044] Figure 14 An application scenario of a positioning method provided by an embodiment of the present application Figure 8 ;
[0045] Figure 15 A composition schematic diagram of a positioning device provided by an embodiment of the present application
[0046] Figure 16 A composition schematic diagram of a positioning system provided by an embodiment of the present application DETAILED DESCRIPTION
[0047] Cameras play an important role in the acquisition of dynamic information in cities, and they extract the position information of targets such as pedestrians and vehicles from video images through visual positioning technology. The position information of targets in cities is the basis for constructing a dynamic information library of cities. Cameras can be widely used in intelligent transportation, safe cities, and smart parks. At present, the positioning methods for positioning targets in cities by using cameras can be divided into the following two types:
[0048] The first type is a target positioning method based on a mobile camera, that is, a monocular, binocular, or multi-camera mounted on a mobile platform is used to estimate the relative position of a target in a city to the mobile platform by using a stereoscopic measurement method, and then a global navigation satellite system (GNSS) and an inertial navigation system mounted on the mobile platform are used to perform absolute orientation on the target to realize spatial positioning of the target.
[0049] Exemplarily, a monocular camera carried by a micro four-rotor unmanned aerial vehicle is used to acquire a video in real time by using a structure from motion (SfM) or a simultaneous localization and mapping (SLAM) technology, a visual odometer construction, 2D pedestrian detection and tracking are completed on the unmanned aerial vehicle platform, a 3D trajectory of the camera and the pedestrian is recovered in real time, and a vehicle target detection method is combined with the SfM algorithm, a boundary line of the vehicle in 2D target detection is determined by tracking sparse features of the SfM algorithm, and a monocular video sensor pose is recovered, so that 3D vehicle position information in a scene is determined in real time.
[0050] Secondly, a target positioning method based on a fixed camera is used, that is, a static camera fixed in a city or a community is used, a distance depth of an image point coordinate of a target in a video to the camera is estimated by training a large amount of sample data, and then a spatial position of the target is inferred. Exemplarily, a monocular image depth is recovered by relying on a deep learning method, pedestrians and vehicles in a road scene are detected and tracked, and position information of the pedestrians and the vehicles in the scene is determined in real time.
[0051] However, in the first positioning method, a feature of an image needs to be extracted and matched in real time, and the amount of calculation is large; in the second positioning method, a large amount of sample data needs to be acquired in advance, the sample data is difficult to acquire, and the amount of calculation is large by using pixel-by-pixel estimation. Therefore, the existing positioning method consumes more resources for real-time positioning.
[0052] To solve the technical problem, the embodiment of the present application provides a positioning method, which comprises the following steps: determining a first positioning pixel of a target object in a first image; determining a pixel coordinate of the first positioning pixel of the target object in the first image; and searching a grid of a physical space corresponding to the first positioning pixel according to an image grid corresponding to the pixel coordinate of the first positioning pixel and a mapping relationship between the image grid and the grid of the physical space, wherein the geographical position information represented by the grid is the geographical position information corresponding to the first positioning pixel of the target object, so that the positioning of the target object is realized, the table lookup positioning replaces the pixel-by-pixel intersection positioning, the amount of calculation is effectively reduced, the resource consumption is reduced, and the positioning efficiency is improved.
[0053] The positioning method provided by the embodiment of the present application is described below with reference to the drawings in the embodiment of the present application.
[0054] The positioning method provided by the embodiment of the present application can be applied to Figure 1 The electronic device shown in the left part of FIG. 1 (the electronic device can include a camera), the electronic device shown in the right part of FIG. 1 (the electronic device can not include a camera), and a positioning system composed of the electronic device and the camera. Figure 2 The electronic device shown in the left part of FIG. 1 (the electronic device can include a camera), the electronic device shown in the right part of FIG. 1 (the electronic device can not include a camera), and a positioning system composed of the electronic device and the camera.Figure 3 The positioning system shown includes a server, a camera, and an electronic device (which can include a display screen).
[0055] As shown in Figure 1 The electronic device 100 can include a processor 110, a camera 120, a universal serial bus (USB) interface 130, a memory 140, a sensor module 150, a display screen 160, and the like. The sensor module 150 can include a pressure sensor 150A and a touch sensor 150B.
[0056] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0057] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.
[0058] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0059] The memory in the processor 110 can also be provided for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0060] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, and the like. The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to the camera 120 and the like through the I2C bus interface.
[0061] The MIPI interface can be used to connect the processor 110 and peripheral devices such as the camera 120. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like. In some embodiments, the processor 110 and the camera 120 communicate through the CSI interface to implement the photographing function of the electronic device 100.
[0062] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 120 and the like. The GPIO interface can also be configured as an I2C interface, a MIPI interface, and the like.
[0063] The USB interface 130 is an interface that complies with the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, and the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transmit data between the electronic device 100 and a peripheral device.
[0064] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0065] The electronic device 100 can implement the function of locating an object through an ISP, a camera 120, a pressure sensor 150A, a touch sensor 150B, a video codec, a GPU, a display 160, and an application processor, and the like.
[0066] ISP is used to process the data fed back by the camera 120. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and conversion into a visible image. The ISP can also optimize the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 120.
[0067] The camera 120 is used to capture a static image or a video of the collection area and a static image or a video of the target object in the collection area. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or other format image signal. In some embodiments, the electronic device 100 can include one or N cameras 120, where N is a positive integer greater than 1.
[0068] In some embodiments of the present application, the camera 120 is used to collect a first image, and a first positioning pixel of the target object in the first image is used to represent the position of the target object in the physical space.
[0069] The processor 110 is used to determine the pixel coordinates of the first positioning pixel of the target object in the first image collected by the camera 120, divide the first image into a plurality of image grids according to the grid of the physical space and the parameter information of the camera 120, determine the image grid corresponding to the pixel coordinates of the first positioning pixel, and determine the geographic location information corresponding to the first positioning pixel according to the image grid corresponding to the pixel coordinates of the first positioning pixel and the mapping relationship between the image grid and the grid of the physical space, and send the geographic location information corresponding to the first positioning pixel to the display screen 160 for display, so as to realize the positioning of the target object.
[0070] In some embodiments of the present application, the camera 120 is used to collect a second image, and the second image is an image taken by the same camera at the same pose as the first image, and the second image is divided into a plurality of grids, and each grid in the second image corresponds to a network body in the physical space.
[0071] The processor 110 is configured to determine a pixel coordinate of a first positioning pixel of a target object in a first image captured by the camera 120, divide a second image captured by the camera 120 into a plurality of image grids according to a grid of a physical space and parameter information of the camera 120, determine an image grid corresponding to the pixel coordinate of the first positioning pixel, determine geographical position information corresponding to the first positioning pixel according to the image grid corresponding to the pixel coordinate of the first positioning pixel and a mapping relationship between the image grid and the grid of the physical space, and send the geographical position information corresponding to the first positioning pixel to the display screen 160 for display, so as to realize positioning of the target object.
[0072] A video codec is used for compressing or decompressing digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, and the like.
[0073] An NPU is a neural-network (NN) computing processor that quickly processes input information by drawing on the structure of a biological neural network, such as the transmission mode between human brain neurons, and can also continuously self-learn. Through the NPU, the electronic device 100 can implement intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, and the like.
[0074] The memory 140 can be configured to store computer-executable program code including instructions. The memory 140 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, and the like), and the like. The data storage area can store data created during use of the electronic device 100 (such as audio data, a phonebook, and the like), and the like. In addition, the memory 140 can include a high-speed random access memory, and can also include a nonvolatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 110 executes various function applications and data processing of the electronic device 100 by running instructions stored in the memory 140 and / or instructions stored in a memory disposed in the processor.
[0075] The pressure sensor 150A is configured to sense a pressure signal and convert the pressure signal to an electrical signal. In some embodiments, the pressure sensor 150A can be disposed on the display 160. The pressure sensor 150A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 150A, the capacitance between the electrodes changes. The electronic device 110 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display 160, the electronic device 110 detects the intensity of the touch operation based on the pressure sensor 150A. The electronic device 110 can also calculate the location of the touch based on the detection signal of the pressure sensor 150A. In some embodiments, touch operations applied to the same touch location but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a target object in the first image, a target object selection operation instruction is executed.
[0076] The touch sensor 150B, also referred to as a "touch device". The touch sensor 150B can be disposed on the display 160, and the touch sensor 150B and the display 160 together form a touch screen, also referred to as a "touch panel". The touch sensor 150B is configured to detect a touch operation applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operation to the application processor to determine the touch event type. Visual output related to the touch operation can be provided through the display 160. For example, when the touch sensor detects a touch operation on a target object in the first image, the application processor responds to the touch operation and executes an identification operation on the target object. In other embodiments, the touch sensor 150B can also be disposed on the surface of the electronic device 110, which is different from the location of the display 160.
[0077] The display screen 160 is used to display images, videos, and geographic location information of targets, etc. The display screen 160 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 110 may include one or N display screens 160, where N is a positive integer greater than 1.
[0078] It should be noted that electronic device 100 can be a desktop computer, laptop computer, mobile phone, tablet computer, wireless terminal, embedded device, chip system, or something else. Figure 1 Equipment with a similar structure. Furthermore... Figure 1 The structural composition shown herein does not constitute a limitation on the electronic device, except... Figure 1 In addition to the components shown, the electronic device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0079] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.
[0080] like Figure 2 As shown, the positioning system 200 may include an electronic device 210 and at least one camera 220. The camera 220 may be the camera 120 described in the above embodiments, and its functionality will not be repeated here. The electronic device 210 may include the processor 110, universal serial bus (USB) interface 130, memory 140, sensor module 150, display screen 160, etc., as described in the above embodiments. Details of each component are provided in the above embodiments and will not be repeated here.
[0081] In addition, in some embodiments of this application, the electronic device 210 may also include: a wireless communication module, an antenna, etc.
[0082] The wireless communication module can provide a wireless communication solution including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), and the like, which is applied to the electronic device 210. The wireless communication module can be one or more devices integrated with at least one communication processing module. The wireless communication module receives electromagnetic waves via an antenna, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the electronic device 210. The wireless communication module can also receive signals to be transmitted from the processor, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna.
[0083] In some embodiments, the antenna and wireless communication module are coupled, enabling the electronic device 210 to communicate with a network and other devices (such as camera 220) via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0084] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 210. In other embodiments of this application, the electronic device 210 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0085] like Figure 3 As shown, the positioning system 300 may include at least one camera 310, a server 320, and an electronic device 330. The camera 310 may be the camera 120 described in the above embodiments, and its functions will not be repeated here. The server 320 may include the processor 110, universal serial bus (USB) interface 130, memory 140, etc., as described in the above embodiments. Details of each component are provided in the above embodiments and will not be repeated here.
[0086] The electronic device 330 can include the display 160 and the sensor module 150 in the above embodiments. The electronic device 330 can be configured to perform a selection operation of the target object in response to a click operation on the target object in the first image. The electronic device 330 can also be configured to display geographical location information of the target object, such as "the target object is located in XX country XX province XX city XX district XXX street", according to the geographical location information corresponding to the first positioning pixel obtained by the server 320.
[0087] Of course, in some embodiments of the present application, the server 320 can include the processor 110, the universal serial bus (USB) interface 130, the memory 140, the display 160, and the sensor module 150 in the above embodiments. The server 320 can be configured to perform a selection operation of the target object in response to a click operation on the target object in the first image. The details of the components in the server 320 are described in the above embodiments, and will not be repeated here.
[0088] The electronic device 330 can include the display 160 in the above embodiments. The electronic device 330 can be configured to display geographical location information of the target object, such as "the target object is located in XX country XX province XX city XX district XXX street", according to the geographical location information corresponding to the first positioning pixel obtained by the server 320.
[0089] In addition, in some embodiments of the present application, the server 320 can further include a wireless communication module, an antenna, and the like.
[0090] The wireless communication module can provide a wireless communication solution applied to the server 320, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), and the like. The wireless communication module can be one or more devices integrated with at least one communication processing module. The wireless communication module receives electromagnetic waves via an antenna, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the server 320. The wireless communication module can also receive signals to be sent from the processor, perform frequency modulation, amplification, and convert the signals to electromagnetic wave radiation via the antenna.
[0091] In some embodiments, the antenna and wireless communication module are coupled, enabling the server 320 to communicate with the network and other devices (such as camera 310 and electronic device 330) via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0092] It should be noted that electronic device 330 can be a desktop computer, laptop, mobile phone, tablet computer, or wireless terminal system. Furthermore, Figure 3 The structural composition shown does not constitute a limitation on this positioning system, except... Figure 3 In addition to the components shown, the positioning system may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0093] The following is based on Figure 3 Taking the illustrated architecture as an example, the positioning method provided in this application embodiment will be described. Each network element in the following embodiments may possess... Figure 3The components shown are not described again. It should be noted that the message names or parameter names in the messages exchanged between various devices in the embodiments of the present application are only examples, and other names can also be used in specific implementations. For example, the positioning pixels described in the embodiments of the present application can be replaced by anchor points, positioning points, etc. The determination in the embodiments of the present application can also be understood as creation or generation, and the "includes" in the embodiments of the present application can also be understood as "carries", which are uniformly described here. The embodiments of the present application do not make specific limitations.
[0094] Figure 4 A flowchart of a positioning method provided in the embodiments of the present application is shown in Figure 4 The method can include the following steps:
[0095] Step 401, the server determines a first positioning pixel of the target object in the first image according to the target object.
[0096] The first image can be obtained by the camera collecting the target object in the collection area. The first image can be a photo or a video.
[0097] The first positioning pixel is used to represent the position of the target object in the first image. In other words, the first positioning pixel is the pixel of the target object in the first image. For example, the first positioning pixel can be the center pixel of the bounding box of the target object in the first image, which can be understood as the pixel corresponding to the geometric center point of the bounding box of the target object. The first positioning pixel can also be any pixel on the bounding box of the target object in the first image.
[0098] Step 402, the server determines the pixel coordinates of the first positioning pixel of the target object in the first image.
[0099] In a specific implementation, the step can be implemented as: step 4021, the server determines the pixel coordinates of the reference pixel of the target object in the first image according to the type of the target object, and the reference pixel is used to determine the pixel coordinates of the first positioning pixel of the target object in the first image. Wherein, the type of the target object can refer to the category of the target object, which is used to indicate which object the target object is. For example, the type of the target object can be a pedestrian or a vehicle.
[0100] For example, the camera 1 collects the first image and sends it to the server 2. The server 2 receives the first image collected by the camera 1, and displays the first image on the display screen of the electronic device 3. When the electronic device 3 detects a click operation on the target object A on the first image, or the electronic device 3 detects a cursor (such as a mouse cursor) on the target object A on the first image, the electronic device 3 can determine the first positioning pixel of the target object A in the first image. Figure 5When the electronic device 3 detects the target object A on the first image (as indicated by the arrow in FIG. 1), the electronic device 3 performs an operation of selecting the target object A in response to an operation for the target object A on the first image, and displays a selected region of the target object A (as indicated by the selected region of the target object A in FIG. 1). The electronic device 3 sends information of the selected target object A to the server 2, and the server 2 determines a type of the target object A according to the information of the target object A, and determines a pixel coordinate of a reference pixel of the target object A in the first image. The server 2 determines a pixel coordinate of a first locating pixel of the target object A according to the pixel coordinate of the reference pixel. Figure 6
[0101] In example 1, if the type of the target object is a vehicle, the server determines a pixel coordinate of a reference pixel of the vehicle in the first image, a height value and a width value of the vehicle, and determines a pixel coordinate of a first locating pixel of the vehicle in the first image according to the pixel coordinate of the reference pixel, the height value and the width value.
[0102] For example, as shown in FIG. 2, the pixel coordinate of the first locating pixel Q0 is (xc, yc), the width value is w, and the height value is h. Figure 7
[0103] Since the vehicle is a cuboid structure, a top-left corner point Q1 of the vehicle can be determined as the first reference pixel, and a bottom-right corner point Q2 of the vehicle can be determined as the second reference pixel. The pixel coordinate of the first reference pixel is (topx, topy), and the pixel coordinate of the second reference pixel is (bottomx, bottomy).
[0104] The server determines the pixel coordinate (xc, yc) of the first locating pixel according to the pixel coordinate (topx, topy) of the first reference pixel, the pixel coordinate (bottomx, bottomy) of the second reference pixel, the width value w, and the height value h of the vehicle. In an implementation, the following relationships can be satisfied: topx = xc - 0.5 * w, topy = yc - 0.5 * h, bottomx = xc + 0.5 * w, and bottomy = yc + 0.5 * h.
[0105] In example 2, if the type of the target object is a pedestrian, the server determines a pixel corresponding to a pair of feet of the pedestrian in the first image as a reference pixel, and determines a first locating pixel according to the pixel corresponding to the pair of feet of the pedestrian in the first image.
[0106] For example, as shown in FIG. 3, the pixel coordinate of the first locating pixel Q0 is (xc, yc), the width value is w, and the height value is h. Figure 8 As shown, the server can determine that the pixel corresponding to the right foot of the pedestrian on the first image is a first reference pixel, and the pixel corresponding to the left foot of the pedestrian on the first image is a second reference pixel. It is assumed that the pixel coordinates of the first reference pixel P1 are (X1, Y1), and the pixel coordinates of the second reference pixel P2 are (X2, Y2). The server can determine that the positioning pixel P0 of the pedestrian is (X0, Y0). In a specific implementation, the following relationship can be met: X0 = 0.5 * (X2-X1), Y0 = 0.5 * (Y2-Y1).
[0107] Of course, the server can also determine the center pixel of the pedestrian from the head to the foot in the first image as the reference pixel. The pixel coordinates of the center pixel can be the pixel coordinates of the first positioning pixel, or the pixel coordinates of the first positioning pixel can be calculated according to other algorithms.
[0108] Of course, since the postures of pedestrians are various, the above embodiment only represents one implementation, and other implementations can also exist according to different postures of pedestrians, which will not be listed one by one in the embodiment of the application.
[0109] In some embodiments, before determining the pixel coordinates of the reference pixel of the target object in the first image according to the type of the target object, the positioning method provided in the embodiment of the application further includes: the server determines the type of the target object. In a specific implementation, determining the type of the target object can be specifically implemented as: the server inputs the first image and the sample image of the first real object into a deep learning model, outputs a confidence degree that the target object in the first image is the first real object, and in the case that the confidence degree is greater than or equal to a threshold value, determines that the target object is the first real object, that is, determines the type of the target object. The confidence degree is used to represent the confidence degree that the target object in the first image is the first real object. The deep learning model is trained based on the first image and the sample image of the first real object. The threshold value can be set according to requirements, for example, the threshold value can be 96%, 85%, or 92%.
[0110] In step 403, the server determines the physical space grid corresponding to the first positioning pixel according to the image grid corresponding to the pixel coordinates of the first positioning pixel, and the mapping relationship between the image grid and the physical space grid.
[0111] The image grid can be a grid obtained by grid division on the first image, or a grid obtained by grid division on a second image. The second image and the first image are images captured by the same camera at the same pose, and the second image is the base image (or initial image) of the first image.
[0112] In a specific implementation, if the image grid is a grid of the first image, determining the image grid corresponding to the pixel coordinate of the first locating pixel can be specifically implemented as follows: the server determines each image grid of the first image according to the grid of the physical space and the parameter information of the camera. The server determines the image grid of the first image to which the first locating pixel belongs according to the pixel coordinate of the first locating pixel and the coverage area of each image grid.
[0113] In a specific implementation, if the image grid is a grid of the second image, determining the image grid corresponding to the pixel coordinate of the first locating pixel can be specifically implemented as follows: the server determines the image grid to which the first locating pixel belongs according to the pixel coordinate of the first locating pixel and the second image. That is, the camera collects the second image, the server receives the second image collected by the camera, and determines each image grid of the second image according to the grid of the physical space and the parameter information of the camera. The server performs feature comparison on the first image and the second image, and establishes a same-name object set of the pixels of the first image and each image grid of the second image. The server determines the image grid of the second image to which the first locating pixel belongs according to the locating pixel and the same-name object set.
[0114] In a specific implementation, the server performs feature comparison on the first image and the second image, and establishes a same-name object set of the pixels of the first image and each image grid of the second image. The specific implementation can be as follows: the server extracts the region features of the image grid of the second image, and extracts the region features of each search region of the first image. The server compares the region features of the second image grid with the region features of each search region to determine the correlation coefficients of the second image grid and each search region. The server takes the center point of the search region with the largest correlation coefficient in each search region as the same-name object of the second image grid to construct the same-name object set of each image grid of the second image.
[0115] The specific implementation can satisfy the following expression: assuming that the pixel coordinate of the first locating pixel in the first image is x', the center pixel coordinate of the first image grid of the second image is x, and the second image is the base image of the first image, each pixel in the first image and each image grid in the second image satisfy the following expression:
[0116] x' = Hx
[0117] where the homography matrix H is which can be obtained by verifying a plurality of groups of data.
[0118] Therefore, the server determines the corresponding pixel coordinate in the second image according to the pixel coordinate of the first locating pixel in the first image and the mapping relationship between the first image and the second image, and determines the image grid corresponding to the first pixel.
[0119] In practical application, it is assumed that if the target object is at the first locating pixel a in the first image 1, the pixel coordinate of the point a is (X0, Y0). The second image is divided into a plurality of image grids (such as the grid G Figure 9 in the figure 11 -G mn ) in advance. Since the first image 1 corresponds to the second image 2, the pixel A corresponding to the first locating pixel a in the second image 2 is in the coverage area of the grid G 32 , and the server can determine that the first locating pixel a of the target object in the first image 1 corresponds to the grid G 32 in the second image 2.
[0120] Since the number of the first images collected by the camera in real time is large, the embodiment of the present application can effectively avoid the grid processing of a large number of first images, greatly simplify the processing program and reduce resource consumption by using the second image, gridizing the second image, establishing the same object set of the pixels of the first image and the image grids of the second image, and establishing the mapping relationship between each grid of the second image and each grid of the physical space.
[0121] In some embodiments of the present application, the positioning method provided by the embodiment of the present application can further include: step 405, the server gridizes the grid of the physical space to obtain each grid of the physical space. In a specific implementation, the physical space can be divided into any level according to the Discrete Global Grid (DGG) subdivision rule according to actual needs, so as to obtain each grid of the physical space. For example, the physical space can be divided into the 20th level, and the side length of each grid is 2.26 m. Since the DGG is a kind of global grid based on the sphere (or ellipsoid) which can be infinitely subdivided without changing the shape of the earth, when subdivided to a certain degree, the purpose of simulating the earth's surface can be achieved. In addition, the DGG has the characteristics of hierarchy and global continuity. Therefore, the physical space in the embodiment of the present application is divided into a grid by using the DGG subdivision rule, which not only avoids the angle, length and area distortion and the discontinuity of spatial data caused by projection, but also overcomes the constraints and uncertainties of many geographic information system (geographic information system or geo-information system, GIS) applications, so that the spatial data of any resolution (different accuracy) obtained at any position on the earth can be standardized to express and analyze, and can be operated at a certain accuracy, and the earth's surface can be accurately simulated.
[0122] The grid of the physical space is used to represent geographic location information, which can include geographic location coordinates or global position encoding. In a specific implementation, taking the geographic location information including global position encoding as an example, the global position encoding of each grid of the physical space is determined, which can be specifically implemented as follows:
[0123] Step 1, determining the longitude and latitude coordinates of the center point of each grid of the physical space, which are expressed in degrees, minutes, seconds and decimal seconds. For example, the longitude and latitude coordinates of the center point are expressed as A°B′C.D″.
[0124] Step 2, converting the longitude and latitude coordinates of the center point into binary numbers in sequence according to degrees, minutes, seconds and decimal seconds, to obtain longitude binary number and latitude binary number.
[0125] Specifically, the degrees |A| is converted from a decimal number into an 8-bit fixed-length binary number (A)2, the minutes B is converted from a decimal number into a 6-bit fixed-length binary number (B)2, the seconds C is converted from a decimal number into a 6-bit fixed-length binary number (C)2, and the decimal seconds D is converted from a decimal number into an 11-bit fixed-length binary number (D)2.
[0126] Step 3, prefixing the longitude binary number and postfixing the latitude binary number, and using the Morden crossover algorithm to obtain a binary mixed code, and converting the binary mixed code into a quaternary code.
[0127] Specifically, it can be implemented as follows: first, (A)2, (B)2, (C)2 and (D)2 obtained are directly spliced into a 31-bit fixed-length binary number (E)2 in sequence, i.e. (E)2=(A)2(B)2(C)2(D)2, to obtain two 31-bit fixed-length numbers of longitude (EL)2 and latitude (EB)2; second, prefixing the latitude (EB)2 and postfixing the longitude (EL)2, using the Morden crossover algorithm to generate a 62-bit mixed code (F)2, for example, if (EB)2 is 100111 and (EL)2 is 011010, (EB)2 is prefixed and (EL)2 is postfixed, and the binary mixed code (F)2 obtained by the Morden crossover algorithm is 100101101110; finally, converting the binary mixed code (F)2 into a quaternary code (F)4, and removing the last 32m quaternary symbols in (F)4 to obtain (F′)4 according to the level m of the grid to be solved.
[0128] Step 4, determining the global position encoding of each grid of the physical space according to the encoding of the region to which the center point of each grid of the physical space belongs in the coordinate system, the level of each grid of the physical space and the quaternary code.
[0129] Specifically, according to the longitude and latitude, the global position encoding of each grid of the physical space is determined according to the following formula: Figure 10The global position coding of each grid can be obtained by adding G0, G1, G2 or G3 before (F') 4 in the arrow direction shown in the figure.
[0130] Of course, other implementation manners can also be adopted, and the embodiments of the present application are not limited specifically.
[0131] In some embodiments of the present application, after the server receives the first image (or the second image) collected by the camera, the positioning method provided by the embodiments of the present application can further include: step 406, the server determines an image grid according to the grid of the physical space and the parameter information of the camera. The grid of the physical space is mapped to the image grid. The parameter information of the camera can include focal length, principal point position, video CCD size and pose parameters. The pose parameters can include geographical position coordinates of the camera, pitch angle of the camera, roll angle and side view angle of the camera.
[0132] In an implementable manner, the server determines the image grid according to the grid of the physical space and the parameter information of the camera, to obtain a plurality of image grids of different sizes. The specific implementation can be: the server determines the included angle between the field of view line close to the camera and the vertical line, and the included angle between the line connecting the target object and the center of the view point of the camera and the field of view line close to the camera, according to the focal length of the camera, the principal point position of the camera, the video CCD size of the camera, the geographical position coordinates of the camera, the pitch angle of the camera, the roll angle of the camera and the side view angle of the camera. The server determines the size of the image grid according to the included angle between the field of view line close to the camera and the vertical line, the included angle between the line connecting the target object and the center of the view point of the camera and the field of view line close to the camera, and the size of the grid of the physical space. The size of the grid of the physical space can be determined according to the global discrete grid division level, and different sizes of the grid of the physical space correspond to different sizes of the image grid.
[0133] Example 3, as shown in Figure 11 The parameters of the camera are known as C={X, Y, Z, ω, θ, γ}, wherein (X, Y, Z) represents the geographical position coordinates of the camera, ω represents the pitch angle of the camera, θ represents the roll angle of the camera, and γ represents the side view angle of the camera.
[0134] It is assumed that the focal length of the camera is f, the video CCD size is m×n, the opening angle of the video image domain (i.e. the included angle between the field of view line S1 and Sn) is β, and the included angle between the field of view line S1 close to the camera and the vertical line H is The pitch angle ω of the camera, the opening angle β of the video image domain (i.e. the included angle between the field of view line S1 and Sn) and the included angle between the field of view line S1 close to the camera and the vertical line H satisfy the following relationship:
[0135]
[0136] According to the principle of perspective imaging and the trigonometric function formula, the following can be obtained:
[0137]
[0138] That is:
[0139] β=2*arctan(m / 2f)
[0140]
[0141] Assuming that the angle between the line D connecting the target object A and the viewpoint center of the camera and the field of view line S1 is μ, and the size of the grid of the physical space is L1=L2=L m As shown in Figure 11 , the length L A of the image grid determined by the server can satisfy the following relationship:
[0142]
[0143] It can be seen that, with the change of the angle μ between the line D connecting the target object A and the viewpoint center of the camera and the field of view line S1, the length L A of each grid of the first image (or the second image) also changes, so that different size grid division of the first image (or the second image) can be realized, and multiple image grids of different sizes can be obtained.
[0144] In actual application, assuming that the first image has a first size and a second size of image grid, the first size is larger than the second size, and the second size of image grid can be referred to as a sub-image grid of the first size of image grid. As shown in Table 1, the first image of different size grid corresponds to the physical space of different size grid.
[0145] Table 1
[0146]
[0147] Among them, the grid G ij -ij is a sub-grid of the grid G ij , and similarly, the grid g ij -ij is a sub-grid of the grid g ij . The global position code DGG_Code ij -ij is a lower-level code of the global position code DGG_Code ij .
[0148] Step 404, the server determines the geographic location information of the target object according to the physical space grid corresponding to the first positioning pixel.
[0149] In practical applications, it is assumed that the mapping relationship table of each image grid of the first image and each grid of the physical space is table 2, as shown in table 2.
[0150] Table 2
[0151]
[0152]
[0153] In combination Figure 12 As shown in the figure, the server determines the pixel coordinates of the first positioning pixel of the target object C in the first image as (c m , t m ), and determines the image grid G m1 to which the first positioning pixel belongs according to the pixel coordinates of the first positioning pixel in the first image; the server looks up the mapping relationship table of each image grid of the first image and each grid of the physical space shown in table 2 according to the image grid G m1 of the first image, determines the grid g m1 of the physical space corresponding to the image grid G m1 of the first image; the server can determine the geolocation code of the first positioning pixel of the target object C in the first image as 0003 according to the global position code 0003 represented by the grid g m1 of the physical space, and sends it to the electronic device for display, thereby realizing the positioning of the target object C.
[0154] Similarly, the server determines the pixel coordinates of the target pixel in the first image as (c n , t n ), and determines the image grid G n1 -n1 to which the target pixel belongs according to the pixel coordinates of the target pixel in the first image; the server looks up the mapping relationship table of each image grid of the first image and each grid of the physical space shown in table 2 according to the grid G n1 -n1 of the first image, determines the grid g n1 -n1 of the physical space corresponding to the image grid G n1 -n1 of the first image; the server can determine the geolocation code of the target object taking the target pixel in the first image as the first positioning pixel as 000301 according to the global position code 000301 represented by the grid g n1 -n1 of the physical space, and sends it to the electronic device for display, thereby realizing the positioning of the target object.
[0155] Wherein, the image grid G m1 -n1 is a sub-image grid of the image grid G m1 , and the grid g m1-n1 is a grid, gm1 is a sub-grid, if the global position code 0003 is a two-level code, the global position code 000301 is a lower-level code of the global position code 0003, and the global position code 000301 is a three-level code.
[0156] The embodiment of the application can obtain multiple image grids of different sizes, establish mapping relationships between the image grids of different sizes of the first image and multiple grids of different sizes of the physical space, and realize positioning of target objects of different sizes. Therefore, the positioning method provided by the embodiment of the application is widely applied to positioning of target objects of various sizes and has strong generalization ability.
[0157] In an implementable manner, the server determines the image grid according to the grid of the physical space and the parameter information of the camera, to establish the mapping relationship between the image grid and the grid of the physical space, which can be specifically implemented as follows:
[0158] According to the geographical position coordinates of the grid of the physical space, the projection matrix, and the geographical position coordinates of the viewpoint center of the camera, the pixel coordinates of the second positioning pixel of the image grid are determined.
[0159] The projection matrix is used to represent the conversion relationship between the pixel coordinates of the pixels in the image and the geographical position coordinates in the physical space.
[0160] The projection matrix can be determined according to the world coordinate system and the angle of the camera in the world coordinate system, or the geographical position coordinates of multiple sample points in a local area are converted into geographical position coordinates in the world coordinate system after translation and rotation, and the same conversion relationship matrix used in the conversion process is the projection matrix. In the embodiment of the application, the projection matrix is a fixed parameter by default.
[0161] In Example 4, step 404 can be implemented by using the following expression: assuming that the geographical position coordinates of the viewpoint center of the camera are (X c ,Y c ,Z c ), which can be determined according to the parameter information of the camera.
[0162] The focal length of the camera is f, and the projection matrix is The geographical position coordinates (C m ,T m ,A m ) of a grid of the physical space,
[0163] The pixel coordinates of the second positioning pixel of the image grid of the first image (or the second image) are (c m ,t m ), and the expression can be as follows:
[0164]
[0165] wherein λ is a scale factor determined by the size of the grid of the physical space corresponding to the image grid.
[0166] According to the mapping relationship between the image grid and the grid of the physical space described in the above embodiments, the geographical position coordinates corresponding to the second positioning pixels of each image grid of the first image (or the second image) can be determined. According to the geographical position coordinates corresponding to the second positioning pixels of each image grid of the first image (or the second image) and the coverage area of each grid of the physical space, the grid of the physical space corresponding to the second positioning pixels of each image grid can be obtained.
[0167] Example 5, each grid of the second image is G 11 -G mn , and taking the geographical position information as a global position code as an example, assuming that each grid of the physical space is g 11 -g mn , the grid g 11 is used to represent the global position code DGG Code 11 , the grid g ij is used to represent the global position code DGG Code ij , the grid g mn is used to represent the global position code DGG Code mn , and the mapping relationship table between each image grid of the second image and each grid of the physical space can be obtained as shown in Table 3.
[0168] Table 3
[0169] Image grid of second image Grid of physical space Global position encoding G 11 ]]> g 11 ]]> DGG_Code 11 ]]> …… …… …… G ij ]]> g ij ]]> DGG_Code ij ]]> …… …… …… G mn ]]> g mn ]]> DGG_Code mn ]]>
[0170] In actual application, it is assumed that if the target object is at the first positioning pixel a of the first image 1, the coordinates of the point a are (X0, Y0). And the second image is divided into multiple grids (such as Figure 9 the grid G 11 -G mn in the second image 2) in advance. Since the first image 1 corresponds to the second image 2, the positioning pixel A corresponding to the first positioning pixel a in the second image 2 is in the coverage area of the grid G 32 , the server can determine that the first positioning pixel a of the target object in the first image 1 corresponds to the grid G 32 in the second image 2, and the server can determine the grid G 32 in the second image 2 corresponding to the first positioning pixel a according to the second image 2.The mapping relationship table of each grid of the first image and each grid of the physical space, as shown in Table 3, can obtain the global position code represented by the grid of the physical space corresponding to the first positioning pixel a in the first image 1, and send to the electronic device, and the display screen of the electronic device displays the geographic location information of the target object (such as Figure 3 the display result of the display screen of the electronic device in the middle).
[0171] Therefore, by using the positioning method provided in the embodiments of the present application, when positioning the target object, according to the grid code of the image grid corresponding to the first positioning pixel of the target object in the first image, the mapping relationship table of the image grid and each grid of the physical space is queried, and the grid of the physical space corresponding to the first positioning pixel is determined, and the geographic location information represented by the grid is the geographic location information of the first positioning pixel of the target object, so that the positioning of the target object is realized, and the table query is used instead of the pixel-by-pixel matching intersection, so as to reduce the amount of calculation, reduce the resource consumption, and improve the positioning efficiency.
[0172] In some embodiments of the embodiments of the present application, as shown in Figure 13 the positioning method in the above embodiments can further include:
[0173] Step 1301, determining the geographic location coordinates corresponding to the first positioning pixel of the target object in the third image and the geographic location coordinates corresponding to the first positioning pixel of the target object in the fourth image.
[0174] Wherein, the third image and the fourth image are located in the same plane and are collected by different two cameras, and the focal lengths of the two cameras are the same.
[0175] Example 6, as shown in Figure 14 the first positioning pixel T1 of the target object T0 in the third image collected by the first camera 1 corresponds to the geographic location coordinates (Xr, Yr), and the first positioning pixel T2 of the target object in the fourth image collected by the second camera 2 corresponds to the geographic location coordinates (Xl, Yl).
[0176] Step 1302, determining the parallax of the two cameras according to the geographic location coordinates corresponding to the first positioning pixel of the target object in the third image and the geographic location coordinates corresponding to the first positioning pixel of the target object in the fourth image.
[0177] Specifically, according to the geographic location coordinates (Xr, Yr) corresponding to the first positioning pixel T1 of the target object T0 in the third image collected by the first camera 1 and the geographic location coordinates (Xl, Yl) corresponding to the first positioning pixel T2 of the target object in the fourth image collected by the second camera 2, the parallax d = |Xr-Xl| of the first camera 1 and the second camera 2 is determined.
[0178] Step 1303, determining the height information of the target object according to the parallax of the two cameras, the focal length of the cameras, and the baseline distance between the two cameras.
[0179] As shown in the first camera 1 and the second camera 2 in the baseline distance b, and in combination with the parallax d obtained in step 1302, and the focal length of the first camera 1 and the focal length of the second camera 2 are both f, the height information of the target object T0 can be obtained Figure 14
[0180] The embodiment of the present application adopts the principle of triangulation to determine the height information of the target object according to the image of the target object collected by the binocular camera and the relationship between the binocular cameras, which can accurately determine the spatial height of the target object, thereby obtaining the three-dimensional coordinates of the target object, and make the positioning of the target object more accurate.
[0181] Figure 15 A positioning device is provided for the embodiment of the present application, and the positioning device 1500 can include:
[0182] The first positioning pixel determination unit 1501 is configured to determine the first positioning pixel of the target object in the first image according to the target object.
[0183] The first positioning pixel coordinate determination unit 1502 is configured to determine the pixel coordinates of the first positioning pixel.
[0184] The physical space grid determination unit 1503 is configured to determine the physical space grid corresponding to the first positioning pixel according to the image grid corresponding to the pixel coordinates of the first positioning pixel and the mapping relationship between the image grid and the grid of the physical space, and the image grid of the first image is determined according to the grid of the physical space and the parameter information of the camera collecting the first image.
[0185] The geographic location information determination unit 1504 is configured to determine the geographic location information of the target object according to the physical space grid corresponding to the first positioning pixel.
[0186] The grid of the physical space has corresponding geographic location information.
[0187] It should be noted that the physical space needs to be gridded in advance, so the positioning device 1500 can include a physical space gridding unit 1506 configured to divide the physical space into grids to obtain the grid of the physical space.
[0188] Further, the positioning device 1500 can include:
[0189] The image grid determination unit 1505 is used to determine the pixel coordinates of the second positioning pixel of the image grid based on the geographical coordinates of the grid in physical space, the projection matrix, and the geographical coordinates of the viewpoint center of the camera. The projection matrix is used to characterize the transformation relationship between the pixel coordinates of the pixels in the image and the geographical coordinates of the grid in physical space.
[0190] Furthermore, the camera's parameter information includes: focal length, main image point position, video CCD size, and pose parameters. The pose parameters include the camera's geographical coordinates, pitch angle, roll angle, and side view.
[0191] Furthermore, the second positioning pixel is the pixel corresponding to the geometric center of the image grid.
[0192] Furthermore, geolocation information includes: global location codes or geolocation coordinates.
[0193] Furthermore, the first positioning pixel coordinate determination unit 1502 may include:
[0194] The first determining subunit is used to determine the pixel coordinates, height value, and width value of the reference pixel of the vehicle in the first image if the type of the target object is a vehicle.
[0195] The second determining subunit is used to determine the pixel coordinates of the first positioning pixel of the vehicle in the first image based on the pixel coordinates, height value, and width value of the reference pixel.
[0196] Specifically, in this possible design, the above Figures 2-12 All relevant details regarding the steps involving the electronic device in the illustrated method embodiment can be found in the functional descriptions of the corresponding functional modules, and will not be repeated here. The electronic device described in this possible design is used to perform... Figures 2-12 The positioning method shown incorporates the functionality of the electronic device, thus achieving the same effect as the positioning method described above.
[0197] Figure 16 A positioning system 1600 provided in this application embodiment may include: at least one camera 1601 and a positioning device 1602; wherein, the camera 1601 is used to acquire a first image of a target object located within a acquisition area and a second image of the acquisition area, and sends the first image and the second image to the positioning device 1602. The positioning device 1602 is used to receive the first image and the second image acquired by the camera 1601, and perform... Figures 2-13 The positioning method is shown.
[0198] Specifically, in this possible design, the above Figures 2-12All the relevant contents of the steps involving the positioning device in the method embodiments can be referred to the function description of the corresponding function module, and will not be repeated here. The electronic device described in the possible design is used to execute Figures 2-12 The function of the positioning device in the positioning method shown, and thus the same effect as the above positioning method can be achieved.
[0199] The electronic device provided by the embodiments of the present application comprises a processor and a memory, the memory is coupled with the processor, the memory is used to store computer program code, the computer program code comprises computer instructions, when the processor reads the computer instructions from the memory, the electronic device executes Figures 2-12 The positioning method shown.
[0200] The computer program product provided by the embodiments of the present application, when the computer program product runs on the computer, makes the computer execute Figures 2-13 The positioning method shown.
[0201] The computer readable storage medium provided by the embodiments of the present application comprises computer instructions, when the computer instructions run on the terminal, the network equipment executes Figures 2-12 The positioning method shown.
[0202] The chip system provided by the embodiments of the present application comprises one or more processors, when the one or more processors execute instructions, the one or more processors execute Figures 2-12 The positioning method shown.
[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0204] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed each other can be through some interface, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.
[0205] The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0206] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0207] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for making a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk and various program code storage media.
[0208] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A positioning method, characterized by, The method comprises: determining a first positioning pixel of a target object in a first image according to the target object; determining a pixel coordinate of the first positioning pixel; determining a physical space grid corresponding to the first positioning pixel according to the pixel coordinate of the first positioning pixel and a second image, the second image being a base image of the first image; determining geographical position information of the target object according to the physical space grid corresponding to the first positioning pixel, wherein the grid of the physical space has corresponding geographical position information; the determining of the physical space grid corresponding to the first positioning pixel according to the pixel coordinate of the first positioning pixel and the second image comprises: determining each image grid of the second image according to the grid of the physical space and parameter information of the camera, to establish a mapping relationship between the image grid of the second image and the grid of the physical space; establishing a same-name object set of pixels of the first image and each image grid of the second image; determining an image grid corresponding to the first positioning pixel in the second image according to the first positioning pixel and the same-name object set; determining the physical space grid corresponding to the first positioning pixel according to the mapping relationship between the image grid of the second image and the grid of the physical space.
2. The method of claim 1, wherein, The determining of the image grid according to the grid of the physical space and the parameter information of the camera comprises: determining a pixel coordinate of a second positioning pixel of the image grid according to a geographical position coordinate of the grid of the physical space, a projection matrix and a geographical position coordinate of a viewpoint center of the camera, the projection matrix being used to represent a conversion relationship between pixel coordinates of pixels in the image and geographical position coordinates of grids in the physical space, the second positioning pixel being a pixel corresponding to a geometric center of the image grid, the geographical position coordinate of the viewpoint center of the camera being determined according to the parameter information of the camera.
3. The method of claim 1, wherein, The parameter information of the camera comprises a focal length, a principal point position, a video CCD size and a pose parameter, the pose parameter comprising a geographical position coordinate of the camera, a pitch angle, a roll angle and a side view angle of the camera.
4. The method according to any one of claims 1 to 3, characterized in that, The geographical position information comprises a global position code or a geographical position coordinate.
5. The method according to any one of claims 1 to 3, characterized in that, The determining of the pixel coordinate of the first positioning pixel of the target object in the first image comprises: if the type of the target object is a vehicle, determining a pixel coordinate of a reference pixel of the vehicle, a height value of the vehicle and a width value of the vehicle in the first image; determining the pixel coordinate of the first positioning pixel of the vehicle in the first image according to the pixel coordinate of the reference pixel, the height value and the width value.
6. A positioning device, characterized in that The apparatus comprises: a first positioning pixel determination unit configured to determine a first positioning pixel of a target object in a first image according to the target object; a first positioning pixel coordinate determination unit configured to determine a pixel coordinate of the first positioning pixel; a physical space grid determination unit configured to determine a physical space grid corresponding to the first positioning pixel according to the pixel coordinate of the first positioning pixel and a second image, the second image being a base image of the first image; The geographic position information determination unit is configured to determine geographic position information of the target object according to a physical space grid corresponding to the first positioning pixel; wherein the grid of the physical space has corresponding geographic position information. The physical space grid determination unit is specifically configured to determine each image grid of the second image according to the grid of the physical space and parameter information of the camera, to establish a mapping relationship between the second image grid and the grid of the physical space; establish a same-name object set of pixels of the first image and each image grid of the second image; determine the image grid corresponding to the first positioning pixel in the second image according to the first positioning pixel and the same-name object set; and determine the physical space grid corresponding to the first positioning pixel according to the mapping relationship between the image grid of the second image and the grid of the physical space.
7. The apparatus of claim 6, wherein, The image grid determination unit is configured to determine pixel coordinates of a second positioning pixel of the image grid according to geographic position coordinates of the grid of the physical space, a projection matrix, and geographic position coordinates of a viewpoint center of the camera; the projection matrix is used to represent a conversion relationship between pixel coordinates of the image and geographic position coordinates of the grid of the physical space; the second positioning pixel is a pixel corresponding to a geometric center of the image grid; and the geographic position coordinates of the viewpoint center of the camera are determined according to parameter information of the camera. The parameter information of the camera includes a focal length, a principal point position, a video CCD size, and a pose parameter; the pose parameter includes geographic position coordinates of the camera, a pitch angle, a roll angle, and a side view angle of the camera.
8. The apparatus of claim 6, wherein, The geographic position information includes a global position code or geographic position coordinates.
9. The device of any of claims 6-8, wherein, The first positioning pixel coordinate determination unit includes:
10. The device of any one of claims 6-8, wherein, A first determination subunit configured to, if the type of the target object is a vehicle, determine pixel coordinates of a reference pixel of the vehicle in the first image, a height value of the vehicle, and a width value of the vehicle; A second determination subunit configured to determine pixel coordinates of a first positioning pixel of the vehicle in the first image according to the pixel coordinates of the reference pixel, the height value, and the width value. The system includes at least one camera and a positioning device; wherein 11. A positioning system, characterized by The camera is configured to collect an image and send the image to the positioning device; The positioning device is configured to receive the image collected by the camera and perform the positioning method according to any one of claims 1-5. The system includes a processor and a memory; the memory is coupled to the processor; the memory is configured to store computer program codes; the computer program codes include computer instructions; when the processor reads the computer instructions from the memory, the electronic device performs the positioning method according to any one of claims 1-5. The computer program product includes computer instructions; when the computer instructions run on a computer, the computer performs the positioning method according to any one of claims 1-5.
12. An electronic device, comprising: 13. A computer program product, characterised in that, 14. A computer-readable storage medium, characterized in that, A computer-readable storage medium comprising computer instructions that, when executed on a computer, cause the computer to perform the positioning method of any one of claims 1-5.
15. A chip system, characterized by One or more processors that, when executing instructions, perform the positioning method of any one of claims 1-5.
Citation Information
Patent Citations
Unmanned aerial vehicle long-distance real-time positioning mapping display interconnection type control method
CN107367262A
Target tracking method and device
CN110636248A
Object positioning method and device, electronic equipment and storage medium
CN111046762A