Target object positioning method, device, equipment and storage medium
By combining target images, remote sensing images and three-dimensional point cloud data, using pre-trained depth of field estimation model and image positioning model, high-precision target object positioning is achieved, solving the problem of precise positioning of traditional positioning technology in urban areas, and reducing measurement costs and time.
Patent Information
- Application Number
- CN202510144842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Traditional positioning technology is difficult to achieve accurate detailed positioning in urban dense building areas, resulting in low accuracy of positioning results. Ground measurement equipment consumes a lot of manpower and material resources, making it difficult to complete large-scale measurements in a short time.
By acquiring the target image collected by the shooting device for the target object and the remote sensing image of the shooting device's shooting position, the three-dimensional point cloud data is obtained based on the pre-trained depth-estimation model, and the positioning result of the target object is determined based on the output offset of the image positioning model.
It improves the accuracy of positioning results, solves the problem of precise positioning of traditional positioning techniques in urban areas, and reduces measurement costs and time.
Smart Images

Figure CN119600458B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of positioning technology, and in particular to a method, device, equipment and storage medium for positioning a target object. Background Art
[0002] Target positioning technology achieves accurate positioning and identification of target objects through image analysis, processing and algorithm application, and has a wide range of applications in many fields.
[0003] Taking the application in the field of geographic information systems and remote sensing as an example, traditional positioning technology mainly relies on satellite remote sensing data and ground measurement equipment to jointly determine the spatial coordinates of the target object to achieve the positioning of the target object. However, this method also has limitations. For example, the resolution of satellite remote sensing data is limited by the satellite altitude and the performance of the imaging sensor. It is difficult to achieve accurate and detailed positioning in densely built-up areas in cities, resulting in low accuracy of positioning results. In addition, given that ground measurement equipment (such as total stations, lidar, etc.) requires a lot of manpower and material resources in the actual measurement process, it is difficult to complete large-scale measurements in a short time, which will also have an adverse effect on the positioning results.
[0004] Therefore, how to effectively improve the accuracy of positioning results is an urgent problem to be solved by those skilled in the art. Summary of the invention
[0005] The present application provides a method, apparatus, device and storage medium for locating a target object, which can effectively improve the accuracy of the positioning result when locating the target object.
[0006] The present application provides a method for locating a target object, comprising:
[0007] Acquire a target image captured by a photographing device for a target object and a remote sensing image of a photographing position where the photographing device is located;
[0008] Based on a pre-trained depth estimation model, obtaining three-dimensional point cloud data corresponding to the target image; wherein the depth estimation model is trained based on a predicted depth map corresponding to the first image sample and a laser radar data sample corresponding to the first image sample;
[0009] Based on the three-dimensional point cloud data, the target image and the remote sensing image, a positioning result of the target object is determined.
[0010] According to a method for positioning a target object provided by the present application, determining a positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image includes:
[0011] Inputting the target image and the remote sensing image into a pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; wherein the image positioning model is trained based on the predicted offset and the offset label obtained from the second image sample and the remote sensing image sample;
[0012] Determine, based on the three-dimensional point cloud data, a first distance between the target object and the shooting location in a longitude direction and a second distance between the target object and the shooting location in a latitude direction;
[0013] A positioning result of the target object is determined based on the offset, the first distance, and the second distance.
[0014] According to a method for positioning a target object provided by the present application, the offset includes a yaw angle offset, and the first distance between the target object and the shooting position in the longitude direction and the second distance between the target object and the shooting position in the latitude direction are determined based on the three-dimensional point cloud data, including:
[0015] Based on the rotation matrix and translation matrix corresponding to the shooting device, the three-dimensional point cloud data is converted into target three-dimensional point cloud data in a world coordinate system; wherein the yaw angle in the rotation matrix is determined based on the yaw angle offset;
[0016] Determine the first distance between the target object and the shooting position in the longitude direction based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data and the initial latitude of the shooting position;
[0017] The second distance between the target object and the shooting position in the latitude direction is determined based on the second point cloud data in the vertical direction in the target three-dimensional point cloud data.
[0018] According to a method for positioning a target object provided by the present application, determining the first distance between the target object and the shooting position in the longitude direction based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data and the initial latitude of the shooting position includes:
[0019] based on Determine the first distance between the target object and the shooting position in the longitude direction;
[0020] Correspondingly, determining the second distance between the target object and the shooting position in the latitude direction based on the second point cloud data in the vertical direction of the target three-dimensional point cloud data includes:
[0021] based on Determining the second distance between the target object and the shooting position in the latitude direction;
[0022] in, represents the first distance, represents the first point cloud data in the horizontal direction, represents the initial latitude of the shooting location, represents the radius of the Earth, represents the second distance, Represents second point cloud data in the vertical direction.
[0023] According to a method for positioning a target object provided by the present application, the offset includes a longitude offset and a latitude offset, and determining a positioning result of the target object based on the offset, the first distance, and the second distance includes:
[0024] Based on the first distance and the longitude offset, correct the initial longitude of the shooting position to obtain the target longitude of the target object;
[0025] Based on the second distance and the latitude offset, the initial latitude of the shooting position is corrected to obtain the target latitude where the target object is located; wherein the positioning result includes the target longitude and the target latitude.
[0026] According to a method for locating a target object provided by the present application, the method of acquiring three-dimensional point cloud data corresponding to the target image based on a pre-trained depth of field estimation model includes:
[0027] Inputting the target image into a pre-trained depth estimation model to obtain a target depth map corresponding to the target image output by the depth estimation model;
[0028] The coordinates of the pixels in the target image are converted based on the target depth map to obtain the three-dimensional point cloud data.
[0029] The present application also provides a target object positioning device, comprising:
[0030] A first acquisition unit, used to acquire a target image captured by a shooting device for a target object, and a remote sensing image of a shooting position where the shooting device is located;
[0031] A second acquisition unit is used to acquire the three-dimensional point cloud data corresponding to the target image based on a pre-trained depth estimation model; wherein the depth estimation model is trained based on the predicted depth map corresponding to the first image sample and the laser radar data sample corresponding to the first image sample;
[0032] A processing unit is used to determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image.
[0033] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a method for locating a target object as described above is implemented.
[0034] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for locating a target object as described in any one of the above is implemented.
[0035] The present application also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the method for locating a target object as described in any one of the above is implemented.
[0036] The target object positioning method, device, equipment and storage medium provided in the present application can first obtain the target image collected by the shooting device for the target object and the remote sensing image of the shooting position of the shooting device when positioning the target object; and based on the pre-trained depth of field estimation model, obtain the three-dimensional point cloud data corresponding to the target image; wherein the depth of field estimation model is obtained by training based on the predicted depth map corresponding to the first image sample and the laser radar data sample corresponding to the first image sample; and then determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image. In this way, when the target object is positioned by combining the three-dimensional point cloud data, the target image and the remote sensing image, the macroscopic geographic space background information can be obtained through the remote sensing image, and the microscopic detail perspective can be obtained through the target image collected by the shooting device. The organic combination of the three can solve the limitations of the traditional positioning technology in the prior art, thereby effectively improving the accuracy of the positioning result. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 A flowchart of a method for locating a target object provided in an embodiment of the present application.
[0039] Figure 2 A schematic diagram of a target image captured by a shooting device for a target object and a remote sensing image of a shooting position where the shooting device is located provided in an embodiment of the present application.
[0040] Figure 3 A schematic flow chart of a method for obtaining three-dimensional point cloud data corresponding to a target image based on a pre-trained depth of field estimation model provided in an embodiment of the present application.
[0041] Figure 4 A flowchart of a method for determining a positioning result of a target object based on three-dimensional point cloud data, a target image and a remote sensing image is provided in an embodiment of the present application.
[0042] Figure 5 A schematic diagram of the structure of a target object positioning device provided in an embodiment of the present application.
[0043] Figure 6 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0045] In the embodiments of the present application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In the text description of the present application, the character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0046] The technical solution provided by the embodiment of the present application can be adapted to a variety of scenarios, such as urban planning scenarios, disaster monitoring scenarios, ecological protection scenarios, geographic information systems and remote sensing scenarios, etc. In the urban planning scenario, the use of positioning technology can provide more accurate and comprehensive geographic information support to help optimize the urban spatial layout; in the disaster monitoring scenario, positioning technology can quickly obtain three-dimensional information of the disaster site, providing important reference for emergency response and rescue decisions; in the ecological protection scenario, the use of positioning technology helps to better monitor changes in the natural environment, thereby providing a scientific basis for the formulation of protection strategies.
[0047] Taking the application in the field of geographic information systems and remote sensing as an example, traditional positioning technology mainly relies on satellite remote sensing data and ground measurement equipment to jointly determine the spatial coordinates of the target object to achieve the positioning of the target object. However, this method also has limitations. For example, the resolution of satellite remote sensing data is limited by the satellite altitude and the performance of the imaging sensor. It is difficult to achieve accurate and detailed positioning in densely built-up areas in cities, resulting in low accuracy of positioning results. In addition, given that ground measurement equipment requires a lot of manpower and material resources in the actual measurement process, it is difficult to complete large-scale measurements in a short period of time, which will also have an adverse effect on the positioning results.
[0048] In order to effectively improve the accuracy of positioning results, considering that camera devices, such as smart phones, drones, etc., are generally equipped with high-resolution cameras, built-in attitude sensors and global positioning system (GPS) modules, they can capture rich image data and obtain corresponding position and attitude information, thus providing the possibility for flexible and low-cost spatial positioning methods. Therefore, this application proposes a three-dimensional positioning method based on combining mobile device photos and remote sensing images, which can be combined with the target image collected by the shooting device for the target object and the remote sensing image of the shooting position of the shooting device to jointly locate the target object, which can solve the limitations of traditional positioning technology in the prior art, thereby effectively improving the accuracy of positioning results.
[0049] It can be understood that the executor of the present method can be a photographing device, a specially set positioning device, an electronic device such as a computer or a server, or a positioning device of the target object set in the electronic device. The positioning device of the target object can be implemented by software, hardware or a combination of both, and can be specifically set according to actual needs.
[0050] The following specific embodiments will be used to describe the target object positioning method provided by the present application in detail. It is understandable that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0051] Figure 1 A flow chart of a method for locating a target object provided in an embodiment of the present application, for example, see Figure 1 As shown, the method for locating the target object may include:
[0052] S101, acquiring a target image captured by a shooting device for a target object and a remote sensing image of a shooting position where the shooting device is located.
[0053] For example, the photographing device may be a mobile photographing device or a photographing device at a fixed position, etc., and may be specifically configured according to actual needs.
[0054] For example, the target object may be a person, a vehicle or other object, and may be set according to actual needs.
[0055] For example, see Figure 2 As shown, Figure 2 A schematic diagram of a target image captured by a photographing device for a target object and a remote sensing image of a photographing position of the photographing device provided in an embodiment of the present application, wherein: Figure 2 The (1)th image is the target image captured by the shooting device for the target object. Figure 2 The second image in the figure is the remote sensing image of the shooting position of the shooting device, and the object selected by the rectangular frame is recorded as the target object. It should be noted that in Figure 2 In the figure, the starting point of the solid arrow represents the shooting position, and the direction of the arrow represents the shooting direction, so that the target object can be located in combination with the target image collected for the target object and the remote sensing image of the shooting position of the shooting device.
[0056] S102. Based on a pre-trained depth estimation model, obtain three-dimensional point cloud data corresponding to the target image; wherein the depth estimation model is trained based on a predicted depth map corresponding to the first image sample and a lidar data sample corresponding to the first image sample.
[0057] The laser radar data sample corresponding to the first image sample can be understood as the real depth map label corresponding to the first image sample. The first image sample can include the posture information of its shooting device, such as shooting angle, orientation, etc., and GPS location data of the location of the shooting device, etc., which can be set according to actual needs.
[0058] For example, when obtaining the first image sample for training the depth of field estimation model, the first image sample can be obtained from a public KITTI dataset, or from other databases, etc., and the specific settings can be made according to actual needs.
[0059] Taking the acquisition of the first image sample from the public KITTI dataset as an example, the sensor used to collect the first image sample in the KITTI dataset generally includes a forward-looking binocular camera, a lidar, a GPS or a posture sensor, etc., and the corresponding lidar data sample is obtained for training the depth estimation model. For example, the depth estimation model can adopt the DepthFormer model, etc., which can be set according to actual needs.
[0060] It can be understood that in the embodiment of the present application, the depth of field estimation model is mainly used to generate a depth map of the image, which can help understand the distance between each pixel in the image and the shooting device, and provide support for the subsequent construction of three-dimensional point cloud data. In this way, the depth map information generated by the depth of field estimation model is combined with the positioning technology to achieve high-precision three-dimensional reconstruction and positioning.
[0061] For example, in an embodiment of the present application, when training a depth estimation model based on a predicted depth map corresponding to a first image sample and a lidar data sample corresponding to the first image sample, the training process can be completed by minimizing the error between the predicted depth map corresponding to the first image sample and the lidar data sample, that is, the true depth map label corresponding to the first image sample, as shown in the following formula 1, to ensure that the depth estimation model can accurately predict the depth map in different scenarios, thereby improving the accuracy of the depth estimation model.
[0062] Formula 1
[0063] Where N represents the number of first image samples, represents the first of the N first image samples image samples, represents the model parameters of the depth estimation model, Indicates The loss function between the real depth map labels corresponding to the first image samples is used to measure the error between the real depth map labels corresponding to the first image samples. Indicates The lidar data sample corresponding to the first image sample is the real depth map label.
[0064] In combination with the above description, after the depth of field estimation model is pre-trained, the three-dimensional point cloud data corresponding to the target image can be obtained based on the depth of field estimation model, so that the three-dimensional point cloud data and the remote sensing image of the shooting position of the shooting device can be combined to achieve high-precision three-dimensional positioning of the target object, that is, execute the following S103.
[0065] S103: Determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image.
[0066] It can be seen that in the embodiment of the present application, when positioning the target object, the target image captured by the shooting device for the target object and the remote sensing image of the shooting position of the shooting device can be obtained first; and based on the pre-trained depth of field estimation model, the three-dimensional point cloud data corresponding to the target image is obtained; wherein the depth of field estimation model is trained based on the predicted depth map corresponding to the first image sample and the laser radar data sample corresponding to the first image sample; and then based on the three-dimensional point cloud data, the target image and the remote sensing image, the positioning result of the target object is determined. In this way, when the target object is positioned by combining the three-dimensional point cloud data, the target image and the remote sensing image, the macroscopic geographic space background information can be obtained through the remote sensing image, and the microscopic detail perspective can be obtained through the target image captured by the shooting device. The organic combination of the three can solve the limitations of the traditional positioning technology in the prior art, thereby effectively improving the accuracy of the positioning result.
[0067] Based on the above Figure 1 In the embodiment shown, for example, in the above S102, based on the pre-trained depth estimation model, the specific implementation of obtaining the three-dimensional point cloud data corresponding to the target image can be seen below Figure 3 The embodiment shown.
[0068] Figure 3 A schematic diagram of a method flow chart for obtaining three-dimensional point cloud data corresponding to a target image based on a pre-trained depth estimation model provided in an embodiment of the present application. For example, see Figure 3 As shown, the method may include:
[0069] S301 , inputting a target image into a pre-trained depth estimation model to obtain a target depth map corresponding to the target image output by the depth estimation model.
[0070] S302 : transforming the coordinates of the pixels in the target image based on the target depth map to obtain three-dimensional point cloud data.
[0071] Based on the target depth map corresponding to the target image output by the depth estimation model, the plane coordinates of each pixel in the target image can be converted into three-dimensional coordinates to generate three-dimensional point cloud data corresponding to the target image. For example, in an embodiment of the present application, the generation process of the three-dimensional point cloud data can be shown in the following formula 2.
[0072] Formula 2
[0073] in, Represents the three-dimensional coordinates of the pixel points in the target image, Represents the plane coordinates of the pixel points in the target image, Indicates the optical center position of the shooting device, Indicates the horizontal focal length of the camera. represents the focal length of the shooting device in the vertical direction, then the data set consisting of the three-dimensional point cloud data corresponding to the target image can be recorded as , I represents the target image, Indicates the first Pixels.
[0074] Combined with the above description, after obtaining the three-dimensional point cloud data corresponding to the target image, the positioning result of the target object can be determined based on the three-dimensional point cloud data, the target image and the remote sensing image. In this way, when the target object is positioned by combining the three-dimensional point cloud data, the target image and the remote sensing image, the macroscopic geographic spatial background information can be obtained through the remote sensing image, and the microscopic detail perspective can be obtained through the target image captured by the shooting device. The organic combination of the three can solve the limitations of the traditional positioning technology in the prior art, thereby effectively improving the accuracy of the positioning result.
[0075] Based on any of the above embodiments, for example, in the above S103, the specific implementation of determining the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image can be seen below. Figure 4 The embodiment shown.
[0076] Figure 4 A flowchart of a method for determining a positioning result of a target object based on three-dimensional point cloud data, a target image and a remote sensing image provided in an embodiment of the present application is provided. For example, see Figure 4 As shown, the method may include:
[0077] S401. Input the target image and the remote sensing image into a pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; wherein the image positioning model is trained based on the predicted offset and offset label obtained from the second image sample and the remote sensing image sample.
[0078] For example, the offset may include longitude offset, latitude offset and yaw angle offset, which may be set according to actual needs. The longitude offset may be expressed as , the latitude offset can be expressed as , the yaw angle offset can be expressed as .
[0079] For example, when obtaining the second image sample for training the image localization model, the second image sample can be obtained from a public KITTI dataset, or from other databases, etc., and the specific settings can be made according to actual needs.
[0080] Taking the acquisition of the second image sample from the public KITTI dataset as an example, the sensors used to collect the second image sample in the KITTI dataset usually include a forward-looking binocular camera, a lidar, a GPS or a posture sensor, etc., and the corresponding high-resolution remote sensing image can be downloaded according to the longitude and latitude of the second image sample in the KITTI dataset as its corresponding remote sensing image sample for training the image positioning model. For example, the image positioning model can use the CNN+Transformer fusion strategy to build a regression model, etc., which can be set according to actual needs. In this way, by combining the image positioning model and the depth of field estimation model, the generalization ability of the model and the high-precision positioning ability in different scenarios can be ensured.
[0081] It is understandable that in the embodiment of the present application, the image positioning model is mainly used to determine the spatial relationship between the image and the remote sensing image, and then correct the shooting position of the shooting device, so as to improve the positioning accuracy. In order to achieve this goal, the features of the image and the remote sensing image can be extracted through the feature extraction network in the image positioning model, and the extracted features can be matched to obtain the offset of the shooting position of the shooting device, thereby providing support for the subsequent construction of three-dimensional point cloud data.
[0082] For example, in an embodiment of the present application, when training the image positioning model based on the predicted offset and offset label obtained based on the second image sample and the remote sensing image sample, it can be completed during the training process by minimizing the error between the predicted offset and the offset label obtained based on the second image sample and the remote sensing image sample, as shown in the following formula 3, to ensure that the image positioning model can accurately predict the offset of the shooting position of the shooting device in different scenarios, thereby improving the accuracy of the image positioning model.
[0083] Formula 3
[0084] Wherein, M represents the number of second image samples, j represents the jth image sample among the M second image samples, represents the model parameters of the image localization model, represents the loss function between the predicted offset and the offset label obtained by the j-th second image sample and the remote sensing image sample, which is used to measure the error between the predicted offset and the offset label. They represent the longitude offset label, latitude offset label and yaw angle offset label corresponding to the j second image samples, respectively. They respectively represent the predicted longitude offset, predicted latitude offset and predicted yaw angle offset obtained from the j-th second image sample and remote sensing image sample.
[0085] S402: Determine a first distance between the target object and the shooting position in the longitude direction and a second distance between the target object and the shooting position in the latitude direction based on the three-dimensional point cloud data.
[0086] For example, in the case where the offset includes a yaw angle offset, a first distance between the target object and the shooting position in the longitude direction, and a second distance between the target object and the shooting position in the latitude direction are determined respectively based on the three-dimensional point cloud data. The three-dimensional point cloud data can be first converted into target three-dimensional point cloud data in the world coordinate system based on the rotation matrix and translation matrix corresponding to the shooting device; wherein the yaw angle in the rotation matrix is determined based on the yaw angle offset; and based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data, and the initial latitude of the shooting position, the first distance between the target object and the shooting position in the longitude direction is determined; and based on the second point cloud data in the vertical direction in the target three-dimensional point cloud data, the second distance between the target object and the shooting position in the latitude direction is determined.
[0087] For example, in an embodiment of the present application, when converting the three-dimensional point cloud data into target three-dimensional point cloud data in the world coordinate system based on the rotation matrix and translation matrix corresponding to the shooting device, please refer to the following formula 4 for details.
[0088] Formula 4
[0089] in, Represents the rotation matrix corresponding to the camera device, which is a 3 3 matrix, Represents the translation matrix corresponding to the camera device, which is a 3 1 matrix, Represents the three-dimensional coordinates of the pixel points in the target image, Represents the target 3D point cloud coordinates in the world coordinate system.
[0090] It is understandable that the rotation matrix can usually be derived from the posture information of the camera, such as yaw angle, pitch angle and roll angle, to ensure that each pixel point is properly positioned in the world coordinates. The yaw angle used in can be determined based on the yaw angle offset predicted by the above-mentioned image positioning model, as shown in the following formula 5.
[0091] Formula 5
[0092] in, Represents the rotation matrix The yaw angle used in Indicates the yaw angle obtained by the camera. Represents the yaw angle offset of the shooting position output by the image localization model.
[0093] For example, in the embodiment of the present application, based on the first point cloud data in the horizontal direction of the target three-dimensional point cloud data and the initial latitude of the shooting position, when determining the first distance between the target object and the shooting position in the longitude direction, it can be based on A first distance between the target object and the shooting location in the longitude direction is determined.
[0094] Correspondingly, when determining the second distance between the target object and the shooting position in the latitude direction based on the second point cloud data in the vertical direction of the target three-dimensional point cloud data, the second distance between the target object and the shooting position in the latitude direction can be determined based on Determine a second distance between the target object and the shooting position in the latitude direction. represents the first distance, Represents the first point cloud data in the horizontal direction, Indicates the initial latitude of the shooting location. represents the radius of the Earth, represents the second distance, Indicates the second point cloud data in the vertical direction.
[0095] It is worth noting that in the embodiments of the present application, , and The units need to remain consistent.
[0096] After the offset, the first distance, and the second distance are determined respectively, the following S403 may be executed.
[0097] S403: Determine a positioning result of the target object based on the offset, the first distance, and the second distance.
[0098] By way of example, in the case where the offset includes a longitude offset and a latitude offset, when determining the positioning result of the target object based on the offset, the first distance, and the second distance, the initial longitude of the shooting location can be corrected based on the first distance and the longitude offset to obtain the target longitude where the target object is located; and the initial latitude of the shooting location can be corrected based on the second distance and the latitude offset to obtain the target latitude where the target object is located; wherein the positioning result includes the target longitude and the target latitude.
[0099] For example, when correcting the initial longitude of the shooting location based on the first distance and the longitude offset, refer to the following formula 6.
[0100] Formula 6
[0101] in, Indicates the target longitude where the target object is located. Indicates the initial longitude of the shooting location. represents the first distance, Represents the longitude offset output by the image localization model.
[0102] For example, when correcting the initial latitude of the shooting location based on the second distance and the latitude offset, refer to the following formula 7.
[0103] Formula 7
[0104] in, Indicates the target latitude where the target object is located. Indicates the initial latitude of the shooting location. represents the second distance, Represents the latitude offset output by the image localization model.
[0105] In combination with the above description, the target image and the remote sensing image are first input into the pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; and based on the three-dimensional point cloud data, the first distance between the target object and the shooting position in the longitude direction and the second distance between the target object and the shooting position in the latitude direction are respectively determined; and then the positioning result of the target object is determined based on the offset, the first distance and the second distance. In this way, when the offset obtained by combining the target image and the remote sensing image, the first distance between the target object and the shooting position in the longitude direction and the second distance between the target object and the shooting position in the latitude direction are used to locate the target object, the macroscopic geographic space background information can be obtained through the remote sensing image, and the microscopic detail perspective can be obtained through the target image captured by the shooting device, which can solve the limitations of the traditional positioning technology in the prior art, thereby effectively improving the accuracy of the positioning result.
[0106] The target object positioning device provided in the present application is described below. The target object positioning device described below and the target object positioning method described above can be referenced to each other.
[0107] Figure 5 A schematic diagram of a target object positioning device provided in an embodiment of the present application, for example, see Figure 5 As shown, the target object positioning device 50 may include:
[0108] The first acquisition unit 501 is used to acquire a target image captured by a shooting device for a target object and a remote sensing image of a shooting position where the shooting device is located;
[0109] A second acquisition unit 502 is used to acquire the three-dimensional point cloud data corresponding to the target image based on a pre-trained depth estimation model; wherein the depth estimation model is trained based on the predicted depth map corresponding to the first image sample and the laser radar data sample corresponding to the first image sample;
[0110] The processing unit 503 is used to determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image.
[0111] For example, in the embodiment of the present application, the processing unit 503 is used to determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image, including:
[0112] Inputting the target image and the remote sensing image into a pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; wherein the image positioning model is trained based on the predicted offset and the offset label obtained from the second image sample and the remote sensing image sample;
[0113] Determine, based on the three-dimensional point cloud data, a first distance between the target object and the shooting location in a longitude direction and a second distance between the target object and the shooting location in a latitude direction;
[0114] A positioning result of the target object is determined based on the offset, the first distance, and the second distance.
[0115] For example, in the embodiment of the present application, the offset includes a yaw angle offset, and the processing unit 503 is used to determine, based on the three-dimensional point cloud data, a first distance between the target object and the shooting position in the longitude direction and a second distance between the target object and the shooting position in the latitude direction, including:
[0116] Based on the rotation matrix and translation matrix corresponding to the shooting device, the three-dimensional point cloud data is converted into target three-dimensional point cloud data in a world coordinate system; wherein the yaw angle in the rotation matrix is determined based on the yaw angle offset;
[0117] Determine the first distance between the target object and the shooting position in the longitude direction based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data and the initial latitude of the shooting position;
[0118] The second distance between the target object and the shooting position in the latitude direction is determined based on the second point cloud data in the vertical direction in the target three-dimensional point cloud data.
[0119] For example, in the embodiment of the present application, the processing unit 503 is configured to determine the first distance between the target object and the shooting position in the longitude direction based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data and the initial latitude of the shooting position, including:
[0120] based on Determine the first distance between the target object and the shooting position in the longitude direction;
[0121] Correspondingly, determining the second distance between the target object and the shooting position in the latitude direction based on the second point cloud data in the vertical direction of the target three-dimensional point cloud data includes:
[0122] based on Determining the second distance between the target object and the shooting position in the latitude direction;
[0123] in, represents the first distance, represents the first point cloud data in the horizontal direction, represents the initial latitude of the shooting location, represents the radius of the Earth, represents the second distance, Represents second point cloud data in the vertical direction.
[0124] For example, in an embodiment of the present application, the offset includes a longitude offset and a latitude offset, and the processing unit 503 is used to determine the positioning result of the target object based on the offset, the first distance, and the second distance, including:
[0125] Based on the first distance and the longitude offset, correct the initial longitude of the shooting position to obtain the target longitude of the target object;
[0126] Based on the second distance and the latitude offset, the initial latitude of the shooting position is corrected to obtain the target latitude where the target object is located; wherein the positioning result includes the target longitude and the target latitude.
[0127] For example, in the embodiment of the present application, the second acquisition unit 502 is used to acquire the three-dimensional point cloud data corresponding to the target image based on the pre-trained depth of field estimation model, including:
[0128] Inputting the target image into a pre-trained depth estimation model to obtain a target depth map corresponding to the target image output by the depth estimation model;
[0129] The coordinates of the pixels in the target image are converted based on the target depth map to obtain the three-dimensional point cloud data.
[0130] The target object positioning device 50 provided in the embodiment of the present application can execute the technical solution of the target object positioning method in any of the above-mentioned embodiments, and its implementation principle and beneficial effects are similar to the implementation principle and beneficial effects of the target object positioning method. Please refer to the implementation principle and beneficial effects of the target object positioning method, which will not be repeated here.
[0131] Figure 6 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a method for positioning a target object, the method comprising: obtaining a target image captured by a shooting device for a target object and a remote sensing image of a shooting position where the shooting device is located; based on a pre-trained depth estimation model, obtaining three-dimensional point cloud data corresponding to the target image; wherein the depth estimation model is obtained by training based on a predicted depth map corresponding to a first image sample and a laser radar data sample corresponding to the first image sample; based on the three-dimensional point cloud data, the target image and the remote sensing image, determining the positioning result of the target object.
[0132] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0133] On the other hand, the present application also provides a computer program product, which includes a computer program, and the computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target object positioning method provided by the above-mentioned methods, and the method includes: obtaining a target image captured by a shooting device for the target object, and a remote sensing image of the shooting position where the shooting device is located; based on a pre-trained depth of field estimation model, obtaining three-dimensional point cloud data corresponding to the target image; wherein the depth of field estimation model is trained based on a predicted depth map corresponding to a first image sample and a lidar data sample corresponding to the first image sample; based on the three-dimensional point cloud data, the target image and the remote sensing image, determining the positioning result of the target object.
[0134] On the other hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the target object positioning method provided by the above-mentioned methods, the method comprising: obtaining a target image captured by a shooting device for the target object, and a remote sensing image of a shooting position at which the shooting device is located; based on a pre-trained depth of field estimation model, obtaining three-dimensional point cloud data corresponding to the target image; wherein the depth of field estimation model is trained based on a predicted depth map corresponding to a first image sample and a lidar data sample corresponding to the first image sample; and determining the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image.
[0135] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for locating a target object, characterized in that: include: Acquire a target image captured by a photographing device for a target object and a remote sensing image of a photographing position where the photographing device is located; Based on a pre-trained depth estimation model, obtaining three-dimensional point cloud data corresponding to the target image; wherein the depth estimation model is trained based on a predicted depth map corresponding to the first image sample and a laser radar data sample corresponding to the first image sample; Based on the three-dimensional point cloud data, the target image and the remote sensing image, the positioning result of the target object is determined; wherein the target image and the remote sensing image are input into a pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; wherein the image positioning model is trained based on the predicted offset and the offset label obtained from the second image sample and the remote sensing image sample; based on the three-dimensional point cloud data, a first distance between the target object and the shooting position in the longitude direction and a second distance between the target object and the shooting position in the latitude direction are determined respectively; based on the offset, the first distance and the second distance, the positioning result of the target object is determined.
2. The method for locating a target object according to claim 1, characterized in that: The offset includes a yaw angle offset, and determining a first distance between the target object and the shooting position in a longitude direction and a second distance between the target object and the shooting position in a latitude direction based on the three-dimensional point cloud data respectively includes: Based on the rotation matrix and translation matrix corresponding to the shooting device, the three-dimensional point cloud data is converted into target three-dimensional point cloud data in a world coordinate system; wherein the yaw angle in the rotation matrix is determined based on the yaw angle offset; Determine the first distance between the target object and the shooting position in the longitude direction based on the first point cloud data in the horizontal direction in the target three-dimensional point cloud data and the initial latitude of the shooting position; The second distance between the target object and the shooting position in the latitude direction is determined based on the second point cloud data in the vertical direction in the target three-dimensional point cloud data.
3. The method for locating a target object according to claim 2, characterized in that: The determining, based on the first point cloud data in the horizontal direction of the target three-dimensional point cloud data and the initial latitude of the shooting position, the first distance between the target object and the shooting position in the longitude direction comprises: based on Determine the first distance between the target object and the shooting position in the longitude direction; Correspondingly, determining the second distance between the target object and the shooting position in the latitude direction based on the second point cloud data in the vertical direction of the target three-dimensional point cloud data includes: based on Determining the second distance between the target object and the shooting position in the latitude direction; in, represents the first distance, represents the first point cloud data in the horizontal direction, represents the initial latitude of the shooting location, represents the radius of the Earth, represents the second distance, Represents second point cloud data in the vertical direction.
4. The method for locating a target object according to claim 1, characterized in that: The offset includes a longitude offset and a latitude offset, and determining a positioning result of the target object based on the offset, the first distance, and the second distance includes: Based on the first distance and the longitude offset, correct the initial longitude of the shooting position to obtain the target longitude of the target object; Based on the second distance and the latitude offset, the initial latitude of the shooting position is corrected to obtain the target latitude where the target object is located; wherein the positioning result includes the target longitude and the target latitude.
5. The method for locating a target object according to any one of claims 1 to 3, characterized in that: The step of obtaining three-dimensional point cloud data corresponding to the target image based on a pre-trained depth estimation model includes: Inputting the target image into a pre-trained depth estimation model to obtain a target depth map corresponding to the target image output by the depth estimation model; The coordinates of the pixels in the target image are converted based on the target depth map to obtain the three-dimensional point cloud data.
6. A device for locating a target object, characterized in that: include: A first acquisition unit, used to acquire a target image captured by a shooting device for a target object, and a remote sensing image of a shooting position where the shooting device is located; A second acquisition unit is used to acquire the three-dimensional point cloud data corresponding to the target image based on a pre-trained depth estimation model; wherein the depth estimation model is trained based on the predicted depth map corresponding to the first image sample and the laser radar data sample corresponding to the first image sample; A processing unit is used to determine the positioning result of the target object based on the three-dimensional point cloud data, the target image and the remote sensing image; wherein the target image and the remote sensing image are input into a pre-trained image positioning model to obtain the offset of the shooting position output by the image positioning model; wherein the image positioning model is trained based on the predicted offset and offset label obtained from the second image sample and the remote sensing image sample; based on the three-dimensional point cloud data, a first distance between the target object and the shooting position in the longitude direction and a second distance between the target object and the shooting position in the latitude direction are determined respectively; based on the offset, the first distance and the second distance, the positioning result of the target object is determined.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for locating a target object as claimed in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for locating a target object as claimed in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for locating a target object as claimed in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Laser and camera fused target object identification and positioning method
CN112598729A
Vector retrieval positioning method and system for panoramic image of outdoor scene and medium
CN117292268A