Camera point location determination method, device and equipment based on three-dimensional map

By creating a virtual camera in a three-dimensional map and matching image features, the actual camera points are automatically determined, which solves the problems of low efficiency and low accuracy of manual labeling in the prior art, and efficient and accurate camera point labeling is achieved.

CN120279220AActive Publication Date: 2025-07-08ZHEJIANG UNIVIEW TECH CO LTD +1

Patent Information

Application Number
CN202510757379.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The camera point labeling method in the existing three-dimensional map relies on manual labor, resulting in low labeling efficiency and low accuracy and consistency.

Method used

By creating a virtual camera in a three-dimensional map, the points of the actual camera are automatically determined using the image feature extraction and matching collected by the actual camera and the virtual camera, including the implementation of feature extraction, image matching and point determination modules.

Benefits of technology

It improves the accuracy and consistency of camera point labeling, reduces human errors, and realizes efficient labeling in real-time and large-scale data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279220A_ABST
    Figure CN120279220A_ABST
Patent Text Reader

Abstract

The invention discloses a camera point location determination method, device and equipment based on a three-dimensional map. The method comprises the following steps: determining an actual image collected by an actual camera deployed in an actual scene, and determining a plurality of first virtual images collected by a plurality of virtual cameras corresponding to the actual camera; performing feature extraction on the actual image and each first virtual image to obtain key point information of each image; determining an image matching result between the actual image and each first virtual image according to a matching result of the key point information of the actual image and each first virtual image; and determining a second virtual image from the first virtual images according to an image matching result, and determining a point location determination result of the actual camera according to the second virtual image. According to the method, the accuracy and the consistency of point location determination are remarkably improved, so that the possibility of manual marking errors is reduced; and moreover, the point position marking speed is increased, and real-time marking is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method, apparatus, and device for determining camera positions based on a three-dimensional map. Background Art

[0002] In recent years, the applications of three-dimensional maps in the fields of geographic information systems (GIS), urban planning, unmanned driving, and virtual reality have been increasing continuously. These applications require accurate annotation of camera positions in the three-dimensional map for further data analysis and decision support.

[0003] Existing position annotation methods mostly rely on manual annotation, which is not only inefficient but also easily affected by human factors, resulting in low accuracy and consistency of annotation. Summary of the Invention

[0004] The present invention provides a method, apparatus, and device for determining camera positions based on a three-dimensional map to solve the problems of low accuracy and low efficiency in annotating camera positions in the three-dimensional map.

[0005] According to one aspect of the present invention, there is provided a method for determining camera positions based on a three-dimensional map, including:

[0006] Determining actual images collected by an actual camera deployed in an actual scene, and determining a plurality of first virtual images collected by a plurality of virtual cameras corresponding to the actual camera; wherein the virtual cameras are created in a three-dimensional map corresponding to the actual scene according to the camera parameters of the actual camera;

[0007] Performing feature extraction on the actual images and each of the first virtual images to obtain key point information of each image;

[0008] Determining an image matching result between the actual image and each of the first virtual images according to a matching result of the key point information of the actual image and each of the first virtual images;

[0009] Determining a second virtual image from each of the first virtual images according to the image matching result, and determining a position determination result of the actual camera according to the second virtual image.

[0010] According to another aspect of the present invention, there is provided a device for determining camera positions based on a three-dimensional map, including:

[0011] An image determination module, configured to determine actual images collected by an actual camera deployed in an actual scene, and determine a plurality of first virtual images collected by a plurality of virtual cameras corresponding to the actual camera; wherein the virtual cameras are created in a three-dimensional map corresponding to the actual scene according to the camera parameters of the actual camera;

[0012] A key point extraction module, configured to extract features from the actual image and each of the first virtual images to obtain key point information of each image;

[0013] An image matching module, configured to determine an image matching result between the actual image and each of the first virtual images according to a matching result of the key point information of the actual image and each of the first virtual images;

[0014] A point position determination module, configured to determine a second virtual image from each of the first virtual images according to the image matching result, and determine a point position determination result of the actual camera according to the second virtual image.

[0015] According to another aspect of the present invention, there is provided an electronic device, including:

[0016] At least one processor; and a memory communicatively connected to the at least one processor;

[0017] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the camera point position based on a three-dimensional map according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the method for determining the camera point position based on a three-dimensional map according to any embodiment of the present invention when executed.

[0019] The technical solution of this embodiment creates corresponding virtual cameras according to actual cameras, and obtains the camera point position determination result through the detection and matching results of key points in the images collected by each camera. According to this camera point position determination result, the actual camera can be marked in the three-dimensional map, significantly improving the accuracy and consistency of point position determination, thereby reducing the possibility of manual marking errors; and through the technical solution of this embodiment, the marking speed can be improved to achieve real-time marking to adapt to real-time and large-scale data scenarios.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0022] Figure 1 is a flowchart of a method for determining camera positions based on a three-dimensional map according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart of another method for determining camera positions based on a three-dimensional map according to an embodiment of the present invention;

[0024] Figure 3 is a schematic diagram of the model architecture of a key point extraction model according to an embodiment of the present invention;

[0025] Figure 4 is a flowchart of another method for determining camera positions based on a three-dimensional map according to an embodiment of the present invention;

[0026] Figure 5 is an architecture diagram of a graph neural network according to an embodiment of the present invention;

[0027] Figure 6 is a flowchart of another method for determining camera positions based on a three-dimensional map according to an embodiment of the present invention;

[0028] Figure 7 is a schematic diagram of the structure of a device for determining camera positions based on a three-dimensional map according to an embodiment of the present invention;

[0029] Figure 8 is a schematic diagram of the structure of an electronic device for implementing the method for determining camera positions based on a three-dimensional map in the embodiments of the present invention. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention shall fall within the protection scope of the present invention.

[0031] It should be noted that in the description of the present invention, the terms "candidate", "target", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Figure 1 The embodiment of the present invention provides a flowchart of a method for determining a camera position based on a three-dimensional map. This embodiment is applicable to the situation where the accurate position of camera deployment in the actual scene is uncertain, and accurately marks the camera position in the three-dimensional map corresponding to the actual scene. This method can be executed by a device for determining a camera position based on a three-dimensional map. The device for determining a camera position based on a three-dimensional map can be implemented in the form of hardware and / or software, and can be configured in a server with communication and computing capabilities. As Figure 1 shown, the method includes:

[0033] S110. Determine the actual images collected by the actual cameras deployed in the actual scene, and determine the multiple first virtual images collected by the multiple virtual cameras corresponding to the actual cameras.

[0034] The actual scene is a real platform where camera positions are deployed, such as a shopping mall or an industrial park. In order to achieve monitoring coverage of these actual scenes, monitoring cameras will be deployed at various positions in the scene. And in order to facilitate the monitoring of the environment in the actual scene, a corresponding three-dimensional map will be established for the actual scene. This three-dimensional map is constructed according to the real scenery in the actual scene, that is, the three-dimensional map at least includes all the fixed objects in the actual scene; at the same time, in order to facilitate the management of the monitoring cameras in the actual scene, virtual cameras will be set at the places corresponding to the actual cameras in the three-dimensional map. However, in order to ensure that the field of view of the virtual camera is the same as that of the corresponding actual camera, it is necessary to ensure the accuracy of the creation position information and rotation information of the virtual camera in the three-dimensional map. Therefore, it is necessary to accurately obtain the specific position information of the actual camera, and create a virtual camera based on the accurate specific position information of the actual camera, which ensures the accuracy and efficiency of managing the virtual camera corresponding to the actual camera through the three-dimensional map.

[0035] Among them, the virtual camera is created in the three-dimensional map corresponding to the actual scene according to the camera parameters of the actual camera. Since the position information of the actual camera cannot be obtained initially, only the rough position information of the actual camera can be obtained, that is, the approximate position information is a regional range, and the virtual camera corresponding to the actual camera is determined according to the rough position information.

[0036] Specifically, the camera parameters of the actual camera at least include the regional range corresponding to the camera deployment position in the three-dimensional map and the camera internal parameters. Traverse within this regional range to obtain multiple candidate position information. Multiple virtual cameras are created according to the multiple candidate position information and multiple candidate rotation angle information. At the same time, the internal parameters of the virtual camera are the same as those of the actual camera. The candidate rotation angles are determined according to the sampling results of the angle range that the camera can rotate. Exemplarily, different combinations of the candidate position information and the candidate rotation angle information are determined according to the multiple candidate position information and the multiple candidate rotation angle information, and the virtual camera corresponding to each combination of the candidate position information and the candidate rotation angle information is created. That is, the multiple virtual cameras created are a set of various possible results of the position where the actual camera is deployed.

[0037] The camera parameters of the actual camera also include the camera IP information. The remote access and control of the actual camera can be realized through the camera IP information to collect the actual images of the actual camera in the actual scene. Render the virtual images collected by each created virtual camera according to the position information and rotation angle information of the virtual camera in the three-dimensional map. Each virtual camera collects a corresponding first virtual image.

[0038] Exemplarily, a plurality of actual cameras are included in an actual scenario. Therefore, a camera data table for all actual cameras is established to manage and schedule relevant information of each actual camera according to the camera data table. The camera data table at least includes camera IP information. Each actual camera has an independent IP address, and remote access and control of the actual camera can be achieved through the IP address of the actual camera, so as to obtain real-time actual images. The camera data table also includes the area range corresponding to the camera deployment location in the three-dimensional map. To improve the calculation efficiency and reduce unnecessary resource consumption, the cameras are divided according to the areas where they are installed to obtain the area range of the actual camera in the actual scenario and the corresponding area range in the three-dimensional map. Through this area range, the rendering range of the virtual camera can be restricted and reduced, avoiding excessive calculation burden caused by unnecessary global rendering. The camera data table also includes camera internal parameters. Each camera has different internal parameters, and the internal parameters directly affect the creation and matching of the virtual camera. The internal parameter information includes field of view, aspect ratio, focal length, and aperture information. After the camera data table is prepared, real-time actual images of each actual camera are automatically obtained in batches through a Python script. The specific operation steps are as follows: Using the Python script, by sequentially accessing the IP addresses of the actual cameras, the script automatically connects to the cameras to retrieve the image data captured in real time as the actual images of the actual cameras. For each actual camera in the camera data table, the script loops through all the cameras, sequentially obtains each actual image and stores it in a local folder. When determining the point position information of each actual camera, the actual images in the folder are obtained correspondingly in sequence.

[0039] Generate a series of movement parameters for the virtual camera according to the area range of the actual camera whose point position is to be determined currently. The movement parameters specifically include position information and rotation angle information. And according to the internal parameter information of the actual camera in the camera data table, create a virtual camera corresponding to the actual camera in the UE (Unreal Engine) model to ensure that the camera internal parameters of the virtual camera are the same as those of the actual camera. Among them, there are two situations in the UE model of the three-dimensional map. One is that there are built-in point positions in the UE model. The built-in point positions are the possible deployment point positions of the actual cameras pre-set in the UE model. In this case, only the corresponding relationship and viewing angle between the actual camera and the virtual camera need to be determined. The movement parameters of the virtual camera only need to set the viewing angle, that is, the rotation angle information. The viewing angle setting will be adjusted in steps of a preset unit angle to generate different viewing angle combinations by rotating the angle around the Y axis and the angle around the Z axis. The viewing angle change formula is shown as follows: . Among them, is the rotation angle vector, and is an integer representing the number of rotation steps, a represents the preset unit angle. In another case, if there is no built-in position point in the UE model, it is necessary to dynamically generate the movement parameters of the virtual camera according to the actual area range where the camera is located. Specifically, the position parameters (X, Y, Z) of the virtual camera change from the initial position ( in the limited area range each time with the preset distance unit as the step length. The initial position can be any position within the actual area range where the camera is located. At the same time, the rotation angle around the Y-axis and the rotation angle around the Z-axis change by 60 degrees each time to ensure coverage of different perspectives and positions. The position change formula is as follows: . Among them, is the position vector of the virtual camera, , , represent the position steps, b is the preset distance unit, which can be set according to the actual scenario, and the specific value is not limited here. For example, b = 0.2 meters.

[0040] Since the camera data table stores the area range corresponding to the camera, and there may be at least two actual cameras within the same area range, in order to improve the rendering efficiency of the virtual camera, the cameras are divided according to the area where they are located. The virtual cameras are divided according to the area range of the cameras in the camera data table to obtain the virtual camera sub-table corresponding to each area range. In each virtual camera sub-table, the virtual camera sub-table is further divided according to the camera type and camera internal parameters to obtain the first virtual camera sub-table corresponding to each camera type and camera internal parameter. In this first virtual camera sub-table, the first virtual images corresponding to each virtual camera are batch-rendered and stored.

[0041] S120. Extract the features of the actual image and each first virtual image to obtain the key point information of each image.

[0042] Among them, the key point is the recognizable feature point in the image. For example, the key point is the point of interest. Extract the features of the actual image to obtain the actual key point information in the actual image, and extract the features of each first virtual image to obtain the virtual key point information in each virtual image.

[0043] Exemplarily, pre-train the point of interest extraction model, and input the actual image and each first virtual image into the point of interest extraction model respectively to obtain the key point information in each image. The key point information includes at least the key point position information and may also include the key point category information.

[0044] S130. Determine the image matching result between the actual image and each first virtual image according to the matching result of the key point information of the actual image and each first virtual image.

[0045] Perform key-point information matching on the actual key points in the actual image and the virtual key points in each first virtual image respectively. Determine the image matching result between the actual image and the first virtual image according to the key-point information matching degree between the key points in the actual image and the first virtual image. Among them, the image matching result is the key-point matching degree between the actual image and the first virtual image. The higher the key-point matching degree, the higher the similarity between the actual image and the first virtual image, that is, the higher the matching degree between the virtual camera corresponding to the first virtual image and the actual camera.

[0046] Specifically, traverse and match the actual key points and the virtual key points according to the key-point position information in the actual image and the first virtual image, determine the number of key points with successful position matching, and determine the image matching result between the actual image and the first virtual image according to the ratio of the number of key points with successful matching to the number of actual key points in the actual image. Among them, successful position matching means that the gap between the position of the virtual key point and the position of the actual key point is less than the preset gap threshold. Further, perform further matching judgment on the key points with successful position matching according to the key-point type information, determine the number of key points with both successful position and type matching, and determine the image matching result between the actual image and the first virtual image according to the ratio of the number of key points with both successful position and type matching to the number of actual key points in the actual image. Among them, successful type matching means that the type of the virtual key point is the same as the type of the actual key point.

[0047] Exemplarily, pre-train a key-point matching model, input the actual image, each first virtual image and the key-point information into the key-point matching model, and the output of the model is the image matching result between the actual image and each first virtual image.

[0048] S140. Determine a second virtual image from each first virtual image according to the image matching result, and determine the point position determination result of the actual camera according to the second virtual image.

[0049] Determine the first virtual image with the highest matching degree from all the first virtual images as the second virtual image according to the image matching result, and determine the point position determination result of the actual camera according to the movement parameters corresponding to the creation of the virtual camera corresponding to the second virtual image.

[0050] Specifically, the image matching result represents the image similarity. Sort the first virtual images in descending order according to the image similarity, determine the first virtual image ranked first as the second virtual image, determine the position information and rotation angle information of the virtual camera corresponding to the second virtual image at the time of creation, and use the position information and rotation angle information of the virtual camera as the point position determination result of the actual camera.

[0051] Optionally, the image matching result represents the image similarity. Sort the first virtual images in descending order according to the image similarity, and select a preset number of the first virtual images ranked at the front as the second virtual images. Determine the virtual cameras corresponding to the second virtual images as the second virtual cameras. Statistically analyze the position information of each second virtual camera, and use the position information of the second virtual camera with the largest proportion as the target position information of the actual camera. And determine the second virtual camera with the highest image similarity among the second virtual cameras corresponding to the target position information as the target virtual camera, and determine the rotation angle information of the target virtual camera as the target rotation angle information of the actual camera. Finally, use the target position information and the target rotation angle information as the point position determination result of the actual camera.

[0052] The technical solution of this embodiment creates corresponding virtual cameras according to the actual camera, and obtains the camera point position determination result based on the detection and matching results of the key points in the images collected by each camera. According to this camera point position determination result, the actual camera can be marked on the 3D map, significantly improving the accuracy and consistency of point position determination, thereby reducing the possibility of manual marking errors; and through the technical solution of this embodiment, the point position marking speed can be improved to achieve real-time marking to adapt to real-time and large-scale data scenarios.

[0053] The implementation of the technical solution of this embodiment provides a more intelligent and efficient 3D map point position marking solution, enhancing the practicality and analysis ability of 3D map data, promoting the application of 3D maps in various fields, and enhancing decision support and data analysis capabilities. Furthermore, it promotes the development and innovation of related industries and can effectively support various applications, such as applications in the fields of unmanned driving, urban planning, and virtual reality.

[0054] Figure 2 This is a flowchart of another method for determining the camera point position based on a 3D map provided by an embodiment of the present invention. In this embodiment, the extraction of key point information in the above embodiment is further refined. As Figure 2 shown, the method includes:

[0055] S210. Determine the actual image collected by the actual camera deployed in the actual scene, and determine the multiple first virtual images collected by multiple virtual cameras corresponding to the actual camera.

[0056] S220. Based on the pre-trained key point extraction model, perform feature extraction on the actual image and each first virtual image to obtain the key point detection results and corresponding descriptors of each image as the key point information.

[0057] The key point extraction model is a pre-trained model for extracting key point features from images. In order to improve the accuracy of key point matching, the output of the key point extraction model includes, in addition to the detection results of the key points themselves in the image, the descriptors corresponding to each key point. The descriptors are used to feature-describe the overall information of the local area where each key point is located, that is, the descriptors are used to characterize the overall features of the local area where the corresponding key points are located.

[0058] Specifically, the actual image and each first virtual image are respectively input into the key point extraction model, and the key point extraction model extracts features from each image. The output of the model is the detection result of each key point and the corresponding descriptor. Among them, the key point detection result includes at least the key point position information, and the descriptor is the regional feature description information of the local area formed by expanding a preset range centered on the key point position information. This regional feature description information can be the compression result of the feature extraction result of the key point extraction model for this local area. Exemplarily, the key point extraction model can adopt the PointNet network model.

[0059] In a feasible embodiment, the key point extraction model includes a shared feature encoder, a key point detection decoder, and a descriptor generation decoder; the output of the shared feature encoder is the input of the key point detection decoder and the descriptor generation decoder, the output of the key point detection decoder is the key point detection result, and the output of the descriptor generation decoder is the descriptors corresponding to each key point;

[0060] Among them, the shared feature encoder includes at least one convolutional layer and at least one FasterNet block module; the key point detection decoder includes at least one convolutional layer and at least one FasterNet block module; the descriptor generation decoder includes a feature synthesis layer, at least one convolutional layer, and at least one FasterNet block module.

[0061] The key point extraction model consists of three parts: a shared feature encoder, a key point detection decoder, and a descriptor generation decoder. Among them, the shared feature encoder is used to extract the feature details in the image, the key point detection decoder is used to determine the key point detection result according to the result output by the shared feature encoder, and the descriptor generation decoder is used to determine the feature information of the local area where each key point is located as the descriptor output result according to the result output by the shared feature encoder and the key point position information output by the key point detection encoder.

[0062] Such as Figure 3The following is a schematic diagram of the model architecture of the key point extraction model. The feature embedding layer and two FasterNet block modules constitute the shared feature encoder. One FasterNet block module and one convolutional layer constitute the key point detection decoder. The feature synthesis layer, two FasterNet block modules, and the convolutional layer constitute the descriptor generation decoder. In the key point extraction model, the initial feature embedding layer uses a 4×4 convolutional kernel and a convolution with a stride of 4 to quickly extract high-level image features, and the shared feature encoder reduces the spatial dimension of the input image to 1 / 8, reducing the subsequent computational pressure. Similarly, in the feature synthesis layer, convolutional kernels of other sizes can also be used for convolutional processing and deconvolutional processing to synthesize the feature extraction results output by the shared feature encoder and improve the accuracy of local region feature extraction. As for the convolutional kernel sizes of other involved convolutional layers, they can be adjusted according to the actual situation and are not restricted here.

[0063] In this embodiment, by using the structure of the FasterNet block in the key point extraction model, the model parameters are reduced, the model computational amount is reduced, and thus the model processing speed is improved, facilitating the real-time processing of the camera position.

[0064] In a feasible embodiment, when training the key point extraction model, the first loss function is used in the key point detection decoder to predict the key point position information;

[0065] The first loss function is determined according to the cross-entropy loss between the region vectors of multiple local regions in the feature map obtained by the key point extraction model and the image region labels corresponding to the local regions; the region vector is composed of the pixel point features of the local region and the key point extraction results of the local region.

[0066] Specifically, in the training stage of the key point extraction model, the pixel point features in the feature map finally output by the model are expanded into multiple region vectors, and each region vector represents a local region in the image. The multiple local regions can overlap. To improve the accuracy of the first loss function in predicting the key point position information, in addition to calculating the value of the first loss function according to the pixel point features of the local region, it also includes calculating the value of the first loss function in combination with the key point extraction results of the local region, so as to strengthen the key point feature extraction results of each local region in the calculation of the first loss function, and further improve the accuracy of predicting the key point position information through the first loss function.

[0067] Exemplarily, in the training stage of the key point extraction model, the pixel point features in the feature map finally output by the model are expanded into multiple region vectors of a size, and determine the key point extraction results of the local regions corresponding to the region vectors, and also add the key point extraction results to the region vectors to participate in the calculation of the first loss function. The key point extraction results of the local regions include that there are key points in the region and that there are no key points in the region, that is, multiple region vectors of a size. Exemplarily, by adding a "unkey" category in the region vector to represent the local region without key points.

[0068] For each region vector of a size, after performing normalization processing, perform cross-entropy loss calculation with the image region label corresponding to the region vector. The cross-entropy loss calculation result is the calculation result of the first loss function. Similarly, the dimension of the image region label corresponding to the region vector is also , determined according to the label of the sample image. In addition to including the feature information of the image region, the image region label corresponding to the region vector also includes whether there are key points in the image region. For example, the first loss function is determined according to the following formula: , where represents performing cross-entropy loss calculation on the vector and label carrying region key point information, p represents each local region in the image, is the region vector of the local region output by the model, is the image region label corresponding to the local region, and each label vector has a dimension of , that is, includes whether there is a key point label in this local region.

[0069] In a feasible embodiment, the key point extraction model is obtained by distillation training according to the teacher model; when performing distillation training on the key point extraction model and the teacher model, the second loss function is used in the descriptor generation decoder to align the descriptors output by the key point extraction model and the teacher model;

[0070] The second loss function is determined according to the following formula:

[0071] ;

[0072] where represents the value of the second loss function, represents the first descriptor obtained by the key point extraction model, represents the second descriptor obtained by the teacher model, is an orthogonal matrix, determined according to the singular value decomposition results of the first descriptor and the second descriptor.

[0073] In this embodiment, by introducing the orthogonal matrix R, the descriptor results output by the teacher model and the student model are aligned, and the high-dimensional descriptor output by the teacher model and the low-dimensional descriptor output by the student model are mapped through the orthogonal matrix R to ensure the accuracy of loss calculation based on the mapping result, thereby improving their similarity. The student model is the key point extraction model. The orthogonal matrix for its descriptor is calculated according to the following formula: Calculate U and V in through singular value decomposition, where and are both orthogonal matrices, is a diagonal matrix, and R is determined according to the calculated U and V, .

[0074] When optimizing the model through the above second loss function, the goal is to make the descriptor of the student model and the descriptor of the teacher model as close to alignment as possible after being mapped through . That is, minimize to make the descriptor of the student model as close as possible to the descriptor of the teacher model.

[0075] Optionally, in the descriptor generation decoder, the descriptor is compressed by the LocalPCA technique to generate a compact multi-dimensional feature descriptor for each detected key point. Traditional PCA compression usually performs statistical analysis on the entire sample data set. In the embodiment of the present invention, the descriptor generation decoder performs PCA compression of the descriptor separately on each sample image, avoiding information loss that may occur during compression on a large-scale data set, and at the same time ensuring the compactness of the descriptor of each image.

[0076] During the training process of the key point extraction model, the knowledge distillation method is adopted to improve the performance of the model. The total loss of the key point extraction model consists of and , and the formula is: . is used to measure the performance of the student model in the detection task, while is used to ensure the consistency of the output of the student model and the descriptor output by the teacher model in the high-dimensional space. For example, the teacher model selects images with at least 128 key points in the Google Landmark Dataset v2 (GLDv2) dataset for distillation training in SuperPoint detection.

[0077] S230. Determine the image matching result between the actual image and each first virtual image according to the matching result of the key point information of the actual image and each first virtual image.

[0078] S240. Determine the second virtual image from each first virtual image according to the image matching result, and determine the position determination result of the actual camera according to the second virtual image.

[0079] The technical solution of this embodiment realizes efficient key point detection and descriptor generation through the key point extraction model, can extract high-quality features from the actual image and the virtual image, and improves the accuracy of key point matching according to the feature extraction result. And the structure of the FasterNet block is used in the key point extraction model to reduce the model parameters, and LocalPCA is used to further compress the descriptor, and the first loss function and the second loss function are created to make the model training converge faster. At the same time, during training, learn from the teacher model through the method of knowledge distillation, reducing the model size without reducing the model accuracy.

[0080] Figure 4 It is a flowchart of another method for determining the camera position based on a 3D map provided by an embodiment of the present invention. This embodiment further refines the image matching process in the above embodiment. As Figure 4 shown, the method includes:

[0081] S410. Determine the actual image collected by the actual camera deployed in the actual scene, and determine the multiple first virtual images collected by the multiple virtual cameras corresponding to the actual camera.

[0082] S420. Extract features from the actual image and each first virtual image to obtain the key point information of each image.

[0083] S430. Based on the pre-trained graph neural network, determine the global relationship of the key points in the actual image and each first virtual image, and obtain the matching probability between each actual key point in the actual image and each virtual key point in each first virtual image.

[0084] Learn and capture the global relationship between all key points in the actual image and the first virtual image through the graph neural network, so as to realize high-precision image feature matching. The main architecture of the graph neural network mainly includes a key point encoder, a graph neural network message passing layer, and a matching layer. As Figure 5The figure shows the architecture diagram of the graph neural network. Among them, the key point encoder is used to construct the graph structure and perform feature encoding based on the key point detection results and corresponding descriptors of the actual image and the first virtual image. The graph neural network message passing layer is used to perform self-attention and cross-attention processing on the feature encoding results of the key point encoder to obtain the feature vectors of the actual image and the first virtual image. The matching layer is used to perform feature matching based on the feature vectors of the actual image and the first virtual image to obtain the matching probabilities between each actual key point in the actual image and each virtual key point in each first virtual image.

[0085] Specifically, the output result of the graph neural network is the matching probability between each actual key point in the actual image and each virtual key point in the first virtual image, which can be represented in the form of a matching probability matrix. Each value in the matrix represents the similarity between every two key points, and this similarity is represented by the inner product. Additionally, the matching probability matrix includes the results of non-matching of two key points, that is, there are special values in the matrix, such as represented by "unkey", to indicate that the corresponding two key points do not match, which is used to handle the partially occluded or missing matching key points in the image. By adding this category to the matching probability matrix, the accuracy of subsequent key point matching based on the matching probability matrix is improved, and the error influence of the occluded or missing matching key points is avoided.

[0086] S440. Process the matching probabilities between the target actual key point in the actual image and each virtual key point in any first virtual image, and determine the virtual key point that matches successfully with the target actual key point according to the processing results.

[0087] Since the orders of magnitude of the matching probabilities between the target actual key point in the actual image obtained by the graph neural network and each virtual key point in any first virtual image may be different, in order to ensure the accuracy of subsequent calculations, the processing operation of the matching probabilities at least includes a normalization operation. And in order to avoid the interference of abnormal matching probability values in the matching probabilities to key point matching, the processing operation also includes a non-maximum suppression operation. After processing the matching probabilities, determine the virtual key point with the highest matching probability with the target actual key point among each virtual key point according to the processing results. If this matching probability is greater than the preset probability threshold, it is determined that this virtual key point matches successfully with this target actual key point.

[0088] Exemplarily, a matching probability matrix is obtained based on the matching probabilities between the actual key points in the actual image and each virtual key point in any first virtual image. The matching probability matrix is normalized by the Sinkhorn algorithm to generate a soft assignment probability matrix, ensuring that the sum of the matching probabilities between any key point to be matched in the actual image and each virtual key point in the first virtual image in the soft assignment probability matrix is 1. Then, the soft assignment probability matrix is processed by the non-maximum suppression algorithm to remove duplicate or unreliable matching key point pairs. Finally, the key point pairs with a matching probability greater than the preset probability threshold in the key point pairs to be matched are regarded as successfully matched.

[0089] S450. Determine the image matching probability between the actual image and the first virtual image based on the sum of the matching probabilities between each actual key point in the actual image and the successfully matched virtual key points in the first virtual image, as the image matching result.

[0090] Sum the matching probabilities of all successfully matched key point pairs in the actual image and the first virtual image. The sum result is the image matching probability between the actual image and the first virtual image, as the image matching result. This image matching probability represents the similarity between the two images.

[0091] Exemplarily, the actual image includes 10 key points, the first virtual image includes 8 key points. After the above key point matching, 6 pairs of successfully matched key point pairs are obtained, that is, 6 key points in the actual image establish a one-to-one correspondence with 6 key points in the first virtual image. Determine the image matching probability between the actual image and the first virtual image according to the sum of the matching probability values of these 6 pairs of successfully matched key point pairs.

[0092] S460. Determine the second virtual image from each first virtual image according to the image matching result, and determine the positioning result of the actual camera according to the second virtual image.

[0093] The technical solution of this embodiment performs image feature matching based on the method of graph neural network, can capture the global relationship between key points, and achieve high-precision image matching. Effectively utilizes the attention mechanism and message passing mechanism, and improves the accuracy and robustness of image matching.

[0094] Figure 6 This is a flowchart of another method for determining the camera position based on a three-dimensional map provided by an embodiment of the present invention. This embodiment further refines the determination of the second virtual image in the above embodiment. As Figure 6 shown, the method includes:

[0095] S610. Determine the actual image captured by the actual camera deployed in the actual scenario, and determine the multiple first virtual images captured by the multiple virtual cameras corresponding to the actual camera.

[0096] S620. Extract features from the actual image and each first virtual image to obtain the key point information of each image.

[0097] S630. Determine the image matching result between the actual image and each first virtual image according to the matching result of the key point information of the actual image and each first virtual image.

[0098] S640. Sort each first virtual image according to the image matching result to obtain a sorting result.

[0099] Among them, the first virtual image with a high matching degree with the actual image in the sorting result is located in front of the first virtual image with a low matching degree with the actual image.

[0100] Among them, the image matching result represents the similarity between the first virtual image and the actual image. Sort each first virtual image from high to low according to the similarity to obtain the sorting result of the first virtual image. That is, according to this sorting result, the similarity between the first virtual image ranked in front and the actual image is greater than that of the first virtual image ranked behind.

[0101] S650. Traverse each first virtual image based on the sorting result, and convert the virtual image coordinates of each virtual key point in the first virtual image to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image to obtain each virtual world coordinate.

[0102] Traverse each first virtual image according to the sorting result, calculate the matching error of the key points between each first virtual image and the actual image in turn, and select the second virtual image closest to the actual image from the first virtual images according to the matching error, avoiding the error caused by simply determining the second virtual image according to the image matching result, and improving the accuracy of camera position determination.

[0103] Since the image matching result is determined according to the image position, when determining the matching error of the key points between the first virtual image and the actual image, project the key points to the world coordinate system, and determine the matching error according to the matching situation in the world coordinate system to improve the accuracy of the matching error determination.

[0104] Specifically, the virtual image coordinates are determined according to the position information of each virtual key point in the first virtual image, and the virtual image coordinates are converted to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image, so as to obtain the virtual world coordinates of each key point. The corresponding relationship between the image coordinate system and the world coordinate system is determined according to the camera parameters of the virtual camera, and the image coordinates are converted to the world coordinates according to this corresponding relationship.

[0105] Optionally, on the basis of the above embodiments, the key point pairs where the actual image and the first virtual image are successfully matched are determined, the virtual key points in the key point pairs are determined as the reference virtual key points, and the actual key points are determined as the reference actual key points. The virtual image coordinates of each reference virtual key point in the first virtual image are converted to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image, so as to obtain the virtual world coordinates. The redundant calculation of the virtual key points and actual key points that are not successfully matched is reduced, the amount of calculation is reduced, and thus the point position determination efficiency is improved.

[0106] In a feasible embodiment, converting the virtual image coordinates of each virtual key point in the first virtual image to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image to obtain the virtual world coordinates includes:

[0107] Converting the virtual image coordinates of each virtual key point in the first virtual image to the camera coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image to obtain the virtual camera coordinates;

[0108] Determining the external camera parameters of the virtual camera according to the camera parameters of the virtual camera corresponding to the first virtual image and the virtual camera coordinates;

[0109] Converting the virtual camera coordinates to the world coordinate system according to the external camera parameters of the virtual camera to obtain the virtual world coordinates.

[0110] In the process of converting the image coordinates to the world coordinates, the image coordinates are first converted to the camera coordinates, that is, the relationship between the image coordinate system and the camera coordinate system is first determined, and then the camera coordinates are converted to the world coordinates, that is, the relationship between the camera coordinate system and the world coordinate system is determined.

[0111] Specifically, the relationship between the image coordinate system and the camera coordinate system is determined according to the camera parameters of the virtual camera, and the coordinates of each virtual key point in the virtual camera coordinate system are calculated. The virtual camera coordinates are determined according to the following formula:

[0112] ;

[0113] Wherein, is the focal length of the virtual camera, is the coordinate of the center point in the virtual image, are the coordinates of each virtual key point in the virtual image, is the depth value of each virtual key point, and finally the camera coordinates of the virtual key points in the camera coordinate system are obtained .

[0114] Determine the extrinsic parameters of the virtual camera according to the virtual camera coordinates and camera parameters of each key point. The extrinsic parameters of the camera include a rotation matrix and a translation vector. Specifically, determine the rotation angle parameters according to the camera parameters of the virtual camera: the rotation angle α around the X-axis, the rotation angle β around the Y-axis, and the rotation angle γ around the Z-axis. Calculate the rotation matrix according to the rotation angle parameters. The calculation formula is:

[0115] .

[0116] Determine the position parameters according to the camera parameters of the virtual camera , and calculate the translation vector according to the position parameters and the rotation matrix , and the calculation formula is: . Finally, the extrinsic parameter matrix of the virtual camera can be obtained as .

[0117] According to the extrinsic parameter matrix of the virtual camera, the key point coordinates in the camera coordinate system can be converted to the world coordinate system. The formula is: ; where is the virtual camera coordinate of the i-th virtual key point, is the virtual world coordinate of the i-th virtual key point.

[0118] S660. Select a preset number of first key point pairs from the key point pairs successfully matched in the actual image and the first virtual image, and determine the actual camera pose parameters according to the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the preset number of first key points.

[0119] Select some key point pairs from the key point pairs successfully matched in the actual image and the first virtual image as the first key point pairs, and determine the pose parameters of the actual camera according to the virtual world coordinates of the virtual key points and the actual image coordinates of the actual key points in the first key point pairs. Since the actual camera pose parameters are determined according to the successfully matched first key point pairs, and since the first key point pairs are successfully matched, that is, the virtual key points and the actual key points are the same point, the virtual world coordinates of the virtual key points can also be considered as the world coordinates corresponding to the actual image coordinates. Therefore, the actual camera pose parameters can be determined according to the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the first key points.

[0120] Exemplarily, randomly select 6 pairs of successfully matched first key point pairs , where are the actual image coordinates of the actual key points, and are the virtual world coordinates of their corresponding virtual key points. The actual camera pose parameters are calculated using the EPnP algorithm , and their corresponding representation relationship is: . The number of the selected first key pairs is determined according to the unknown parameter values in the actual camera pose parameters, and there is no limitation here.

[0121] S670. Determine the projection errors of other successfully matched second key point pairs according to the actual camera pose parameters, and determine whether the first virtual image meets the error condition according to the projection errors.

[0122] Among them, the other successfully matched second key point pairs are the remaining key point pairs except the first key point pairs among all the key point pairs successfully matched in the actual image and the first virtual image. Since the actual camera pose parameters are determined under the condition that the first virtual image is considered to match the actual image, and the matching of the first virtual image and the actual image is determined according to the matched key point pairs, if the precondition for successful matching is accurate, the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the second key points also meet the corresponding relationship of the actual camera pose parameters. If the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the second key points do not meet the corresponding relationship of the actual camera pose parameters, that is, the projection error is large, it is considered that the precondition for the successful matching of the actual image and the first virtual image is inaccurate, that is, the actual camera pose parameters are inaccurate.

[0123] Specifically, convert the virtual world coordinates of the virtual key points in the second key point pairs according to the actual camera pose parameters to obtain the reference image coordinates, determine the error between the reference image coordinates and the actual image coordinates of the actual key points in the second key point pairs. This error can be determined by the absolute value of the difference between the reference image coordinates and the actual image coordinates. Take this error as the projection error of the second key point pair, and determine the projection errors of all second key point pairs according to this method. If the number of second key point pairs with projection errors less than the error threshold is greater than the number threshold, it is determined that the first virtual image and the actual image match successfully, and the actual camera pose parameters are the pose parameters of the actual camera.

[0124] Exemplarily, according to the calculated actual camera pose parameters, calculate the matching key points in each second key point pair of the projection error , and its formula can be expressed as: .

[0125] S680. If not satisfied, continue to traverse the first virtual image based on the sorting result until the error condition is met or the preset traversal number is reached.

[0126] Among them, the error condition is whether the number of key point pairs with projection error less than the error threshold in the second key point pair is greater than the quantity threshold. If it is less than or equal to the quantity threshold, it is determined that the first virtual image does not meet the error condition, that is, the first virtual image does not match the actual image. Then, continue to traverse according to the sorting result of the first virtual image to determine the next first virtual image until the first virtual image traversed meets the error condition; or reach the preset traversal quantity, that is, the aforementioned preset traversal quantity of first virtual images do not meet the error condition according to the sorting result. Since the image matching degree between the subsequent first virtual images and the actual image is relatively low, no further calculation is performed, and the average value is directly obtained from the actual camera pose parameters calculated from the aforementioned preset traversal quantity of first virtual images as the final camera position determination result.

[0127] S690. If it is satisfied, determine that the first virtual image is the second virtual image, and determine the position determination result of the actual camera according to the actual camera pose parameters corresponding to the second virtual image.

[0128] If the number of key point pairs with projection error less than the error threshold in the second key point pair in the first virtual image is greater than the quantity threshold, it is determined that the first virtual image meets the error condition, that is, the first virtual image matches the actual image, and the first virtual image is determined to be the second virtual image.

[0129] Determine the position determination result of the actual camera according to the actual camera pose parameters corresponding to the second virtual image. The position determination result includes the position parameters and rotation angle parameters of the actual camera. And store the ID, camera IP, location range, position parameters, and rotation angle parameters of the actual camera in the database for use by the front-end for binding and display.

[0130] Exemplarily, convert the pose parameters of the actual camera into UE parameters: position parameters and rotation angle parameters. The pose parameters of the actual camera are Converted into position parameters , and the rotation angle parameter The calculation formula is: ; ; where represents the value in the first row and first column of the rotation matrix, represents the value in the second row and first column of the rotation matrix, represents the value in the third row and first column of the rotation matrix, represents the value in the third row and second column of the rotation matrix, represents the value in the third row and third column of the rotation matrix.

[0131] Remove the successfully matched second virtual image from the set of virtual images to be matched, select the next actual image to be matched, and continue to determine the position until all actual images are successfully matched.

[0132] The technical solution of this embodiment realizes the accurate determination of the actual camera pose through the projection conversion in the image coordinate system and the world coordinate, and fine-tunes the camera position parameters and rotation angle parameters through the final determination of the camera position by the actual camera pose, avoiding the error problem caused by simply determining the camera position according to the position parameters and rotation angle parameters of the virtual camera, and realizing high-precision position determination.

[0133] Figure 7 It is a schematic structural diagram of a camera position determination device provided by an embodiment of the present invention. As Figure 7 shown, the device includes:

[0134] An image determination module 710, configured to determine the actual images collected by the actual cameras deployed in the actual scene, and determine a plurality of first virtual images collected by a plurality of virtual cameras corresponding to the actual cameras; wherein, the virtual cameras are created in a three-dimensional map corresponding to the actual scene according to the camera parameters of the actual cameras;

[0135] A key point extraction module 720, configured to perform feature extraction on the actual images and each of the first virtual images to obtain the key point information of each image;

[0136] An image matching module 730, configured to determine the image matching result between the actual image and each of the first virtual images according to the matching result of the key point information of the actual image and each of the first virtual images;

[0137] A position determination module 740, configured to determine a second virtual image from each of the first virtual images according to the image matching result, and determine the position determination result of the actual camera according to the second virtual image.

[0138] The technical solution of this embodiment creates corresponding virtual cameras according to the actual cameras, and obtains the camera position determination result according to the detection and matching results of the key points in the images collected by each camera. According to this camera position determination result, the actual cameras can be marked in the three-dimensional map, significantly improving the accuracy and consistency of position determination, thereby reducing the possibility of manual marking errors; and through the technical solution of this embodiment, the position marking speed can be improved, real-time marking can be realized, so as to adapt to real-time and large-scale data scenarios.

[0139] Optionally, the key point extraction module includes:

[0140] Feature extraction is performed on the actual image and each of the first virtual images based on a pre-trained key point extraction model to obtain the key point detection results and corresponding descriptors of each image as the key point information; wherein, the descriptor is used to characterize the overall features of the local area where the corresponding key point is located.

[0141] Optionally, the key point extraction model includes a shared feature encoder, a key point detection decoder, and a descriptor generation decoder; the output of the shared feature encoder is the input of the key point detection decoder and the descriptor generation decoder, the output of the key point detection decoder is the key point detection result, and the output of the descriptor generation decoder is the descriptor corresponding to each key point;

[0142] Among them, the shared feature encoder includes at least one convolutional layer and at least one FasterNet block module; the key point detection decoder includes at least one convolutional layer and at least one FasterNet block module; the descriptor generation decoder includes a feature synthesis layer, at least one convolutional layer, and at least one FasterNet block module.

[0143] Optionally, when training the key point extraction model, the first loss function is used in the key point detection decoder to predict the key point position information;

[0144] The first loss function is determined according to the cross-entropy loss between the region vectors of multiple local regions in the feature map obtained by the key point extraction model and the image region labels corresponding to the local regions; the region vector is composed of the pixel point features of the local region and the key point extraction result of the local region.

[0145] Optionally, the key point extraction model is obtained by distillation training based on a teacher model; when performing distillation training on the key point extraction model and the teacher model, the second loss function is used in the descriptor generation decoder to align the descriptors output by the key point extraction model and the teacher model;

[0146] The second loss function is determined according to the following formula:

[0147] ;

[0148] Wherein, represents the value of the second loss function, represents the first descriptor obtained by the key point extraction model, represents the second descriptor obtained by the teacher model, is an orthogonal matrix, which is determined according to the singular value decomposition results of the first descriptor and the second descriptor.

[0149] Optionally, the image matching module is specifically configured to:

[0150] Determine the global relationship of key points in the actual image and each of the first virtual images based on a pre-trained graph neural network, and obtain the matching probabilities between each actual key point in the actual image and each virtual key point in each of the first virtual images;

[0151] Process the matching probabilities between the target actual key point in the actual image and each virtual key point in any one of the first virtual images, and determine the virtual key point that successfully matches the target actual key point according to the processing result;

[0152] Determine the image matching probability between the actual image and the first virtual image according to the sum of the matching probabilities between each actual key point in the actual image and the virtual key point that successfully matches in the first virtual image, and use it as the image matching result.

[0153] Optionally, the point position determination module includes:

[0154] The first virtual image sorting unit is used to sort each of the first virtual images according to the image matching result to obtain a sorting result; wherein, the first virtual image with a higher matching degree with the actual image in the sorting result is located in front of the first virtual image with a lower matching degree with the actual image;

[0155] The virtual world coordinate determination unit is used to traverse each first virtual image based on the sorting result, and convert the virtual image coordinates of each virtual key point in the first virtual image to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image, so as to obtain each virtual world coordinate;

[0156] The actual camera pose parameter determination unit is used to select a preset number of first key point pairs from the successfully matched key point pairs in the actual image and the first virtual image according to the image matching result, and determine the actual camera pose parameters according to the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the preset number of first key points;

[0157] The projection error determination unit is used to determine the projection error of other successfully matched second key point pairs according to the actual camera pose parameters, and determine whether the first virtual image meets the error condition according to the projection error;

[0158] The first judgment unit is used to, if not satisfied, continue to traverse the first virtual image based on the sorting result until the error condition is met or the preset traversal number is reached;

[0159] A second determination unit, configured to determine that the first virtual image is a second virtual image if the condition is met, and determine a point position determination result of the actual camera according to the actual camera pose parameters corresponding to the second virtual image.

[0160] Optionally, the virtual world coordinate determination unit is specifically configured to:

[0161] Convert the virtual image coordinates of each virtual key point in the first virtual image to the camera coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image, to obtain respective virtual camera coordinates;

[0162] Determine the extrinsic parameters of the virtual camera according to the camera parameters of the virtual camera corresponding to the first virtual image and the respective virtual camera coordinates;

[0163] Convert the respective virtual camera coordinates to the world coordinate system according to the extrinsic parameters of the virtual camera, to obtain respective virtual world coordinates.

[0164] The camera point position determination device based on a three-dimensional map provided by an embodiment of the present invention can execute the camera point position determination method based on a three-dimensional map provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0165] In the technical solution of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations, and do not violate public order and good customs.

[0166] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0167] Figure 8 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0168] Such as Figure 8As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0169] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0170] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for determining the camera position based on a three-dimensional map.

[0171] In some embodiments, the method for determining the camera position based on a three-dimensional map can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for determining the camera position based on a three-dimensional map described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for determining the camera position based on a three-dimensional map by any other appropriate means (e.g., by means of firmware).

[0172] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or a general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0173] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0174] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0176] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including switching components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, switching components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0177] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0178] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication unit 19, or installed from a storage unit 18, or installed from a ROM 12. When the computer program is executed by a processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0179] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0180] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for determining camera positions based on a three-dimensional map, characterized in that The method includes: determining an actual image captured by an actual camera deployed in an actual scenario, and determining a plurality of first virtual images captured by a plurality of virtual cameras corresponding to the actual camera; wherein, the virtual cameras are created in a three-dimensional map corresponding to the actual scenario according to the camera parameters of the actual camera; performing feature extraction on the actual image and each of the first virtual images to obtain key point information of each image; determining an image matching result between the actual image and each of the first virtual images according to a matching result of the key point information of the actual image and each of the first virtual images; determining a second virtual image from each of the first virtual images according to the image matching result, and determining a point position determination result of the actual camera according to the second virtual image.

2. The method according to claim 1, characterized in that, Performing feature extraction on the actual image and each of the first virtual images to obtain key point information of each image, including: performing feature extraction on the actual image and each of the first virtual images based on a pre-trained key point extraction model to obtain key point detection results of each image and corresponding descriptors as the key point information; wherein, the descriptors are used to characterize the overall features of the local regions where the corresponding key points are located.

3. The method according to claim 2, wherein Wherein, the key point extraction model includes a shared feature encoder, a key point detection decoder, and a descriptor generation decoder; an output of the shared feature encoder is an input to the key point detection decoder and the descriptor generation decoder, an output of the key point detection decoder is the key point detection result, and an output of the descriptor generation decoder is a descriptor corresponding to each key point; wherein, the shared feature encoder includes at least one convolutional layer and at least one FasterNet block module; the key point detection decoder includes at least one convolutional layer and at least one FasterNet block module; the descriptor generation decoder includes a feature synthesis layer, at least one convolutional layer, and at least one FasterNet block module.

4. The method according to claim 3, wherein Wherein, when training the key point extraction model, a first loss function is used in the key point detection decoder to predict key point position information; the first loss function is determined according to a cross-entropy loss between region vectors of a plurality of local regions in a feature map obtained by the key point extraction model and image region labels corresponding to the local regions; the region vectors are composed of pixel point features of the local regions and key point extraction results of the local regions.

5. The method according to claim 3, wherein Wherein, the key point extraction model is obtained by distillation training according to a teacher model; when performing distillation training on the key point extraction model and the teacher model, a second loss function is used in the descriptor generation decoder to align descriptors output by the key point extraction model and the teacher model; the second loss function is determined according to the following formula: ; Among them, represents the value of the second loss function, represents the first descriptor obtained by the key point extraction model, represents the second descriptor obtained by the teacher model, is an orthogonal matrix, determined according to the singular value decomposition results of the first descriptor and the second descriptor.

6. The method according to claim 1, wherein Determining an image matching result between the actual image and each of the first virtual images according to a matching result of the key point information of the actual image and each of the first virtual images, including: Based on a pre-trained graph neural network, determine the global relationship of key points in the actual image and each of the first virtual images, and obtain the matching probabilities between each actual key point in the actual image and each virtual key point in each of the first virtual images; Process the matching probabilities between the target actual key point in the actual image and each virtual key point in any one of the first virtual images, and determine the virtual key point that matches successfully with the target actual key point according to the processing results; Determine the image matching probability between the actual image and the first virtual image according to the sum of the matching probabilities between each actual key point in the actual image and the virtual key points that match successfully in the first virtual image, and use it as the image matching result.

7. The method according to claim 1, characterized in that Determine the second virtual image from each of the first virtual images according to the image matching result, and determine the positioning result of the actual camera according to the second virtual image, including: Sort each of the first virtual images according to the image matching result to obtain a sorting result; wherein, the first virtual image with a higher matching degree with the actual image in the sorting result is located in front of the first virtual image with a lower matching degree with the actual image; Traverse each first virtual image based on the sorting result, and convert the virtual image coordinates of each virtual key point in the first virtual image to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image to obtain each virtual world coordinate; Select a preset number of first key point pairs from the key point pairs that match successfully in the actual image and the first virtual image according to the image matching result, and determine the actual camera pose parameters according to the actual image coordinates of the actual key points and the virtual world coordinates of the virtual key points in the preset number of first key points; Determine the projection error of other successfully matched second key point pairs according to the actual camera pose parameters, and determine whether the first virtual image meets the error condition according to the projection error; If not, continue to traverse the first virtual images based on the sorting result until the error condition is met or the preset traversal number is reached; If so, determine the first virtual image as the second virtual image, and determine the positioning result of the actual camera according to the actual camera pose parameters corresponding to the second virtual image.

8. The method according to claim 7, wherein Convert the virtual image coordinates of each virtual key point in the first virtual image to the world coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image, to obtain each virtual world coordinate, including: Convert the virtual image coordinates of each virtual key point in the first virtual image to the camera coordinate system according to the camera parameters of the virtual camera corresponding to the first virtual image to obtain each virtual camera coordinate; Determine the external camera parameters of the virtual camera according to the camera parameters of the virtual camera corresponding to the first virtual image and each virtual camera coordinate; Convert each virtual camera coordinate to the world coordinate system according to the external camera parameters of the virtual camera to obtain each virtual world coordinate.

9. A camera position determination device based on a three-dimensional map, characterized in that, The device includes: An image determination module, configured to determine an actual image captured by an actual camera deployed in an actual scene, and determine a plurality of first virtual images captured by a plurality of virtual cameras corresponding to the actual camera; wherein, the virtual cameras are created in a three-dimensional map corresponding to the actual scene according to the camera parameters of the actual camera; A key point extraction module, configured to perform feature extraction on the actual image and each of the first virtual images to obtain key point information of each image; An image matching module, configured to determine an image matching result between the actual image and each of the first virtual images according to a matching result of the key point information of the actual image and each of the first virtual images; A point position determination module, configured to determine a second virtual image from each of the first virtual images according to the image matching result, and determine a point position determination result of the actual camera according to the second virtual image.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the method for determining a camera point position based on a three-dimensional map according to any one of claims 1-8.

Citation Information

Patent Citations

  • Method and device for acquiring three-dimensional reconstruction training data and electronic equipment

    CN113205591A

  • Camera point position calibration method, device, equipment and medium

    CN113763477A

  • Camera pose determination method and device, electronic equipment and storage medium

    CN115205383A

  • Bullet time video generation method and device, electronic equipment and storage medium

    CN116193158A

  • Calibration parameter generation method of virtual camera and electronic equipment

    CN119919501A

Cited By

  • Flight control method, device and equipment of unmanned aerial vehicle and storage medium

    CN120973053A