Video GIS positioning method and device

By splitting the video into multiple frames and identifying reference objects, combined with camera coordinate correction, high-precision spatial positioning of video data was achieved, solving the problem of insufficient position binding in traditional video systems and improving the efficiency of video retrieval and utilization.

CN121120779APending Publication Date: 2025-12-12BEIJING OUTLOOK CHINA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511296990.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Traditional video systems lack close integration with geographic location, resulting in low efficiency in video retrieval and utilization, and making it difficult for users to quickly locate the location and scope of events through spatial dimensions.

Method used

By splitting the video to be processed into multiple frames, identifying continuous positioning reference objects and auxiliary positioning reference objects, and using camera coordinates and spatial coordinate correction information, accurate positioning of each frame of positioning image is achieved.

Benefits of technology

It achieves high-precision spatial mapping of video data, improves the utilization rate and retrieval efficiency of video data, and enables users to quickly locate relevant images based on geographical location, avoiding the time-consuming and error-prone nature of traditional manual retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120779A_ABST
    Figure CN121120779A_ABST
Patent Text Reader

Abstract

The invention provides a video GIS positioning method and device, and belongs to the technical field of video processing. The method comprises the following steps: firstly, acquiring a to-be-processed video, splitting the to-be-processed video into multiple frames of positioning images, and extracting a continuous positioning reference object and an auxiliary positioning reference object from each frame of image; and respectively determining coordinate values of the two types of reference objects under a camera coordinate system, and correcting the coordinate values of the continuous reference objects by using the auxiliary reference object so as to obtain attitude correction information. And then, determining a spatial coordinate value by taking the continuous reference object as a first index target, determining a spatial coordinate value by taking the auxiliary reference object as a second index target, and correcting the first spatial coordinate value through the auxiliary reference object to obtain a corrected spatial coordinate value of each frame of image. And finally, fusing the corrected space coordinate value with the attitude correction information, and further determining a target space coordinate value of each frame of positioning image. The invention aims to improve the problem of low utilization rate of video data without clear position information in the existing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video GIS positioning method and apparatus. Background Technology

[0002] With the continuous advancement of informatization and intelligentization, video surveillance technology and Geographic Information Systems (GIS) are being used more and more widely in various industries. Video can provide intuitive dynamic image information, realistically reflecting the on-site operating status and event process, and is widely used in urban security, traffic management, emergency command, and industrial inspection.

[0003] However, traditional video systems lack close integration with geographic location. Video data is usually stored and retrieved only in chronological order, making it difficult for users to quickly locate the location and scope of events using spatial dimensions, resulting in low efficiency in video retrieval and utilization. Summary of the Invention

[0004] This application provides a video GIS positioning method and apparatus to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: Firstly, a video GIS positioning method is provided, the method including: The video to be processed is acquired and split into multiple positioning images. Each positioning image is filtered to determine the continuous positioning reference object and the auxiliary positioning reference object in each positioning image. Determine the first camera coordinate value of the continuous positioning reference object in the camera coordinate system and the second camera coordinate value of the auxiliary positioning reference object in the camera coordinate system, and correct the first camera coordinate value based on the second camera coordinate value to determine the pose correction information of each frame of positioning image; Using a continuous positioning reference object as the first index target, the first spatial coordinate value of each frame of positioning image is determined. Using an auxiliary positioning reference object as the second index target, the second spatial coordinate value of each frame of positioning image is determined. The first spatial coordinate value is then corrected based on the second spatial coordinate value to determine the corrected spatial coordinate value of each frame of positioning image. The target spatial coordinates of each frame of the positioning image are determined based on the fusion result of the corrected spatial coordinates and attitude correction information.

[0006] The video GIS positioning method provided in this application achieves precise positioning based on image features by splitting the video to be processed into multiple frames and identifying continuous and auxiliary positioning reference objects in each frame. By determining the coordinates of the reference objects in camera coordinates and using auxiliary reference objects for attitude correction, errors caused by changes in the shooting device's attitude, shaking, or environmental interference are effectively eliminated. Furthermore, by correcting the spatial coordinates of continuous and auxiliary reference objects, precise mapping of each frame of the positioning image to the actual geographic space is achieved. Combining attitude correction information with corrected spatial coordinates, this method can generate high-precision target spatial coordinate values, ensuring that each frame of the video accurately corresponds to its real spatial location. For video data without explicit location information, it can fully utilize the image's own feature information for spatial estimation, thereby significantly improving the utilization rate and value of this type of video data.

[0007] In one possible implementation, the video to be processed is acquired and split into multiple frames of localization images, including: The positioning accuracy of video GIS positioning is obtained, and the image segmentation interval is determined based on the positioning accuracy. The image segmentation interval is negatively correlated with the positioning accuracy. The video to be processed is divided into multiple positioning images according to the image splitting interval.

[0008] In one possible implementation, each frame of the positioning image is filtered to determine continuous positioning reference objects and auxiliary positioning reference objects in each frame of the positioning image, including: Identify candidate reference objects in each frame of the localization image; Based on the evaluation results of the stability and continuity of the candidate reference objects in adjacent frame images, continuously localized reference objects are determined; Based on the evaluation results of the identifiability and positioning accuracy of the candidate reference object in each frame of the image, the auxiliary positioning reference object is determined.

[0009] In one possible implementation, determining continuously located reference objects based on the stability and continuity evaluation results of candidate reference objects in adjacent frame images includes: Determine the stability evaluation score of each candidate reference object in adjacent frame images. The more times a candidate reference object appears in multiple adjacent frame images, the higher the stability evaluation score. The continuity evaluation score of each candidate reference object in adjacent frame images is determined. The smaller the inter-frame positional offset and the smoother the trajectory of the candidate reference object in multiple adjacent frame images, the higher the continuity evaluation score. Based on the weighted fusion result of the stability evaluation score and the continuity evaluation score, a reference object for continuous positioning is determined.

[0010] In one possible implementation, the first camera coordinate values ​​are corrected based on the second camera coordinate values ​​to determine the pose correction information for each frame of the positioning image, including: Obtain the spatial relationship between continuous positioning reference objects and auxiliary positioning reference objects in the previous frame positioning image; A corrected coordinate system is constructed using the coordinate values ​​of the second camera, and the coordinate values ​​of the continuous positioning reference object in the corrected coordinate system are determined based on the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object. The first camera coordinates of the continuously positioned reference object are mapped to the corrected coordinate system, and the pose correction information of each frame of the positioning image is determined based on the difference between the mapping result and the coordinates of the continuously positioned reference object in the corrected coordinate system.

[0011] In one possible implementation, using a continuous positioning reference object as a first index target, determining the first spatial coordinate value of each frame of the positioning image, and using an auxiliary positioning reference object as a second index target, determining the second spatial coordinate value of each frame of the positioning image, includes: Using the feature information of the continuously located reference object as the first index target, find the corresponding first target spatial region, and use the position information of the first target spatial region as the first spatial coordinate value; Using the feature information of the auxiliary positioning reference object as the second index target, find the corresponding second target spatial region, and use the position information of the second target spatial region as the second spatial coordinate value; the feature information includes at least one of contour shape, texture features, or semantic label.

[0012] In one possible implementation, the target spatial coordinates of each frame of the positioning image are determined based on the fusion result of the corrected spatial coordinates and pose correction information, including: Based on the pose correction information of each frame of the positioning image, determine the camera-space correction coefficients for each frame of the positioning image. The corrected spatial coordinates of each frame of the positioning image are adjusted based on the camera-space correction coefficient to determine the target spatial coordinates of each frame of the positioning image.

[0013] In one possible implementation, the method further includes: Based on the target spatial coordinates of each frame of the localization image, generate localization marker results; Add the corresponding positioning marker result to each frame of the positioning image.

[0014] Secondly, a video GIS positioning device is provided, the device comprising: The acquisition module is used to acquire the video to be processed, split the video into multiple frames of positioning images, filter each frame of positioning images, and determine the continuous positioning reference objects and auxiliary positioning reference objects in each frame of positioning images. The first correction module is used to determine the first camera coordinate value of the continuous positioning reference object in the camera coordinate system and the second camera coordinate value of the auxiliary positioning reference object in the camera coordinate system, and to correct the first camera coordinate value based on the second camera coordinate value to determine the pose correction information of each frame of positioning image. The second correction module is used to determine the first spatial coordinate value of each frame of positioning image by taking the continuous positioning reference object as the first index target, and to determine the second spatial coordinate value of each frame of positioning image by taking the auxiliary positioning reference object as the second index target, and to correct the first spatial coordinate value according to the second spatial coordinate value, and to determine the corrected spatial coordinate value of each frame of positioning image. The spatial coordinate determination module is used to determine the target spatial coordinate value of each frame of the positioning image based on the fusion result of the corrected spatial coordinate value and the attitude correction information.

[0015] In one possible implementation, the acquisition module includes: The first acquisition submodule is used to acquire the positioning accuracy of video GIS positioning and determine the image segmentation interval based on the positioning accuracy. The image segmentation interval is negatively correlated with the positioning accuracy. The splitting submodule is used to split the video to be processed into multiple frames of positioning images according to the image splitting interval.

[0016] In one possible implementation, the acquisition module further includes: The determination submodule is used to determine candidate reference objects in each frame of the localization image; The first filtering submodule is used to determine continuously located reference objects based on the evaluation results of the stability and continuity of candidate reference objects in adjacent frame images. The second filtering submodule is used to determine the auxiliary positioning reference object based on the evaluation results of the identifiability and positioning accuracy of the candidate reference object in each frame of the image.

[0017] In one possible implementation, the first filtering submodule includes: The first evaluation unit is used to determine the stability evaluation score of each candidate reference object in adjacent frame images. The more times the candidate reference object appears in multiple adjacent frame images, the higher the stability evaluation score. The second evaluation unit is used to determine the continuity evaluation score of each candidate reference object in adjacent frame images. The smaller the inter-frame position offset and the smoother the trajectory of the candidate reference object in multiple adjacent frame images, the higher the continuity evaluation score. The third evaluation unit is used to determine the continuous positioning reference object based on the weighted fusion result of the stability evaluation score and the continuity evaluation score.

[0018] In one possible implementation, the first correction module includes: The spatial relationship acquisition submodule is used to acquire the spatial relationship between continuous positioning reference objects and auxiliary positioning reference objects in the previous frame positioning image; The reference system determination submodule is used to construct a corrected coordinate system using the coordinate values ​​of the second camera, and to determine the coordinate values ​​of the continuous positioning reference object in the corrected coordinate system based on the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object. The correction submodule is used to map the first camera coordinate values ​​of the continuously positioned reference object to the corrected coordinate system, and determine the pose correction information of each frame of the positioning image based on the difference between the mapping result and the coordinate values ​​of the continuously positioned reference object in the corrected coordinate system.

[0019] In one possible implementation, the second correction module includes: The first search submodule is used to search for the corresponding first target spatial region using the feature information of the continuously located reference object as the first index target, and to use the location information of the first target spatial region as the first spatial coordinate value. The second search submodule is used to find the corresponding second target spatial region by using the feature information of the auxiliary positioning reference object as the second index target, and to use the position information of the second target spatial region as the second spatial coordinate value; the feature information includes at least one of contour shape, texture features, or semantic label.

[0020] In one possible implementation, the spatial coordinate determination module includes: The correction coefficient determination submodule is used to determine the camera-space correction coefficients for each frame of the positioning image based on the pose correction information of each frame of the positioning image. The adjustment submodule is used to adjust the corrected spatial coordinate values ​​of each frame of the positioning image according to the camera-space correction coefficient, and to determine the target spatial coordinate values ​​of each frame of the positioning image.

[0021] Thirdly, an electronic device is provided, including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement the method as described in any of the above aspects.

[0022] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method of any of the above aspects.

[0023] The technical effects of the second to fourth aspects and any of their embodiments are referenced and will not be repeated here. Attached Figure Description

[0024] Figure 1 This is a flowchart of the steps of a video GIS positioning method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the functional modules of a video GIS positioning device provided in an embodiment of the present invention. Detailed Implementation

[0025] To make the purpose, technical solution, and advantages of this application clearer, the following is combined with Figures 1-2 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0026] The terms used in the embodiments of this application, such as "second," are only used to distinguish features of the same type and should not be construed as indicating relative importance, quantity, order, etc.

[0027] The terms "exemplary" or "for example" used in the embodiments of this application are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0028] The terms "coupling" and "connection" used in the embodiments of this application should be interpreted broadly. For example, they can refer to a physical direct connection or an indirect connection achieved through electronic devices, such as a connection achieved through resistors, inductors, capacitors or other electronic devices.

[0029] As the scale of video surveillance continues to expand, existing video surveillance systems are gradually revealing significant limitations in data management and utilization. Traditional video systems typically store and index image data in chronological order, with video files organized solely based on the recording time. In this model, when users need to trace back an event or find a scene, they often have to search segment by segment along a timeline or rely on manual observation of details in the footage to infer the geographical location of the event. This retrieval method is not only cumbersome and time-consuming, but also prone to omissions or misjudgments due to the subjectivity and uncertainty of manual inference. For example, in large-scale monitoring environments, the number of monitoring points may reach hundreds or thousands. If the video data does not carry clear spatial location information, users cannot directly limit the target video by geographical scope during retrieval and must repeatedly search through a large number of irrelevant videos, resulting in a significant decrease in retrieval efficiency. Furthermore, when there are similar scenes, repeated features, or complex lighting in the monitoring footage, it becomes even more difficult to manually identify the video content, further increasing the difficulty of locating the event's location.

[0030] To address the aforementioned issues, this invention proposes the following concept: by establishing spatial identifiers for video data lacking explicit location information, the retrieval of video data is expanded from a single dimension relying solely on temporal order to a dual-dimensional approach encompassing both time and space. In this way, users can not only access videos based on time information but also quickly locate and retrieve relevant images based on geographic location, thus avoiding the problems of manual segment-by-segment searching and inaccurate positioning caused by the lack of location information in traditional systems.

[0031] Reference Figure 1 The present invention provides a video GIS positioning method, which may specifically include the following steps: S101: Acquire the video to be processed, and split the video into multiple positioning images. Filter each positioning image to determine the continuous positioning reference object and auxiliary positioning reference object in each positioning image.

[0032] In this embodiment, after acquiring the video to be processed, it is split into multiple positioning images. On one hand, video is essentially a dynamic sequence of multiple static images. By splitting it frame by frame, the continuous video signal can be converted into a sequence of images that is easy to process, thus using images as the basic processing unit in spatial coordinate calculation and pose estimation. On the other hand, frame-level processing can refine the image to each instant on the timeline, making the detection and positioning of the reference object more accurate and avoiding target blurring, positional shifts, or information loss caused by directly processing the entire video. The specific implementation steps include: S1011: Obtain the positioning accuracy of video GIS positioning, and determine the image segmentation interval based on the positioning accuracy. The image segmentation interval is negatively correlated with the positioning accuracy. S1012: Divide the video to be processed into multiple positioning images according to the image splitting interval.

[0033] In the implementations of S1011 and S1012, the required positioning accuracy for video GIS positioning is first obtained, and the image splitting interval is dynamically determined based on this positioning accuracy. Since the image splitting interval is negatively correlated with positioning accuracy, when the positioning accuracy requirement is high, the system needs to split the video at shorter intervals to obtain a denser sequence of positioning images on the time axis. This ensures sufficient temporal information support during reference object detection and spatial coordinate estimation, reducing error accumulation caused by inter-frame motion or pose changes. Conversely, when the positioning accuracy requirement is relatively low, the image splitting interval can be appropriately extended to reduce the number of frames, alleviating subsequent computation and storage overhead and achieving a balance between accuracy and efficiency. Subsequently, the system splits the video to be processed according to the determined image splitting interval, obtaining multiple frames of positioning images. This adaptive splitting mechanism effectively avoids redundant processing or insufficient accuracy problems caused by fixed splitting intervals while meeting the positioning accuracy requirements in different scenarios, providing a stable and reliable data foundation for subsequent reference object extraction, coordinate correction, and pose information fusion.

[0034] As an example, in road traffic monitoring scenarios, when the system needs to accurately locate vehicles traveling at high speeds, the positioning accuracy requirement is high. The system will automatically shorten the video segmentation interval, for example, generating a positioning image frame every 0.1 seconds, to ensure that it can capture subtle changes in the vehicle's position within a short period. In low-speed scenarios such as farmland inspections or infrastructure checks, the positioning accuracy requirement is relatively low. The system can then appropriately extend the segmentation interval, for example, generating a positioning image frame every 1 second, thereby reducing the generation and processing overhead of a large amount of redundant data while maintaining basic positioning accuracy. Through this differentiated and adaptive processing approach, the system can flexibly respond to the positioning needs of different application scenarios and achieve the optimal balance between accuracy and efficiency.

[0035] The specific steps for determining the continuous localization reference object and the auxiliary localization reference object in each frame of the localization image may include: S1013: Determine candidate reference objects in each frame of the localization image; S1014: Determine continuously located reference objects based on the evaluation results of the stability and continuity of candidate reference objects in adjacent frame images; S1015: Determine the auxiliary positioning reference object based on the evaluation results of the identifiability and positioning accuracy of the candidate reference object in each frame of the image.

[0036] In the implementations of S1013 to S1015, after acquiring each frame of the positioning image, candidate reference objects that meet the preset requirements are first screened from the image using feature detection and semantic recognition methods. These objects are usually static ground features with stable geometric features and clear textures, such as building corners, road markings, or fixed signs. During the screening process, objects with obvious motion or instability, such as pedestrians, vehicles, and leaves, are removed. At the same time, image clarity, occlusion rate, texture richness, and distribution in the image are combined to ensure that the candidate reference objects have good spatial dispersion, thereby avoiding geometric degradation during subsequent positioning. Based on this, the system matches and tracks candidate reference objects in adjacent frames. By analyzing the continuity and stability of their cross-frame trajectories, continuous positioning reference objects are selected. These are objects that can be reliably identified in multiple frames, have smooth position changes, and have high trajectory integrity. Such objects can provide a stable benchmark for overall attitude calculation. At the same time, the system also evaluates the recognition confidence and positioning accuracy of candidate reference objects in single-frame images, thereby selecting targets suitable as auxiliary positioning reference objects. Although these objects may not appear continuously throughout the entire sequence, they have high recognition and positioning accuracy in certain frames and can play a supplementary role when continuous positioning reference objects are insufficient or unevenly distributed. This forms a combination of continuous positioning reference objects and auxiliary positioning reference objects, ensuring that reference information that balances stability and accuracy can be obtained in each positioning image.

[0037] In one feasible implementation, determining continuously located reference objects based on the stability and continuity evaluation results of candidate reference objects in adjacent frame images may include: Determine the stability evaluation score of each candidate reference object in adjacent frame images. The more times a candidate reference object appears in multiple adjacent frame images, the higher the stability evaluation score. The continuity evaluation score of each candidate reference object in adjacent frame images is determined. The smaller the inter-frame positional offset and the smoother the trajectory of the candidate reference object in multiple adjacent frame images, the higher the continuity evaluation score.

[0038] In this embodiment, when evaluating the stability and continuity of candidate reference objects, the system comprehensively analyzes the performance of each candidate reference object in adjacent frame images. Specifically, firstly, it statistically analyzes whether the candidate reference object is continuously identified in multiple adjacent frames. If it maintains a high frequency of occurrence over a long time sequence, it indicates that the object is less affected by occlusion, lighting changes, or viewing angle changes during image acquisition, thus obtaining a higher stability evaluation score. Secondly, it calculates the spatial position changes of the candidate reference object between frames. If its positional shift is small and its motion trajectory is smooth and continuous in consecutive frames, it indicates that the object can provide stable geometric constraints, thus obtaining a higher score in the continuity evaluation. After calculating the above two evaluation scores, the system comprehensively ranks the stability and continuity of each candidate reference object through preset weighting rules or a comprehensive scoring mechanism, and finally selects the object with the best overall performance as the continuous positioning reference object.

[0039] As an example, if candidate reference objects in a scene include building corners, road signs, and trees, analysis of adjacent frames reveals that building corners are almost always present throughout the sequence with minimal positional changes, resulting in high stability and continuity scores. While road signs are clearly visible in some frames, their recognition may be interrupted by changes in lighting, leading to lower stability scores. Trees, due to the swaying of branches and leaves, exhibit significant positional changes between frames, resulting in noticeably lower continuity scores. Through comprehensive evaluation, building corners are ultimately selected as the continuous localization reference object, ensuring higher reliability and accuracy for subsequent pose correction and spatial coordinate calculations.

[0040] It should be noted that the selection process for auxiliary positioning reference objects can be based on a comprehensive evaluation of the identifiability and positioning accuracy of candidate reference objects in each frame of the image. This ensures that the selected auxiliary objects can compensate for the shortcomings of continuous positioning reference objects in certain specific situations. Specifically, the system first determines whether the candidate reference object can be stably detected and accurately identified in a single frame of the image. If the object maintains a high recognition success rate under different lighting conditions and angles, its identifiability score is high. Secondly, the system evaluates the positioning accuracy of the candidate reference object based on indicators such as edge sharpness, feature point richness, and fitting error of the positioning algorithm. If the object can obtain low-error coordinate calculation results in multiple frames, it indicates that the object has high positioning accuracy. After combining the above evaluation results, the system will weight and rank the identifiability and positioning accuracy according to preset weights, and select the best-performing object as the auxiliary positioning reference object.

[0041] S102: Determine the first camera coordinate value of the continuous positioning reference object in the camera coordinate system and the second camera coordinate value of the auxiliary positioning reference object in the camera coordinate system, and correct the first camera coordinate value based on the second camera coordinate value to determine the pose correction information of each frame of positioning image.

[0042] In this embodiment, after determining the continuous positioning reference object and the auxiliary positioning reference object, the system first extracts the spatial position of the continuous positioning reference object in each frame of the positioning image based on image processing and target recognition algorithms, thereby obtaining its first camera coordinate value in the camera coordinate system. Simultaneously, the system also detects and locates the auxiliary positioning reference object, obtaining its second camera coordinate value in the camera coordinate system. Since the continuous positioning reference object is usually located in the main area of ​​interest of the video sequence, it is susceptible to camera shake, pose changes, or environmental interference. Therefore, directly using its coordinate values ​​may lead to accumulated errors in pose estimation. Therefore, this embodiment introduces the auxiliary positioning reference object as a reference standard. By comparing the spatial distribution relationship between the two types of reference objects, the first camera coordinate value is differentially corrected or optimized for registration, thereby generating more accurate pose correction information. In this way, the system can effectively offset the instability caused by a single reference object, improve the robustness and reliability of the overall pose estimation, and provide high-precision input conditions for subsequent spatial coordinate calculation and multi-source information fusion. Specific steps may include: S1021: Obtain the spatial relationship between continuous positioning reference objects and auxiliary positioning reference objects in the previous frame positioning image; S1022: Construct a corrected coordinate system using the coordinate values ​​of the second camera, and determine the coordinate values ​​of the continuous positioning reference object in the corrected coordinate system based on the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object; S1023: Map the first camera coordinates of the continuous positioning reference object to the corrected coordinate system, and determine the pose correction information of each frame of positioning image based on the difference between the mapping result and the coordinates of the continuous positioning reference object in the corrected coordinate system.

[0043] In the embodiments S1021 to S1023, the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object refers to the geometric constraint relationship formed between them in the camera coordinate system. This geometric constraint relationship includes relative positional relationship, relative distance relationship, and directional relationship. The relative positional relationship can be characterized by the offset of the auxiliary positioning reference object relative to the continuous positioning reference object in the horizontal and vertical directions, thereby reflecting their relative arrangement in the coordinate system. The relative distance relationship can be represented by the Euclidean distance or projected distance between them in the camera coordinate system, used to reflect their relative distance. The directional relationship can be reflected by the angle between the line connecting them and the camera principal axis or image coordinate axis, thereby revealing the camera's attitude changes. Furthermore, when the number of candidate reference objects is greater than one, specific geometric structural relationships can also be formed between the continuous positioning reference object and multiple auxiliary positioning reference objects.

[0044] First, the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object in the previous frame's positioning image is obtained. Then, using the second camera coordinates of the auxiliary positioning reference object in camera coordinates, a corrected coordinate system is constructed. The coordinates of the continuous positioning reference object are then redefined in this corrected coordinate system, optimizing the coordinate results by incorporating the stability characteristics of the auxiliary reference object. Finally, the first camera coordinates of the continuous positioning reference object are mapped to the corrected coordinate system, and the difference between this mapped coordinate value and the true coordinate value in the corrected coordinate system is calculated. This difference reflects the camera pose offset between the current frame and the previous frame. Ultimately, pose correction information is generated based on this difference result.

[0045] S103: Using the continuous positioning reference object as the first index target, determine the first spatial coordinate value of each frame of positioning image; using the auxiliary positioning reference object as the second index target, determine the second spatial coordinate value of each frame of positioning image; and correct the first spatial coordinate value according to the second spatial coordinate value to determine the corrected spatial coordinate value of each frame of positioning image.

[0046] In this embodiment, after determining the pose correction information of the positioning image, the target spatial coordinates corresponding to each frame of the positioning image can be further calculated using the corrected spatial coordinates. The target spatial coordinates can specifically include latitude and longitude coordinates, geographic elevation information, or other geographic coordinate data that matches the GIS system. In this way, pixel-level features of the image reference object can be mapped to the actual geographic spatial location. The specific steps may include: S1031: Using the feature information of the continuously positioned reference object as the first index target, find the corresponding first target spatial region, and use the position information of the first target spatial region as the first spatial coordinate value; S1032: Using the feature information of the auxiliary positioning reference object as the second index target, find the corresponding second target spatial region, and use the position information of the second target spatial region as the second spatial coordinate value; the feature information includes at least one of contour shape, texture features, or semantic label.

[0047] In the implementations of S1031 to S1032, the feature information (including at least one or more of contour shape, texture features, or semantic labels) extracted from the image of the continuous positioning reference object and the auxiliary positioning reference object is first normalized and vectorized. The features can be composed of semantic labels output by a pre-trained detection / segmentation model, shape descriptors based on geometric contours, or feature descriptors based on local / global textures. Then, the feature description is used as an index to retrieve matching target spatial regions in a spatial object database or GIS base database. The retrieval process adopts a multi-layer filtering strategy, that is, firstly, attribute filtering is performed according to semantic category, expected scale, and observable orientation, then candidate regions are quickly returned through the spatial index structure, and then fine-grained feature similarity comparison is performed on the candidate regions to obtain matching scores. For each candidate match, spatial rationality (e.g., whether it is consistent with the positional relationship of camera height, viewpoint, and prior map features) and temporal consistency (whether the matching results with adjacent frames are consistent) are also evaluated. Feature similarity, spatial rationality, and temporal consistency are combined into the final matching confidence score. When the highest confidence score exceeds the preset threshold, the system uses the representative location information of the selected candidate region (e.g., region centroid, key point coordinates, or labeled anchor points) as the first or second spatial coordinate value and records its source ID and confidence score. At the same time, the location information is projected and labeled according to the agreed geographic reference system (e.g., WGS84 latitude and longitude and elevation).

[0048] S104: Determine the target spatial coordinates of each frame of the positioning image based on the fusion result of the corrected spatial coordinates and attitude correction information.

[0049] In this embodiment, after determining the corrected spatial coordinates of the positioning image, the corrected spatial coordinates are fused with the attitude correction information to obtain the target spatial coordinates of each frame of the positioning image. Specifically, the corrected spatial coordinates are static position information obtained based on image feature matching and spatial region indexing, while the attitude correction information reflects the attitude parameters of the capturing device at that moment (e.g., pitch angle, yaw angle, roll angle) and possible sensor drift compensation. During the fusion process, the system first projects the corrected spatial coordinates onto a unified reference coordinate system consistent with the attitude information using a coordinate transformation matrix, and then rotates, translates, or scales the position deviation according to the attitude correction information to correct the spatial coordinate offset caused by changes in device attitude. Specific steps may include: S1041: Determine the camera-space correction coefficient for each frame of the positioning image based on the pose correction information of each frame of the positioning image. S1042: Adjust the corrected spatial coordinate values ​​of each frame of the positioning image according to the camera-space correction coefficient to determine the target spatial coordinate values ​​of each frame of the positioning image.

[0050] In the implementations of S1041 to S1042, firstly, the attitude change estimate of the current frame camera is extracted from the attitude correction information, including pitch, yaw, roll, and possible small translations or sensor drift. This attitude change is then projected onto the measurement domain of the camera-to-space mapping to obtain a theoretical attitude-based mapping. Secondly, using the observed coordinates of the identified reference object in the image in the current frame and the reference frame, the residuals between the theoretical mapping and the actual observations are calculated. These residuals directly reflect the deviation of the existing attitude estimate in the spatial mapping. Then, using the residuals as constraints, a set of correction parameters that minimize the residuals, i.e., the camera-space correction coefficients, is solved using robust regression or least squares optimization. The corrected spatial coordinate values ​​are then regarded as the initial position information in the current coordinate system. These coordinate values ​​reflect the mapping result of the image in the target space after preliminary correction. Next, an attitude compensation matrix is ​​constructed using the camera-space correction coefficients. This matrix is ​​typically a combination of rotation matrices and translation vectors, capable of characterizing the deviation relationship between the camera coordinate system and the target space coordinate system. The compensation matrix is ​​applied to the corrected spatial coordinates of each frame of the localization image, and a linear or nonlinear transformation is performed on it to correct the offset caused by pose drift, perspective distortion, or spatial mapping errors. In this process, for cases with multiple corrections, rotation compensation can be performed first, followed by translation correction, to gradually obtain target coordinates consistent with the actual space. Furthermore, weighting factors or confidence information can be combined to differentiate the correction magnitude for different dimensions to avoid overcompensation or the introduction of new errors.

[0051] In one feasible implementation, the method further includes: Based on the target spatial coordinates of each frame of the localization image, generate localization marker results; Add the corresponding positioning marker result to each frame of the positioning image.

[0052] In this embodiment, after obtaining the target spatial coordinates of each frame of the positioning image, these coordinates can be associated with a preset geographic information system or spatial database to generate a positioning marker result corresponding to each coordinate value. This positioning marker result can be represented as a geographic coordinate point, a graphic symbol, a text label, or a combination thereof, used to intuitively indicate the target spatial location. Furthermore, when rendering the positioning image, the generated positioning marker result is superimposed onto the image. The superposition method can include drawing a positioning icon at the location corresponding to the target spatial coordinate point in the image, or adding textual explanatory information near the icon to help the user understand the actual location information of the image. In this way, each frame of the positioning image not only has visual content but also includes annotation information related to the real spatial location, thereby achieving a fusion display of visual information and spatial positioning information, allowing users to intuitively obtain the precise spatial location of the image while viewing it.

[0053] This application achieves video GIS positioning and spatial annotation by establishing spatial identifiers for video data lacking explicit location information. This expands the traditional video system's reliance on a single dimension of temporal retrieval to a dual-dimensional retrieval method encompassing both time and space. Specifically, each frame of the positioning image, through continuous positioning reference object and auxiliary positioning reference object spatial feature extraction, pose correction, and spatial coordinate correction fusion, generates high-precision target spatial coordinates and adds positioning markers to the video, achieving a correspondence between image content and actual geographical location. This method significantly improves the utilization rate of video data without location information, enabling users to quickly locate and retrieve relevant images based on geographical location in large-scale video data environments, avoiding the problems of time-consuming, prone to omissions, and inaccurate positioning associated with traditional manual segment-by-segment retrieval. Furthermore, the method in this application enhances the continuity and robustness of positioning through multi-frame information fusion and pose correction, maintaining stable spatial annotation results even when reference objects are partially missing, scenes are similar, or lighting is complex. This provides a reliable spatial foundation for subsequent event analysis, path reconstruction, and geographic information applications, improving the efficiency and intelligence level of video management and utilization.

[0054] This invention also provides a video GIS positioning device, referring to... Figure 2 The diagram illustrates a functional block diagram of a video GIS positioning device according to the present invention. The device may include the following modules: The acquisition module 201 is used to acquire the video to be processed, split the video to be processed into multiple frames of positioning images, perform filtering processing on each frame of positioning images, and determine the continuous positioning reference object and auxiliary positioning reference object in each frame of positioning images. The first correction module 202 is used to determine the first camera coordinate value of the continuous positioning reference object in the camera coordinate system and the second camera coordinate value of the auxiliary positioning reference object in the camera coordinate system, and to correct the first camera coordinate value according to the second camera coordinate value, thereby determining the pose correction information of each frame of positioning image. The second correction module 203 is used to determine the first spatial coordinate value of each frame of positioning image by taking the continuous positioning reference object as the first index target, and to determine the second spatial coordinate value of each frame of positioning image by taking the auxiliary positioning reference object as the second index target, and to correct the first spatial coordinate value according to the second spatial coordinate value, and to determine the corrected spatial coordinate value of each frame of positioning image. The spatial coordinate determination module 204 is used to determine the target spatial coordinate value of each frame of positioning image based on the fusion result of the corrected spatial coordinate value and the attitude correction information of each frame of positioning image.

[0055] In one possible implementation, the acquisition module includes: The first acquisition submodule is used to acquire the positioning accuracy of video GIS positioning and determine the image segmentation interval based on the positioning accuracy. The image segmentation interval is negatively correlated with the positioning accuracy. The splitting submodule is used to split the video to be processed into multiple frames of positioning images according to the image splitting interval.

[0056] In one possible implementation, the acquisition module further includes: The determination submodule is used to determine candidate reference objects in each frame of the localization image; The first filtering submodule is used to determine continuously located reference objects based on the evaluation results of the stability and continuity of candidate reference objects in adjacent frame images. The second filtering submodule is used to determine the auxiliary positioning reference object based on the evaluation results of the identifiability and positioning accuracy of the candidate reference object in each frame of the image.

[0057] In one possible implementation, the first filtering submodule includes: The first evaluation unit is used to determine the stability evaluation score of each candidate reference object in adjacent frame images. The more times the candidate reference object appears in multiple adjacent frame images, the higher the stability evaluation score. The second evaluation unit is used to determine the continuity evaluation score of each candidate reference object in adjacent frame images. The smaller the inter-frame position offset and the smoother the trajectory of the candidate reference object in multiple adjacent frame images, the higher the continuity evaluation score. The third evaluation unit is used to determine the continuous positioning reference object based on the weighted fusion result of the stability evaluation score and the continuity evaluation score.

[0058] In one possible implementation, the first correction module includes: The spatial relationship acquisition submodule is used to acquire the spatial relationship between continuous positioning reference objects and auxiliary positioning reference objects in the previous frame positioning image; The reference system determination submodule is used to construct a corrected coordinate system using the coordinate values ​​of the second camera, and to determine the coordinate values ​​of the continuous positioning reference object in the corrected coordinate system based on the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object. The correction submodule is used to map the first camera coordinate values ​​of the continuously positioned reference object to the corrected coordinate system, and determine the pose correction information of each frame of the positioning image based on the difference between the mapping result and the coordinate values ​​of the continuously positioned reference object in the corrected coordinate system.

[0059] In one possible implementation, the second correction module includes: The first search submodule is used to search for the corresponding first target spatial region using the feature information of the continuously located reference object as the first index target, and to use the location information of the first target spatial region as the first spatial coordinate value. The second search submodule is used to find the corresponding second target spatial region by using the feature information of the auxiliary positioning reference object as the second index target, and to use the position information of the second target spatial region as the second spatial coordinate value; the feature information includes at least one of contour shape, texture features, or semantic label.

[0060] In one possible implementation, the spatial coordinate determination module includes: The correction coefficient determination submodule is used to determine the camera-space correction coefficients for each frame of the positioning image based on the pose correction information of each frame of the positioning image. The adjustment submodule is used to adjust the corrected spatial coordinate values ​​of each frame of the positioning image according to the camera-space correction coefficient, and to determine the target spatial coordinate values ​​of each frame of the positioning image.

[0061] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; The processor, when executing a program stored in memory, implements the video GIS positioning method of the present invention.

[0062] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned terminal and other devices. The memory can include Random Access Memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory can also be at least one storage system located remotely from the aforementioned processor.

[0063] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0064] Furthermore, to achieve the above objectives, embodiments of the present invention also propose a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the video GIS positioning method of the present invention.

[0065] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable vehicles (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0066] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0069] Finally, it should be noted that in this document, relational terms such as "second" and "other" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. "" and / or "" indicate that either one or both can be selected.

[0070] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A video GIS positioning method, characterized in that, The method includes: The video to be processed is acquired and split into multiple frames of positioning images. Each frame of the positioning image is filtered to determine the continuous positioning reference object and the auxiliary positioning reference object in each frame of the positioning image. The first camera coordinate value of the continuous positioning reference object in camera coordinates and the second camera coordinate value of the auxiliary positioning reference object in camera coordinates are determined, and the first camera coordinate value is corrected according to the second camera coordinate value to determine the pose correction information of each frame of the positioning image. Using the continuous positioning reference object as the first index target, the first spatial coordinate value of each frame of the positioning image is determined; using the auxiliary positioning reference object as the second index target, the second spatial coordinate value of each frame of the positioning image is determined; and the first spatial coordinate value is corrected according to the second spatial coordinate value to determine the corrected spatial coordinate value of each frame of the positioning image. The target spatial coordinates of each frame of the positioning image are determined based on the fusion result of the corrected spatial coordinates and the attitude correction information.

2. The video GIS positioning method according to claim 1, characterized in that, The step of acquiring the video to be processed and splitting the video to be processed into multiple positioning images includes: The positioning accuracy of video GIS positioning is obtained, and the image segmentation interval is determined based on the positioning accuracy. The image segmentation interval is negatively correlated with the positioning accuracy. The video to be processed is divided into multiple positioning images according to the image splitting interval.

3. The video GIS positioning method according to claim 1, characterized in that, The step of filtering each frame of the positioning image to determine the continuous positioning reference object and auxiliary positioning reference object in each frame of the positioning image includes: Identify candidate reference objects in each frame of the localization image; Based on the stability and continuity evaluation results of the candidate reference objects in adjacent frame images, continuously located reference objects are determined; Based on the evaluation results of the identifiability and positioning accuracy of the candidate reference object in each frame of the image, an auxiliary positioning reference object is determined.

4. The video GIS positioning method according to claim 3, characterized in that, The step of determining continuously located reference objects based on the stability and continuity evaluation results of the candidate reference objects in adjacent frame images includes: Determine the stability evaluation score of each candidate reference object in adjacent frame images. The more times the candidate reference object appears in multiple adjacent frame images, the higher the stability evaluation score. The continuity evaluation score of each candidate reference object in adjacent frame images is determined. The smaller the inter-frame positional offset and the smoother the trajectory of the candidate reference object in multiple adjacent frame images, the higher the continuity evaluation score. The continuous positioning reference object is determined based on the weighted fusion result of the stability evaluation score and the continuity evaluation score.

5. The video GIS positioning method according to claim 1, characterized in that, The step of correcting the first camera coordinates based on the second camera coordinates to determine the pose correction information for each frame of the positioning image includes: Obtain the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object in the previous frame positioning image; A corrected coordinate system is constructed using the coordinate values ​​of the second camera, and the coordinate values ​​of the continuous positioning reference object in the corrected coordinate system are determined based on the spatial relationship between the continuous positioning reference object and the auxiliary positioning reference object. The first camera coordinates of the continuous positioning reference object are mapped to the corrected coordinate system, and the attitude correction information of each frame of the positioning image is determined based on the difference between the mapping result and the coordinates of the continuous positioning reference object in the corrected coordinate system.

6. The video GIS positioning method according to claim 1, characterized in that, The step of determining the first spatial coordinate value of each frame of the positioning image using the continuous positioning reference object as the first index target and the second spatial coordinate value of each frame of the positioning image using the auxiliary positioning reference object as the second index target includes: Using the feature information of the continuous positioning reference object as the first index target, find the corresponding first target spatial region, and use the position information of the first target spatial region as the first spatial coordinate value; Using the feature information of the auxiliary positioning reference object as the second index target, find the corresponding second target spatial region, and use the position information of the second target spatial region as the second spatial coordinate value; the feature information includes at least one of contour shape, texture features, or semantic label.

7. The video GIS positioning method according to claim 1, characterized in that, The step of determining the target spatial coordinate value of each frame of the positioning image based on the fusion result of the corrected spatial coordinate value and the pose correction information includes: Based on the pose correction information of each frame of the positioning image, determine the camera-space correction coefficient of each frame of the positioning image; The corrected spatial coordinate values ​​of each frame of the positioning image are adjusted according to the camera-space correction coefficient to determine the target spatial coordinate values ​​of each frame of the positioning image.

8. The video GIS positioning method according to claim 1, characterized in that, The method further includes: Based on the target spatial coordinates of each frame of the positioning image, a positioning marker result is generated; Add the corresponding positioning marker result to each frame of the positioning image.

9. A video GIS positioning device, characterized in that, The apparatus for implementing the method according to any one of claims 1-8, the apparatus comprising: The acquisition module is used to acquire the video to be processed, split the video to be processed into multiple frames of positioning images, perform filtering processing on each frame of the positioning image, and determine the continuous positioning reference object and auxiliary positioning reference object in each frame of the positioning image. The first correction module is used to determine the first camera coordinate value of the continuous positioning reference object in camera coordinates and the second camera coordinate value of the auxiliary positioning reference object in camera coordinates, and to correct the first camera coordinate value according to the second camera coordinate value, thereby determining the pose correction information of each frame of the positioning image. The second correction module is used to determine the first spatial coordinate value of each frame of the positioning image by taking the continuous positioning reference object as the first index target, and to determine the second spatial coordinate value of each frame of the positioning image by taking the auxiliary positioning reference object as the second index target, and to correct the first spatial coordinate value according to the second spatial coordinate value, thereby determining the corrected spatial coordinate value of each frame of the positioning image. The spatial coordinate determination module is used to determine the target spatial coordinate value of each frame of the positioning image based on the fusion result of the corrected spatial coordinate value of each frame of the positioning image and the attitude correction information.

10. The video GIS positioning device according to claim 9, characterized in that, The acquisition module includes: The first acquisition submodule is used to acquire the positioning accuracy of video GIS positioning and determine the image splitting interval based on the positioning accuracy. The image splitting interval is negatively correlated with the positioning accuracy. The splitting submodule is used to split the video to be processed into multiple positioning images according to the image splitting interval.