A method and system for mapping video frame plane coordinates to model latitude and longitude coordinates
By acquiring multi-view video frames and building model data, extracting key frames and generating spatial constraint signals, and combining them with geographic coordinate reference signals for inversion calculation, the problems of low mapping accuracy and poor reliability in existing technologies are solved, and the accurate conversion from video frame planar coordinates to digital twin model latitude and longitude coordinates is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-03-20
AI Technical Summary
Existing methods for mapping video frame planar coordinates to model latitude and longitude coordinates rely on manually marked feature points, which is inefficient and prone to introducing errors. They are also difficult to adapt to the complex spatial topology of buildings, resulting in insufficient mapping accuracy and reliability.
By acquiring multi-view video frames and building model data, key frames are extracted and spatial constraint signals are generated. Combined with geographic coordinate reference signals, inverse calculations are performed to achieve accurate conversion from planar coordinates to latitude and longitude coordinates.
It improves mapping accuracy and reliability, adapts to complex indoor spaces and multi-view scenarios, and achieves precise association between real-world scenarios and digital twin models.
Smart Images

Figure CN121504709B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of building coordinate mapping, and in particular to a video frame plane coordinate and model latitude and longitude coordinate mapping method and system. BACKGROUND
[0002] The video frame plane coordinate and model latitude and longitude coordinate mapping method is one of the key technologies for realizing accurate association between a real scene and a digital twin model. The method can be widely applied in the fields of building operation and maintenance, intelligent security and protection, emergency rescue, etc., and provides core data support for scene restoration and accurate management and control of a digital twin system, and has a very broad application prospect.
[0003] At present, in the existing digital twin field, the mapping of video frame plane coordinates and model latitude and longitude coordinates mainly relies on manual marking of feature points to establish association, or is realized by using a simple image matching algorithm combined with rough geographic reference data. This kind of method usually extracts a small number of feature points from a video frame, manually matches the corresponding positions of a digital twin model, and then completes coordinate conversion.
[0004] However, the existing technical method is limited by a lot of manual intervention, which not only has low efficiency, but also is prone to cause insufficient mapping accuracy due to human errors; at the same time, it has poor adaptability to complex spatial topological structures in a building indoor space, and in the cooperative mapping scene of multi-view video frames, it is difficult to form a stable spatial constraint relationship, resulting in poor accuracy and reliability of coordinate mapping. SUMMARY
[0005] The present application aims to provide a video frame plane coordinate and model latitude and longitude coordinate mapping method and system to solve the problems of low mapping accuracy and poor reliability of video frame plane coordinates and model latitude and longitude coordinates in the prior art.
[0006] To solve the above technical problems, in a first aspect, the present application provides a video frame plane coordinate and model latitude and longitude coordinate mapping method, comprising:
[0007] obtaining multi-view video frames of a target building indoor space and building model data corresponding to the target building, wherein the building model data includes indoor spatial topological data and floor layout data;
[0008] extracting key frames from the multi-view video frames and extracting coordinate reference information of walls and passages from the building model data to form a structured data set;
[0009] processing the key frames in the data set to generate feature data of indoor fixed structures, and associating the feature data of the indoor fixed structures with the coordinate reference information of the walls and passages to generate a spatial constraint signal for defining the spatial positions of pixel points in the key frames;
[0010] Receive the planar coordinates of the keyframe and the spatial constraint signal, and perform inverse calculation on the planar coordinates based on the spatial constraint signal to obtain the coordinate values of the keyframe in the local coordinate system of the building interior;
[0011] Geographic coordinate reference signals are obtained from the building model data, and combined with the geographic coordinate reference signals, the coordinate values in the local coordinate system of the building interior are converted into the latitude and longitude coordinate values of the digital twin model corresponding to the target building.
[0012] Optionally, the feature data of the indoor fixed structure is associated with the coordinate reference information of the wall and passageway to generate a spatial constraint signal for defining the spatial position of pixels in the keyframe, including:
[0013] The feature point set in the feature data of the indoor fixed structure is matched point by point with the vertex coordinate set in the coordinate reference information of the wall and passage to generate a mapping table describing the point-to-point relationship.
[0014] Based on the mapping table, the transformation parameters from the image planar coordinates of the keyframe to the coordinates of the building's interior local coordinate system are calculated. The transformation parameters include a scale adjustment factor and a position translation amount.
[0015] Based on the transformation parameters, a spatial constraint signal is generated, which defines the permissible position boundary of the pixel in the local coordinate system of the building interior.
[0016] Optionally, based on the mapping table, the transformation parameters from the image planar coordinates of the keyframe to the coordinates of the building's interior local coordinate system are calculated, including:
[0017] Based on the mapping table, multiple point pair combinations are extracted, each point pair combination containing the image plane coordinates and the corresponding building interior local coordinate system coordinates;
[0018] Based on the coordinate values in the point pair combination, the length ratio of the line segment between each pair of corresponding points in the image plane coordinate system and the building interior local coordinate system is calculated to generate a scale adjustment factor.
[0019] Based on the coordinate values in the selected point pair combination, calculate the geometric center point coordinates of the image plane coordinates and the geometric center point coordinates of the building interior local coordinate system, respectively, to generate the position translation amount;
[0020] The scale adjustment factor and the position translation amount are combined to form the transformation parameters used for coordinate transformation.
[0021] Optionally, the key frames in the data set are processed to generate feature data of the indoor fixed structure, including:
[0022] By analyzing the image pixel array of each key frame, the contour boundary of the indoor fixed structure is identified, and a sequence of feature point coordinates of the contour boundary is extracted;
[0023] Based on the sequence of feature point coordinates, a feature descriptor describing the shape of the structure is calculated, and the feature descriptor contains the relative position relationship of the point set;
[0024] The feature descriptors of all key frames are integrated to generate feature data of the indoor fixed structure, and the feature data is used to represent the spatial characteristics of the fixed structure.
[0025] Optionally, the planar coordinates of the key frames and the spatial constraint signal are received, and the planar coordinates are inversely calculated based on the spatial constraint signal to obtain the coordinate values of the key frames in the local coordinate system of the building interior, including:
[0026] The planar coordinates composed of the row number and column number of the pixel points of the key frames are received, and the conversion parameters defined in the spatial constraint signal are obtained;
[0027] According to the conversion parameters, a coordinate inverse transformation formula is constructed, the planar coordinates are substituted into the coordinate inverse transformation formula for calculation, and the three-dimensional position data corresponding to each pixel point is obtained;
[0028] The three-dimensional position data of all pixel points is summarized to form the coordinate values of the key frames in the local coordinate system of the building interior.
[0029] Optionally, the geographic coordinate reference signal is obtained from the building model data, and the coordinate values in the local coordinate system of the building interior are converted into the latitude and longitude coordinate values of the digital twin model corresponding to the target building in combination with the geographic coordinate reference signal, including:
[0030] The latitude and longitude values of the pre-defined geographic reference point corresponding to the actual geographic location of the target building are read from the structured attribute field of the building model data, and a reference signal containing latitude and longitude data is generated;
[0031] Based on the reference signal, the coordinate conversion parameters between the local coordinate system of the building interior and the geographic coordinate system are calculated, including the offset and direction angle of the origin of the local coordinate system in the geographic coordinate system;
[0032] Using the coordinate conversion parameters, the point-by-point coordinate transformation operation is performed on the coordinate values in the local coordinate system of the building interior to generate the latitude and longitude coordinate values of the digital twin model.
[0033] Optionally, key frames are extracted from the multi-view video frames, and coordinate reference information of walls and passages is extracted from the building model data to form a structured data set, including:
[0034] Based on the sequence of multi-view video frames, the matching degree of feature points between adjacent video frames is calculated to generate a feature difference value;
[0035] When the feature difference value is greater than a preset threshold, the current video frame is marked as a key frame, and a key frame set is stored;
[0036] From the indoor space topology data of the building model data, the polygon structure of the walls and passages is parsed, and the corner point coordinates of the polygon structure are extracted to form coordinate reference information;
[0037] The key frame set is associated and combined with the coordinate reference information to generate a structured data set, wherein each key frame is mapped to a corresponding subset of coordinate reference information.
[0038] In a second aspect, the application provides a video frame plane coordinate and model latitude and longitude coordinate mapping system, comprising:
[0039] An acquisition module is configured to acquire multi-view video frames of a target building indoor space and building model data corresponding to the target building, wherein the building model data includes indoor space topology data and floor layout data;
[0040] An extraction module is configured to extract key frames from the multi-view video frames and extract coordinate reference information of walls and passages from the building model data to form a structured data set;
[0041] An association module is configured to process the key frames in the data set to generate feature data of indoor fixed structures, and associate the feature data of the indoor fixed structures with the coordinate reference information of the walls and passages to generate a space constraint signal for defining the spatial position of pixel points in the key frames;
[0042] An inversion module is configured to receive plane coordinates of the key frames and the space constraint signal, and perform inversion calculation on the plane coordinates based on the space constraint signal to obtain coordinate values of the key frames in a building indoor local coordinate system;
[0043] A conversion module is configured to obtain a geographic coordinate reference signal from the building model data, and convert the coordinate values in the building indoor local coordinate system into latitude and longitude coordinate values of a digital twin model corresponding to the target building in combination with the geographic coordinate reference signal.
[0044] In a third aspect, the present application provides an electronic device, comprising:
[0045] a memory for storing a computer program;
[0046] a processor for implementing the steps of the mapping method of video frame plane coordinates and model latitude and longitude coordinates according to the first aspect when executing the computer program.
[0047] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program can implement the steps of the mapping method of video frame plane coordinates and model latitude and longitude coordinates according to the first aspect when executed by a processor.
[0048] The mapping method of video frame plane coordinates and model latitude and longitude coordinates provided by the present application can provide comprehensive and basic data support for subsequent coordinate mapping by obtaining multi-view video frames in a target building indoor and building model data containing indoor space topology data and floor layout data; can realize data regularization processing and facilitate subsequent efficient correlation operation by extracting key frames from the multi-view video frames and extracting coordinate reference information of walls and passages from the building model data to form a structured data set; can provide accurate spatial boundary limitation basis for coordinate conversion by processing the key frames to generate feature data of indoor fixed structures and associating the feature data with the coordinate reference information of walls and passages to generate a space constraint signal limiting the spatial position of a pixel point; can realize accurate conversion of plane coordinates to local space coordinates by receiving the plane coordinates of the key frames and performing inversion calculation on the plane coordinates based on the space constraint signal to obtain coordinate values in a local coordinate system; and can complete the final mapping from the image plane to the geographic coordinates of the digital twin model and realize accurate association of the real scene and the digital twin model by converting the coordinate values in the local coordinate system to latitude and longitude coordinate values of the digital twin model in combination with the geographic coordinate reference signal obtained from the building model data.
[0049] Further, the feature point set in the feature data of the indoor fixed structures is matched with the vertex coordinate set in the coordinate reference information of the walls and passages point by point to generate a mapping table, conversion parameters of the image plane coordinates to the indoor local coordinate system coordinates are calculated based on the mapping table, and a space constraint signal defining the allowable position boundary of the pixel point in the indoor local coordinate system is generated according to the conversion parameters. The mapping table generated by point-by-point matching can ensure the accurate correspondence of the feature data and the coordinate reference information; the accurate conversion parameters can provide reliable basis for the generation of the space constraint signal; and the finally generated space constraint signal can further improve the accuracy of the pixel point spatial position limitation and guarantee the accuracy of the subsequent coordinate inversion calculation. BRIEF DESCRIPTION OF DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments or prior art of the present application, the drawings needed to be used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0051] Figure 1 A flowchart of a video frame plane coordinate and model latitude and longitude coordinate mapping method provided by an embodiment of the present application;
[0052] Figure 2 A specific implementation flowchart of a video frame plane coordinate and model latitude and longitude coordinate mapping method provided by an embodiment of the present application;
[0053] Figure 3 A structural schematic diagram of a video frame plane coordinate and model latitude and longitude coordinate mapping system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0054] In the application of video frame plane coordinate and model latitude and longitude coordinate mapping in the field of digital twinning, the existing method generally relies on manual marking of feature points or simple image matching combined with rough geographic reference data, which has significant defects: manual intervention leads to low efficiency and is prone to introduce errors, and at the same time, it has poor adaptability to complex indoor spatial topological structure of buildings, and it is difficult to form stable spatial constraints in multi-view collaborative mapping, which ultimately results in insufficient coordinate mapping precision and reliability. This problem is due to the fact that the existing scheme does not fully combine the building space structure to construct effective constraints, lacks systematic coordinate conversion logic, and cannot meet the actual application requirements, so a more reliable mapping scheme is urgently needed.
[0055] In view of the above problems, the present application proposes a video frame plane coordinate and model latitude and longitude coordinate mapping method, the core of which is to fuse multi-view video information and building space structure data in indoor buildings, to form stable spatial constraints by extracting key frame features and associating them with structure coordinate information such as walls and passages, and then to realize the mapping of video frame plane coordinates to digital twinning model latitude and longitude coordinates through systematic coordinate inversion algorithm and reference conversion. This method discards the mode of a large amount of manual intervention, and the constraint mechanism based on building space structure improves the mapping precision, and at the same time adapts to complex indoor and multi-view scenarios, fundamentally solving the problems of low mapping precision and poor reliability of the existing technology, and providing core support for precise application of digital twinning systems.
[0056] For those skilled in the art to better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0057] The core of the present application is to provide a mapping method of video frame plane coordinates and model latitude and longitude coordinates, and a specific embodiment of the flowchart is shown in Figure 1 The method comprises:
[0058] S101, obtaining a plurality of perspective video frames in a target building, and building model data corresponding to the target building.
[0059] Among them, the multi-perspective video frame refers to the continuous video picture taken from different positions and different angles in the target building, which can fully cover all areas of the indoor area and avoid the observation blind area of single perspective.
[0060] The building model data is a digital description data of the spatial form of the target building, which includes indoor space topology data and floor layout data. Among them, the indoor space topology data includes the connection relationship and distribution logic data of indoor walls, passages, rooms and other structures, and the floor layout data includes the intuitive layout information of room division and functional area distribution of each floor.
[0061] In a specific embodiment, the multi-perspective video frame can be obtained by deploying multiple monitoring cameras in the key areas such as corridors, entrances and exits, and elevator halls in the target building, and synchronously recording by means of clock synchronization technology to ensure the time consistency of each camera picture; the building model data can be extracted from the BIM (Building Information Model) of the target building, if there is no existing BIM model, the information can also be extracted from the audited design planning drawings, completion archives and other materials, and obtained after digitalization and arrangement, to ensure that the extracted data can accurately match the actual indoor situation of the target building.
[0062] S102, extracting key frames from the multi-perspective video frames, and extracting coordinate reference information of walls and passages from the building model data, to form a structured data set.
[0063] The key frame refers to a representative video picture in a sequence of multi-view video frames, can reflect the core features of the indoor scene and contains rich spatial information, and is used to reduce the redundancy of subsequent data processing. The coordinate reference information refers to basic data describing the spatial positions of the wall and the passageway, and provides a spatial positioning basis for the association of the video frame and the building model. The structured data set refers to a data set formed by sequentially organizing the key frame and the corresponding coordinate reference information, has clear semantic association and classification annotation, and is convenient for subsequent association operation and data calling.
[0064] Optionally, step S102 can specifically include the following steps:
[0065] S1021, based on the sequence of multi-view video frames, calculating the matching degree of feature points between adjacent video frames to generate a feature difference value.
[0066] In this step, the feature point refers to a point in the video frame image that has obvious recognition, such as the point of the wall corner, the door and window edge and the like, which can stably reflect the local features of the image, and is the core basis for subsequent frame matching. The feature difference value refers to a numerical value for quantifying the feature change degree between adjacent video frames. The greater the feature difference value, the more obvious the scene change between adjacent video frames, which can be used as a core index for judging whether it is a key frame.
[0067] Specifically, first, the feature points of adjacent video frames are extracted by the SIFT algorithm. The core of the algorithm is to obtain feature points with scale and rotation invariance through multi-scale analysis. The specific steps include scale space construction, extreme point detection, key point positioning, orientation and descriptor generation. The scale space construction needs to be realized by a scale space function, which is obtained by convolving a two-dimensional Gaussian function with the original video frame image. The formula is shown in formula (1):
[0068] (1)
[0069] In the formula, represents the scale space function value, which is used to represent the blur processing result of the original image at different scales; represents the pixel gray value of the original video frame image, and together represent the plane coordinates of the pixels in the original image; represents the two-dimensional Gaussian function value, which is the core function for realizing image scale blur; represents the scale parameter of the Gaussian function, which is used to control the blur degree of the image, the greater the value, the higher the image blur degree; represents the convolution operation, which is the operation mode for fusing and calculating the two-dimensional Gaussian function and the original image. Through this operation, image data at different scales can be obtained.
[0070] The two-dimensional Gaussian function expression is shown in formula (2):
[0071] (2)
[0072] In the formula, represents the function output value of the two-dimensional Gaussian function at the coordinate and the scale parameter ; and is a constant of the Gaussian function, and is approximately 3.14. represents the scale parameter, and has the same meaning as in formula (1); represents the square of the distance from the pixel coordinate to the image origin, and is used to represent the spatial position relationship of the pixel to the image center; exp represents the natural exponential function, and is used to realize the exponential decay operation on the square distance term; the entire fraction is a normalization coefficient, which is used to ensure that the integral value of the Gaussian function on the entire image plane is 1, so that the overall brightness of the image after the blur processing does not change. Through formula (2), different scale blur processing of the original image can be realized, so as to detect stable feature points under multiple scales.
[0073] After the feature points are extracted, a 128-dimensional feature point descriptor is generated, and the Euclidean distance between the feature point descriptors of adjacent frames is calculated to determine the matching degree. The smaller the Euclidean distance, the higher the matching degree. Finally, the number of matching successful feature points is counted, and the feature difference value is generated by subtracting the ratio of the number of matching successful feature points to the total number of feature points in the previous frame from 1, so as to quantify the degree of change of the feature points between adjacent frames.
[0074] For example, in an office building indoor monitoring scene, two adjacent video frames F1 and F2 are processed. First, the value range of the Gaussian function scale parameter σ is set to 1.6 to 6.4, and F1 and F2 are respectively convolved by the scale space function and the two-dimensional Gaussian function to construct a multi-scale space and detect extreme points. After key point positioning and orientation, a 128-dimensional descriptor is generated, and finally 150 feature points are extracted from F1 and 148 feature points are extracted from F2. The Euclidean distance between the feature point descriptors of the two frames is calculated, and the Euclidean distance threshold is set to 0.8. The matching pairs with a distance less than the threshold are selected, and a total of 120 matching successful feature points are obtained. According to the feature difference value generation logic, the ratio of the number of matching successful feature points to the total number of feature points in the previous frame is calculated first, and the calculation process is ; and 1-0.8=0.2, and finally the feature difference value is 0.2, which indicates that the scene change between F1 and F2 is relatively obvious. The above example is only an example of the present application, and in actual application, other feature extraction algorithms such as SURF and ORB can also be selected according to requirements, and the present application does not limit this.
[0075] S1022, when the feature difference value is greater than the preset threshold value, the current video frame is marked as a key frame, and a key frame set is stored.
[0076] In this step, the preset threshold value refers to a critical value set according to the actual situation of the indoor scene of the building, which is used to judge whether the scene change of adjacent video frames reaches the degree of marking a key frame, and is the core determination standard for screening key frames; the key frame set refers to a set formed after all key frames meeting the conditions are sorted and stored according to the time sequence of shooting, which can retain the core scene information in the video sequence with less data amount.
[0077] In the embodiment of the present application, first, a reasonable preset threshold value is set in combination with the complexity of the target indoor scene of the building, the frequency of personnel flow and other actual situations. Second, the feature difference values generated by S1021 are compared one by one with the preset threshold value, if the feature difference value is greater than the preset threshold value, it indicates that the current video frame has obvious scene change compared with the previous frame, and the current video frame needs to be marked as a key frame; if the feature difference value is less than or equal to the preset threshold value, it is not marked. Finally, all adjacent frame pairs in the multi-view video frame sequence are processed in turn according to the above rules, after the determination of all frames is completed, all video frames marked as key frames are sorted in chronological order, and a key frame set is stored.
[0078] Exemplarily, in the indoor monitoring scene of an office building, in combination with the common change situations such as daily personnel flow and object movement of the scene, the preset threshold value is set to 0.15 after comprehensive judgment. For the feature difference value 0.2 corresponding to F1 and F2 obtained in S1021, threshold comparison is performed, since 0.2 is greater than 0.15, the current video frame F2 is marked as a key frame. The subsequent adjacent video frame pairs are processed according to the above process. For example, the feature difference value of F2 and F3 is 0.12, which is less than the preset threshold value of 0.15, and F3 is not marked; the feature difference value of F3 and F4 is 0.18, which is greater than the preset threshold value, and F4 is marked as a key frame. After the determination of all video frames is completed, the marked key frames F2, F4 and the like are stored in chronological order of shooting, and a key frame set is formed.
[0079] S1023, from the indoor space topology data of the building model data, the polygon structure of the wall and the passageway is analyzed, and the corner point coordinates of the polygon structure are extracted to form coordinate reference information.
[0080] In this step, the polygon structure refers to the abstracted representation of the wall and passage space form in the building model, which accurately reflects the actual shape and distribution range of the wall and passage through the outline of the polygon; the corner point coordinate refers to the spatial coordinate value of the vertex of the polygon structure, which can accurately locate the outline boundary of the wall and passage, and is the core positioning data for establishing the spatial correlation between the video frame and the building model; the coordinate reference information refers to the data set formed by the corner point coordinates of the wall and passage, which provides a spatial positioning reference for the mapping of the video frame pixel and the building model coordinate.
[0081] In a specific embodiment, first, the building model data of the target building is obtained, from which the indoor space topology data containing the distribution relationship of the wall and passage space is screened out. Second, the screened data is analyzed by a professional contour extraction algorithm, and the spatial form of the wall and passage is converted into identifiable polygon structures.
[0082] Finally, the polygon structure obtained by analysis is processed by a corner point detection algorithm, the spatial coordinates of the vertex of each polygon, i.e., the corner point, are extracted, and all the corner point coordinates are sorted and verified to form the coordinate reference information after removing invalid or incorrect coordinates. If there is no ready-made building model data, the geometric information of the wall and passage can be extracted from architectural design drawings, completion archives and other materials, the indoor space topology data is obtained after digital modeling, and the coordinate reference information is analyzed and extracted according to the above process.
[0083] In actual application, in the indoor scene of an office building, the indoor space topology data is called from the BIM model of the building, the model API interface is called by a professional data extraction tool, and the topology data corresponding to the wall and passage of the office area, corridor and other areas is screened out. The Canny edge detection algorithm is used for contour extraction, the wall of the office area is analyzed as a rectangular polygon structure, and the corridor passage is analyzed as a rectangular polygon structure. The Harris corner point detection algorithm is used to extract the corner points of the analyzed polygon structure, for example, the rectangular polygon structure of the wall of a certain office area extracts four corner point coordinates as 、 、 、 wherein, represents the horizontal direction coordinate value, represents the vertical direction coordinate value, represents the height direction coordinate value. All the corner point coordinates of the wall and passage are sorted by area to form complete coordinate reference information.
[0084] The above example is only one example of the present application, and in actual application, if there is no ready-made BIM model, information can be extracted from architectural design drawings, completion archives and other materials and indoor space topology data can be obtained after digital modeling, which is not limited by the present application.
[0085] S1024, combine the key frame set and the coordinate reference information to generate a structured data set, wherein each key frame is mapped to a corresponding coordinate reference information subset.
[0086] In this step, the coordinate reference information subset refers to the local data in the coordinate reference information corresponding to the scene photographed by a single key frame, which contains the corner point coordinates of all walls and passages in the scene, and can realize the accurate association of the key frame with a specific building area; the structured data set refers to a data set formed by organizing and storing the key frame and the corresponding coordinate reference information subset in a unified format, which has clear semantic association and classification annotation, and can be directly used for subsequent coordinate mapping operations.
[0087] Specifically, first, the shooting position and angle range of each key frame in the key frame set are determined, and the specific area of the building indoor corresponding to each key frame is matched in combination with the spatial distribution information of the building model. Secondly, the corner point coordinates corresponding to the area are selected from the complete coordinate reference information to form the coordinate reference information subset exclusive to each key frame. Finally, a unified association rule is adopted to establish a mapping relationship between each key frame and the corresponding coordinate reference information subset, and a structured data set is generated after being arranged in a predetermined format. The association combination can also be realized by establishing a data table, the field information of the data table is clear, and the association relationship between the key frame and the coordinate reference information is clear and traceable.
[0088] For example, in an office building indoor scene, the shooting parameters of each key frame in the key frame set are first obtained, wherein the shooting position of the key frame F2 is at the entrance of the office area A, and the angle range covers the east wall of the office area A and the adjacent corridor entrance; the shooting position of the key frame F4 is at the window edge of the office area B, and the angle range covers the south wall of the office area B and the door column; the shooting position of the key frame F6 is at the middle section of the central corridor, and the angle range covers the wall of the middle section of the corridor and the wall around the fire hydrant. In combination with the spatial distribution information of the building model, the building area corresponding to each key frame is determined. The corner point coordinates corresponding to each area are selected from the coordinate reference information to form the coordinate reference information subset exclusive to each key frame. The association is established in the form of a data table, and the specific association table is shown in Table 1:
[0089] Table 1
[0090]
[0091] After completing the association entry of all key frames and the corresponding coordinate reference information subsets in the above table format, a structured data set can be generated. The above example is only one example of the present application, and in actual application, the table fields or association organization form can also be adjusted according to the needs, which is not limited in the present application.
[0092] The present application solves the problems of chaotic association of video frame data and building space data and excessive redundant data in the conventional method. Compared with the conventional method, the present application can accurately screen key frames and extract effective space coordinate information, realize the ordered association of the two, and improve the efficiency and accuracy of data processing.
[0093] S103, processing the key frames in the data set to generate feature data of indoor fixed structures, and associating the feature data of the indoor fixed structures with the coordinate reference information of the walls and passages to generate a space constraint signal for defining the spatial positions of the pixel points in the key frames.
[0094] The indoor fixed structures refer to structural components with fixed indoor positions and forms, including walls, columns, door and window frames, fixed ceiling shapes, etc. The feature data refer to set data capable of representing the spatial characteristics of the indoor fixed structures, including the shape, contour, spatial position, etc. of the structures. The space constraint signal refers to a signal for defining the legal position range of the pixel points in the local coordinate system of the building indoors, which can avoid the deviation of the pixel point coordinate mapping from the actual space.
[0095] Optionally, as shown in Figure 2 S103 can specifically include the following steps:
[0096] S1031, identifying the contour boundary of the indoor fixed structures and extracting the feature point coordinate sequence of the contour boundary by analyzing the image pixel array of each key frame.
[0097] In this step, the image pixel array refers to the ordered arrangement of pixels in the key frame image, each pixel containing gray value or color value information, which is the basic unit of image analysis. The contour boundary refers to the boundary line between the indoor fixed structures and the surrounding environment in the image, which can reflect the basic shape of the structures. The feature point coordinate sequence refers to the image plane coordinate set of the contour boundary feature points arranged in a certain order, which is used to accurately describe the shape and trend of the contour.
[0098] Specifically, first, the image pixel array of each key frame is preprocessed by gray scaling and Gaussian filtering to remove image noise and enhance image contrast to highlight the differences between the fixed structures and the environment. Second, the Canny edge detection algorithm is used to analyze the preprocessed pixel array to identify the contour boundary of the indoor fixed structures. Finally, based on the contour boundary, the Harris corner detection algorithm is used to extract the feature points on the contour, record the image plane coordinates of each feature point in a clockwise order, and form the feature point coordinate sequence.
[0099] For example, in the office building indoor monitoring scene, the key frame F002 corresponding to the office area A and the adjacent corridor entrance is selected for processing. First, the color image of the key frame is processed to obtain a grayscale image; then, the Gaussian filter is used to remove the noise in the image, with the Gaussian kernel size set to 5x5 and σ=1.4. The Canny edge detection algorithm is used to identify the contour boundaries of the east wall of the office area A and the walls on both sides of the corridor entrance from the preprocessed image, with the high threshold set to 150 and the low threshold set to 50. The Harris corner point detection algorithm is used to extract the feature points of these contour boundaries, and the coordinate sequence of the feature points in clockwise order is (200, 150), (300, 150), (300, 240), (200, 240), (300, 150), (320, 150), (320, 240), and (300, 240), wherein the first four coordinates correspond to the contour of the east wall of the office area A, and the last four coordinates correspond to the contour of the walls on both sides of the corridor entrance. The above example is only an example of the present application, and in actual application, the Sobel edge detection algorithm, the SIFT feature point extraction algorithm, etc. can also be selected, and the present application does not limit this.
[0100] S1032, based on the feature point coordinate sequence, calculating a feature descriptor describing the shape of the structure.
[0101] In this step, the feature descriptor refers to the vector data used to quantitatively describe the shape characteristics of the indoor fixed structure, and the relative position relationship of the point set contained therein can reflect the spatial correlation information such as the distance and angle between the feature points, and can realize the matching identification of the same structure in different key frames.
[0102] Specifically, first, a reference feature point is selected from the feature point coordinate sequence, and the relative coordinates of all other feature points in the sequence relative to the reference feature point are calculated. Then, based on the relative coordinates, the Euclidean distances and angles between the feature points are calculated, and these distances and angles are arranged in a fixed order to form a feature vector. Finally, the feature vector is normalized to obtain a standardized feature descriptor, ensuring that it is not affected by image scaling and rotation.
[0103] For example, in the office building indoor monitoring scenario, for the feature point coordinate sequence of the key frame F002, the first feature point (200, 150) is selected as the reference feature point. The relative coordinates of other feature points relative to the reference point are calculated: (300-200, 150-150) = (100, 0), (300-200, 240-150) = (100, 90), (200-200, 240-150) = (0, 90), (300-200, 150-150) = (100, 0), (320-200, 150-150) = (120, 0), (320-200, 240-150) = (120, 90), (300-200, 240-150) = (100, 90). The Euclidean distances between adjacent relative coordinates are calculated, such as the distance between (100, 0) and (100, 90) is , the distance between (100, 90) and (0, 90) is , and the distance sequence [90, 100, 90, 0, 20, 90, 20] is obtained by sequentially calculating. The angles between the lines connecting adjacent feature points and the reference point are calculated to obtain the angle sequence [0°, 90°, 180°, 0°, 0°, 90°, 180°]. The distance sequence and the angle sequence are combined to form a 14-dimensional feature vector, and after normalization, the feature descriptor of the fixed structure in the key frame F002 is obtained.
[0104] S1033, integrate the feature descriptors of all key frames to generate feature data of the indoor fixed structure.
[0105] In this step, integration refers to the process of collecting, deduplicating, and fusing feature descriptors corresponding to the same indoor fixed structure in different key frames; feature data is used to represent the spatial characteristics of the fixed structure, covering shape, contour, and other information of the structure under different viewing angles.
[0106] Specifically, first, similarity matching is performed on the feature descriptors of all key frames, and by calculating the Euclidean distance between the feature descriptors, the feature descriptors with a distance less than a preset similarity threshold are determined as corresponding to the same indoor fixed structure. Second, the feature descriptors determined as the same structure are deduplicated to eliminate repeated or redundant description information. Finally, the deduplicated feature descriptors are fused to supplement the feature information under different viewing angles, forming feature data that can fully represent the spatial characteristics of the fixed structure.
[0107] For example, in the office building indoor monitoring scene, the feature descriptors of the key frames F002, F004, F006 are selected for processing. The feature descriptor similarity threshold is set to 0.2, and the Euclidean distance between the feature descriptors in the three key frames is calculated. It is found that the feature descriptor corresponding to the east wall of office area A in F002 and the feature descriptor corresponding to the south wall of office area B in F004 have a distance of 0.8, which is greater than the threshold, and are not the same structure. The feature descriptor corresponding to the corridor entrance wall in F002 and the feature descriptor corresponding to the central corridor wall in F006 have a distance of 0.3, which is greater than the threshold, and are not the same structure. At the same time, the feature descriptors of different structures in each key frame have a distance greater than the threshold.
[0108] After removing the repeated description information caused by slight changes in the shooting angle, the feature descriptors corresponding to each structure are fused to generate the feature data of the indoor fixed structure, which includes the complete spatial characteristic information of multiple structures such as the east wall of office area A, the walls on both sides of the corridor entrance, and the south wall of office area B.
[0109] S1034, the feature point set in the feature data of the indoor fixed structure is matched with the vertex coordinate set in the coordinate reference information of the wall and the passage, and a mapping table describing the point-to-point association relationship is generated.
[0110] In this step, the feature point set refers to the set of all feature points contained in the indoor fixed structure feature data, each feature point corresponding to a key position on the structure contour; the vertex coordinate set refers to the set of all corner point coordinates of the wall and passage polygon structure extracted from the building model data; and the mapping table refers to a table recording the association relationship between each feature point in the feature point set and the corresponding corner point coordinates in the vertex coordinate set, providing a basis for subsequent coordinate conversion.
[0111] Specifically, first, the feature point set is extracted from the feature data of the indoor fixed structure to determine the image plane coordinates of each feature point; at the same time, the vertex coordinate set is extracted from the coordinate reference information of the wall and the passage to determine the local coordinate system coordinates of each vertex in the building. Second, the RANSAC algorithm is used for point-by-point matching, and the correct matching point pairs are selected by calculating the spatial position similarity between the feature points and the vertices, and the incorrect matching is removed. Finally, the correct matching point pairs are arranged in the format of "feature point image plane coordinates-vertex local coordinate system coordinates" to generate the mapping table.
[0112] For example, in the office building indoor monitoring scene, the feature point set of the east wall of the office area A is extracted from the feature data of the indoor fixed structure, including the image plane coordinates of four feature points: P1(200, 150), P2(300, 150), P3(300, 240), and P4(200, 240); the vertex coordinate set of the wall is extracted from the coordinate reference information, including the local coordinate system coordinates of four vertices: Q1(10, 5, 3), Q2(15, 5, 3), Q3(15, 8, 3), and Q4(10, 8, 3). The RANSAC algorithm is used for matching, the number of iterations is set to 1000, and the inlier threshold is set to 2 pixels. By calculating the spatial position similarity of each feature point and the vertex, the correct matching point pairs P1-Q1, P2-Q2, P3-Q3, and P4-Q4 are obtained. The generated mapping table is shown in Table 2 below:
[0113] Table 2
[0114]
[0115] The above example is only an example of the present application, and other matching algorithms can also be used in actual applications, which are not limited in the present application.
[0116] S1035, based on the mapping table, calculating the conversion parameters from the image plane coordinates of the key frame to the local coordinate system coordinates of the building indoor.
[0117] In the above step, the conversion parameters include the scale adjustment factor and the position translation amount; the scale adjustment factor refers to the ratio of the length of the line segment on the image plane to the actual length of the corresponding line segment in the local coordinate system of the building indoor; the position translation amount refers to the coordinate difference between the geometric center points of the two coordinate systems; the point pair combination refers to the multiple matching point pairs selected from the mapping table.
[0118] In the above step, the conversion parameters include the scale adjustment factor and the position translation amount; the scale adjustment factor refers to the ratio of the length of the line segment on the image plane to the actual length of the corresponding line segment in the local coordinate system of the building indoor; the position translation amount refers to the coordinate difference between the geometric center points of the two coordinate systems; the point pair combination refers to the multiple matching point pairs selected from the mapping table.
[0119] Specifically, firstly, a plurality of point pair combinations are extracted from the mapping table, and the number of selected point pairs is not less than 3 pairs to ensure the calculation accuracy. Secondly, a scale adjustment factor is calculated. For each point pair combination, the length of the line segment between the two points in the image plane coordinate system and the length of the line segment between the corresponding two points in the local coordinate system of the building indoor are calculated respectively, and the ratio of the two is calculated. The average value of all ratios is taken as the final scale adjustment factor. Then, a position translation amount is calculated. The geometric center point coordinates of all image plane coordinates in the selected point pair combination and the geometric center point coordinates of all local coordinate system coordinates of the building indoor are calculated respectively, and the two center point coordinates are subtracted to obtain the position translation amount. Finally, the scale adjustment factor and the position translation amount calculated are combined to form complete conversion parameters.
[0120] For example, in the office building indoor monitoring scene, 3 point pair combinations are extracted from the mapping table generated in S1034: (P1-Q1), (P2-Q2), (P3-Q3). Firstly, the scale adjustment factor is calculated. The length of the line segment of each point pair is calculated according to the distance formula between two points. In the image plane coordinate system, the length of the line segment between P1 (200, 150) and P2 (300, 150) is pixels; the length of the line segment between P2 (300, 150) and P3 (300, 240) is pixels. In the local coordinate system of the building indoor, the length of the line segment between Q1 (10, 5, 3) and Q2 (15, 5, 3) is meters; the length of the line segment between Q2 (15, 5, 3) and Q3 (15, 8, 3) is meters. The length ratio is calculated as follows: meters / pixel, meters / pixel. The average value of and is taken, and the scale adjustment factor is meters / pixel.
[0121] Then, the position translation amount is calculated. The geometric center point coordinates of the image plane coordinates are calculated as follows: , ), , , that is, the center point coordinates are . The geometric center point coordinates of the local coordinate system coordinates of the building indoor are calculated as follows: , , ), , that is, the center point coordinates are (13.33, 6, 3). The position translation amount is , and the numerical calculation is as follows: meters, meters, meters, that is, the position translation amount is . The scale adjustment factor meters / pixel and position translation combining to form a conversion parameter.
[0122] S1036, generating a space constraint signal according to the conversion parameter.
[0123] In this step, the space constraint signal defines the allowable position boundary of the pixel point in the building indoor local coordinate system; the allowable position boundary refers to the reasonable space range in which the pixel point may exist in the building indoor local coordinate system, and the pixel point coordinates beyond this range will be judged as invalid, ensuring the accuracy of coordinate mapping.
[0124] Specifically, first, the conversion model of the image plane coordinates to the building indoor local coordinate system is constructed according to the scale adjustment factor and the position translation in the conversion parameter. Second, the actual area range corresponding to the key frame is analyzed to determine the boundary coordinates of the area in the local coordinate system. Finally, based on the conversion model and the actual area boundary coordinates, the space constraint signal is generated to clearly define the allowable position boundary of each pixel point in the local coordinate system in the key frame, and this signal can be directly used for the validity judgment of the subsequent pixel point coordinates.
[0125] For example, in an office building indoor monitoring scene, based on the conversion parameter obtained in S1035, meters / pixel, meters, meters, meters, the conversion model is constructed: , wherein is the image plane coordinate, is the local coordinate system coordinate. The actual area of the key frame F002 corresponding to the building indoor is the office area and the adjacent corridor entrance, the boundary coordinates of the area in the local coordinate system are meters, meters, meters. Based on the conversion model and the boundary coordinates, the space constraint signal is generated to clearly define the allowable position boundary of the pixel point as meters, meters, meters. For any pixel point in the key frame F002, if the converted coordinate exceeds the boundary range, it is judged as invalid coordinate.
[0126] The present application realizes the effective association from image pixel information to building indoor space information, and constructs a precise coordinate conversion basis. The generated space constraint signal can filter out invalid pixel coordinates in advance, greatly reducing the error of subsequent coordinate inversion calculation, and the integrated fixed structure feature data has good view angle robustness, providing reliable guarantee for the spatial positioning consistency of different key frames, effectively solving the problem of disconnection between traditional image coordinates and actual space coordinates.
[0127] S104, receive the planar coordinates of the key frame and the space constraint signal, and perform inversion calculation on the planar coordinates based on the space constraint signal to obtain coordinate values of the key frame in the local coordinate system of the building indoor.
[0128] Wherein, the inversion calculation refers to the process of converting the image planar coordinates of the key frame to the local coordinate system coordinates of the building indoor based on the inverse operation of the conversion parameter; the coordinate values of the key frame in the local coordinate system of the building indoor refer to the local coordinate system coordinate set corresponding to all effective pixel points in the key frame, which can reflect the position of the key frame scene in the actual building space.
[0129] Optionally, step S104 can specifically include the following steps:
[0130] S1041, receive the planar coordinates of the key frame composed of pixel row numbers and column numbers, and obtain the conversion parameter defined in the space constraint signal.
[0131] In this step, the planar coordinates composed of pixel row numbers and column numbers refer to the positioning coordinates of each pixel in the key frame image, and the row number corresponds to the vertical position of the image, and the column number corresponds to the horizontal position of the image; the obtained conversion parameter is the scale adjustment factor and the position translation amount carried in the space constraint signal, which is the core basis for realizing coordinate inversion calculation.
[0132] Specifically, first, through the output interface of the image acquisition device, the row number and the column number of each pixel point in the key frame image are received, which are converted into standard planar coordinates format, wherein corresponding to the column number, corresponding to the row number. Secondly, the space constraint signal is received through the signal transmission interface, the conversion parameter used for coordinate inversion calculation is parsed from the signal, including the scale adjustment factor and the position translation amount, and the integrity and effectiveness of the parameter are checked to ensure that the parameter can be directly used for subsequent calculation.
[0133] For example, in the office building indoor monitoring scene, the pixel planar coordinates of the key frame F002 are received, and three pixel points are selected as an example, and the planar coordinates converted from the row numbers and the column numbers are respectively: pixel point A (220, 160), pixel point B (280, 200), and pixel point C (350, 250). At the same time, the space constraint signal is received, and the conversion parameter is parsed from the signal, which is the scale adjustment factor ≈0.0417 meters / pixel, and the position translation amount =(2.22, -1.51, 3). The parameter is checked to confirm 、 、 、 All are effective values, no missing or error, can be used for subsequent inversion.
[0134] S1042、According to the conversion parameters, a coordinate inverse transformation formula is constructed, and the plane coordinates are substituted into the coordinate inverse transformation formula for calculation to obtain three-dimensional position data corresponding to each pixel point.
[0135] In this step, the coordinate inverse transformation formula refers to a mathematical formula constructed based on the conversion parameters for converting the image plane coordinates into the three-dimensional coordinates of the local coordinate system of the building indoor; the three-dimensional position data refers to the x, y, z coordinate values of each pixel point in the local coordinate system, which can accurately reflect the actual spatial position corresponding to the pixel point.
[0136] Specifically, first, according to the scale adjustment factor and the position translation amount in the conversion parameters, a coordinate inverse transformation formula is constructed. Since the conversion model is , the inverse transformation formula is , and since the conversion is a linear transformation, the inverse transformation is consistent with the positive transformation form, only the direction is opposite. Secondly, the plane coordinates of each pixel point are substituted into the inverse transformation formula to calculate the coordinate values respectively, and the three-dimensional position data of each pixel point is obtained. Finally, in combination with the allowable position boundary in the space constraint signal, the validity of the calculated three-dimensional position data is determined, and the invalid data beyond the boundary is removed.
[0137] For example, in the office building indoor monitoring scene, based on the conversion parameters obtained by analysis, the coordinate inverse transformation formula , , is constructed. Wherein, is the three-dimensional coordinates of the local coordinate system of the building indoor, is the image plane coordinates, is the scale adjustment factor, is the position translation amount.
[0138] Substitute the plane coordinates of the three pixel points in S1041 into the formula for calculation:
[0139] Pixel point A (100, 100): meters, meters, meters, and the three-dimensional position data is ;
[0140] Pixel point B (280, 200): meters, meters, meters, and the three-dimensional position data is ;
[0141] Pixel point C meters, meters, meters, the three-dimensional position data of the pixel point A and B is within the boundary, which is valid data; the three-dimensional position data of the pixel point C is .
[0142] Combining the allowable position boundary in the spatial constraint signal meters, meters, meters) is invalid data. meters, exceeding the maximum boundary of 16 meters, meters, exceeding the maximum boundary of 8 meters, which is invalid data.
[0143] S1043, aggregate the three-dimensional position data of all the pixel points to form the coordinate value of the key frame under the local coordinate system of the building indoor.
[0144] In this step, aggregation refers to the process of arranging and collecting the three-dimensional position data of all valid pixel points in the order of rows and columns of the pixel points in the key frame; the formed coordinate value is the complete spatial representation of the key frame scene in the local coordinate system of the building indoor, which can fully reflect the actual spatial position and scene form corresponding to the key frame.
[0145] Specifically, first, collect the three-dimensional position data of all valid pixel points obtained in S1042. Second, arrange the valid three-dimensional position data in order according to the row and column order of the pixel points in the key frame image, ensuring that the data order is consistent with the position order of the pixel points in the image. Finally, integrate the ordered valid three-dimensional position data to form a complete coordinate value set of the key frame under the local coordinate system of the building indoor.
[0146] For example, in an office building indoor monitoring scene, after the three-dimensional position data of all pixel points of the key frame F002 is calculated and validity determined, 1920x1080 valid pixel points are collected. According to the order of column number from left to right and row number from top to bottom, the three-dimensional position data is arranged in order. The arranged three-dimensional position data is integrated into a 1920x1080x3 array, the third dimension being x, y, z coordinates, which is the coordinate value of the key frame F002 under the local coordinate system of the building indoor. Through the coordinate value set, the actual spatial scene of the office area A and the adjacent corridor entrance corresponding to the key frame F002 can be clearly restored.
[0147] The present application solves the problems of inaccurate mapping of image coordinates and actual building space coordinates, and large invalid coordinate interference in the traditional method. Compared with the traditional method, the present application accurately extracts the features of fixed structures and establishes a correlation mapping, and realizes effective inversion of coordinates combined with spatial constraints, thereby improving the accuracy and reliability of coordinate conversion, and ensuring the comprehensiveness of spatial representation by integrating multiple key frame features.
[0148] S105, obtaining a geographic coordinate reference signal from the building model data, and converting the coordinate value in the building indoor local coordinate system into the latitude and longitude coordinate value of the digital twin model corresponding to the target building in combination with the geographic coordinate reference signal.
[0149] In this step, the building model data refers to a structured data set containing the geometric shape, spatial layout, attribute information and geographic location correlation data of the target building, such as IFC format data of BIM model; the geographic coordinate reference signal refers to a reference signal carrying the latitude and longitude information of the geographic reference point, used to establish the correlation between the local coordinate system and the geographic coordinate system; the latitude and longitude coordinate value of the digital twin model refers to the coordinate data after mapping the building indoor space position to the geographic coordinate system on the earth's surface, which is the core data for realizing the accurate alignment of the digital twin model and the real geographic space.
[0150] Optionally, step S105 can specifically include the following steps:
[0151] S1051, reading the latitude and longitude values of the pre-defined geographic reference points from the structured attribute field of the building model data, and generating a reference signal containing latitude and longitude data.
[0152] The structured attribute field is a specific attribute storage field defined according to the specification in the building model data. The pre-defined geographic reference point corresponds to the actual geographic location of the target building, and has both indoor local coordinates and real geographic latitude and longitude. It is usually selected from fixed and easily positioned positions such as building corners and entrances.
[0153] In a specific embodiment, first, open the building model data file through a data analysis tool, and locate the structured attribute field storing the geographic reference point information. Then, read the latitude and longitude values of the pre-defined geographic reference points from the field, and check the indoor local coordinates corresponding to each reference point to ensure uniqueness and correlation. Finally, organize the reference point identification, latitude and longitude values, and corresponding local coordinates in a preset format to generate a geographic coordinate reference signal.
[0154] For example, in the office building digital twin model construction project, the staff uses the FME data parsing tool to open the BIM model IFC file of the building, and locates the RefLatitude and RefLongitude fields of the IfcSite entity. From it, the information of 3 pre-defined geographic reference points is read, which are the main entrance outer wall corner RP1, the northwest outer wall corner RP2, and the southeast outer wall corner RP3. The longitude of RP1 is 114.03125°, the latitude is 22.53430°, and the corresponding indoor local coordinates are 10, 5, 3; the longitude of RP2 is 114.03088°, the latitude is 22.53456°, and the corresponding indoor local coordinates are 10, 30, 3; the longitude of RP3 is 114.03162°, the latitude is 22.53404°, and the corresponding indoor local coordinates are 45, 5, 3. These information is arranged in the format of "reference point ID-longitude-latitude-local coordinates" to generate a standardized geographic coordinate reference signal.
[0155] S1052, based on the reference signal, calculating the coordinate conversion parameter between the indoor local coordinate system of the building and the geographic coordinate system.
[0156] Wherein, the offset is the coordinate difference value of the origin of the indoor local coordinate system and the origin of the geographic coordinate system, including the longitude offset and the latitude offset, used to correct the difference of the positions of the origins of the two coordinate systems. The direction angle is the included angle between the X axis of the local coordinate system and the true north direction of the geographic coordinate system, used to correct the difference of the directions of the coordinate axes of the two coordinate systems. The coordinate conversion parameter includes the offset and the direction angle of the origin of the local coordinate system in the geographic coordinate system.
[0157] In the embodiments of the present application, at least two geographic reference points are extracted from the reference signal. The longitude and latitude and local coordinate corresponding data pairs of the reference points are extracted. Based on these data pairs, the longitude and latitude difference and the local coordinate difference are calculated, and the offset is derived through linear relationship. Then the included angle between the connection direction of the reference points in the local coordinate system and the geographic true north direction is calculated through the vector included angle formula to obtain the direction angle, and finally the coordinate conversion parameter is integrated.
[0158] For example, in the office building digital twin model construction scenario, based on the geographic coordinate reference signal generated by S1051, the data pairs of RP1 main entrance outer wall corner and RP2 northwest outer wall corner are selected for coordinate conversion parameter calculation. After sorting the data, it is known that the longitude and latitude of RP1 are , , and the corresponding indoor local coordinates are ; the longitude and latitude of RP2 are , and the corresponding indoor local coordinates are .
[0159] Assuming that the longitude and latitude and indoor local coordinates are in a linear relationship, the conversion relationship formula is Lon = , Lat , wherein Lon is longitude, Lat is latitude, is the office building indoor local coordinate, is the scale factor, is the offset. Since the x coordinates of RP1 and RP2 are the same, the scale factors b and d can be simplified, degrees per unit length, degrees per unit length.
[0160] The vector of the reference points RP1 to RP2 in the office building indoor local coordinate system is , along the positive direction of the axis points to the depth of the office area, and the geographic true north direction is the increasing direction of latitude. The direction angle is calculated by the vector angle. Finally, the coordinate conversion parameters suitable for the office building are the offset and the direction angle . The above example is only one example of the present application, and more reference points can be selected to improve the calculation accuracy according to the requirements in actual application, which is not limited by the present application.
[0161] S1053, using the coordinate conversion parameters, performing point-by-point coordinate transformation operation on the coordinate values in the building indoor local coordinate system to generate the latitude and longitude coordinate values of the digital twin model.
[0162] wherein the point-by-point coordinate transformation operation is a calculation process of applying the conversion parameters to the coordinate mapping of each pixel point or spatial point coordinate in the indoor local coordinate system one by one.
[0163] Finally, the latitude and longitude coordinate values are generated by S1053. First, the offset and the direction angle are combined to construct a conversion model of the local coordinate to the latitude and longitude coordinate, and then all the coordinate values in the indoor local coordinate system are extracted from the results of S104 and arranged in order. Each coordinate value is substituted into the conversion model to complete the direction correction and position offset operation in turn to obtain the corresponding latitude and longitude coordinate. After the validity check eliminates the abnormal values, the latitude and longitude coordinate values required by the digital twin model are generated.
[0164] For example, first convert the local coordinate to the north-east coordinate by the rotation formula, and the rotation formula is shown in equations (3) and (4):
[0165] (3)
[0166] (4)
[0167] wherein, is the north coordinate component, is the east coordinate component, is the direction angle. The longitude and latitude conversion formula is constructed by combining the offset and the scale factor, respectively, , wherein, is the offset, degree per unit length, and degree per unit length are the scale factors calculated before. The effective pixel point local coordinates obtained in S104 are converted, and ; are calculated by substituting the rotation formula. The longitude and latitude formula is substituted by ; and .
[0168] After all the effective points are converted by this method, the longitude and latitude coordinate values of the digital twin model are generated. The above example is only one example of the present application, and batch conversion can also be realized by programming according to the needs in actual application, which is not limited by the present application.
[0169] The present application solves the practical problem that the digital twin model cannot accurately match the real geographical position caused by the disconnection between the traditional indoor coordinate and the geographical coordinate by extracting geographical reference information from the building model data, accurately calculating the coordinate conversion parameters, and completing the point-by-point coordinate transformation. Compared with the traditional conversion scheme, it relies on the inherent properties of the building model to build the reference, reduces the additional geographical measurement cost, and at the same time improves the conversion accuracy through the cooperative calculation of multiple reference points, realizes the efficient and accurate association of indoor space and geographical space.
[0170] Figure 3 is a structural schematic diagram of one specific embodiment of a video frame plane coordinate and model longitude and latitude coordinate mapping system provided by an embodiment of the present application, referring to Figure 3 , the system can include:
[0171] An acquisition module 31 is configured to acquire multi-view video frames in a target building indoor space and building model data corresponding to the target building, wherein the building model data includes indoor space topology data and floor layout data.
[0172] An extraction module 32 is configured to extract key frames from the multi-view video frames and extract coordinate reference information of walls and passages from the building model data to form a structured data set.
[0173] An association module 33 is configured to process the key frames in the data set to generate feature data of indoor fixed structures, and associate the feature data of the indoor fixed structures with the coordinate reference information of the walls and passages to generate a spatial constraint signal for defining the spatial position of the pixel points in the key frames.
[0174] an inversion module 34, configured to receive the planar coordinates of the key frame and the space constraint signal, and perform inversion calculation on the planar coordinates based on the space constraint signal to obtain coordinate values of the key frame in a local coordinate system of the building indoor space;
[0175] a conversion module 35, configured to obtain a geographic coordinate reference signal from the building model data, and convert the coordinate values in the local coordinate system of the building indoor space into latitude and longitude coordinate values of a digital twin model corresponding to the target building in combination with the geographic coordinate reference signal.
[0176] The video frame planar coordinate and model latitude and longitude coordinate mapping system according to the embodiments of the present application is used to implement the video frame planar coordinate and model latitude and longitude coordinate mapping method described above, and therefore the specific implementation of the video frame planar coordinate and model latitude and longitude coordinate mapping system can be seen from the implementation of the video frame planar coordinate and model latitude and longitude coordinate mapping method described above, and the specific implementation can be referred to the description of the corresponding embodiment, which will not be repeated here.
[0177] The present application also provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to execute the computer program to implement the steps of any of the video frame planar coordinate and model latitude and longitude coordinate mapping methods described above.
[0178] The present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of any of the video frame planar coordinate and model latitude and longitude coordinate mapping methods described above.
[0179] In an exemplary embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0180] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of any of the video frame planar coordinate and model latitude and longitude coordinate mapping methods described above.
[0181] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of embodiments of the present application and are not intended to limit the scope of the present application. Accordingly, embodiments as described herein contemplate all modifications that come within the scope of the present application.
[0182] The video frame plane coordinate and model latitude and longitude coordinate mapping method and system provided by the present application are described in detail above. The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A method for mapping video frame planar coordinates to model latitude and longitude coordinates, characterized in that, include: Acquire multi-view video frames of the interior of the target building, as well as the building model data corresponding to the target building, wherein the building model data includes interior space topology data and floor layout data; Keyframes are extracted from the multi-view video frames, and coordinate reference information of walls and passages is extracted from the building model data to form a structured data set. The keyframes within the dataset are processed to generate feature data of the indoor fixed structure, and the feature data of the indoor fixed structure is associated with the coordinate reference information of the wall and the passage to generate spatial constraint signals for defining the spatial position of pixels in the keyframes. Receive the planar coordinates of the keyframe and the spatial constraint signal, and perform inverse calculation on the planar coordinates based on the spatial constraint signal to obtain the coordinate values of the keyframe in the local coordinate system of the building interior; Geographic coordinate reference signals are obtained from the building model data, and combined with the geographic coordinate reference signals, the coordinate values in the local coordinate system of the building interior are converted into the latitude and longitude coordinate values of the digital twin model corresponding to the target building.
2. The method according to claim 1, characterized in that, Associating the feature data of the indoor fixed structure with the coordinate reference information of the wall and passageway to generate spatial constraint signals for defining the spatial position of pixels in the keyframe, including: The feature point set in the feature data of the indoor fixed structure is matched point by point with the vertex coordinate set in the coordinate reference information of the wall and passage to generate a mapping table describing the point-to-point relationship. Based on the mapping table, the transformation parameters from the image planar coordinates of the keyframe to the coordinates of the building's interior local coordinate system are calculated. The transformation parameters include a scale adjustment factor and a position translation amount. Based on the transformation parameters, a spatial constraint signal is generated, which defines the permissible position boundary of the pixel in the local coordinate system of the building interior.
3. The method according to claim 2, characterized in that, Based on the mapping table, the transformation parameters from the image planar coordinates of the keyframe to the coordinates of the building's interior local coordinate system are calculated, including: Based on the mapping table, multiple point pair combinations are extracted, each point pair combination containing the image plane coordinates and the corresponding building interior local coordinate system coordinates; Based on the coordinate values in the point pair combination, the length ratio of the line segment between each pair of corresponding points in the image plane coordinate system and the building interior local coordinate system is calculated to generate a scale adjustment factor. Based on the coordinate values in the selected point pair combination, calculate the geometric center point coordinates of the image plane coordinates and the geometric center point coordinates of the building interior local coordinate system, respectively, to generate the position translation amount; The scale adjustment factor and the position translation amount are combined to form the transformation parameters used for coordinate transformation.
4. The method according to claim 1, characterized in that, Processing the keyframes within the dataset to generate feature data for indoor fixed structures includes: By analyzing the image pixel array of each keyframe, the contour boundary of the indoor fixed structure is identified, and the feature point coordinate sequence of the contour boundary is extracted. Based on the feature point coordinate sequence, a feature descriptor describing the shape of the structure is calculated, and the feature descriptor contains the relative positional relationships of the point set; The feature descriptors of all keyframes are integrated to generate feature data of the indoor fixed structure, which is used to represent the spatial characteristics of the fixed structure.
5. The method according to claim 1, characterized in that, Receiving the planar coordinates of the keyframe and the spatial constraint signal, and performing inverse calculation on the planar coordinates based on the spatial constraint signal to obtain the coordinate values of the keyframe in the local coordinate system of the building interior, including: Receive the planar coordinates formed by the row and column numbers of the pixels in the key frame, and simultaneously obtain the transformation parameters defined in the spatial constraint signal; Based on the transformation parameters, an inverse coordinate transformation formula is constructed. The planar coordinates are substituted into the inverse coordinate transformation formula for calculation to obtain the three-dimensional position data corresponding to each pixel. The three-dimensional position data of all pixels are aggregated to form the coordinate values of the keyframe in the local coordinate system of the building interior.
6. The method according to claim 1, characterized in that, Obtain geographic coordinate reference signals from the building model data, and combine them with the geographic coordinate reference signals to convert the coordinate values in the building's indoor local coordinate system into the latitude and longitude coordinate values of the digital twin model corresponding to the target building, including: Read the latitude and longitude values of a predefined geographic reference point from the structured attribute fields of the building model data. The geographic reference point corresponds to the actual geographical location of the target building, and generate a reference signal containing latitude and longitude data. Based on the reference signal, the coordinate transformation parameters between the building's indoor local coordinate system and the geographic coordinate system are calculated. The coordinate transformation parameters include the offset of the origin of the local coordinate system and the orientation angle in the geographic coordinate system. Using the coordinate transformation parameters, point-by-point coordinate transformation operations are performed on the coordinate values in the local coordinate system of the building interior to generate the latitude and longitude coordinate values of the digital twin model.
7. The method according to claim 1, characterized in that, Keyframes are extracted from the multi-view video frames, and coordinate reference information for walls and passageways is extracted from the building model data to form a structured dataset, including: Based on the sequence of the multi-view video frames, the matching degree of feature points between adjacent video frames is calculated to generate feature difference values; When the feature difference value is greater than a preset threshold, the current video frame is marked as a key frame and the key frame set is stored. From the interior space topology data of the building model data, the polygonal structure of the walls and passages is parsed, and the corner coordinates of the polygonal structure are extracted to form coordinate reference information; The keyframe set is associated and combined with the coordinate reference information to generate a structured data set, wherein each keyframe is mapped to a corresponding subset of coordinate reference information.
8. A mapping system between video frame planar coordinates and model latitude and longitude coordinates, characterized in that, include: The acquisition module is used to acquire multi-view video frames of the interior of the target building, as well as the building model data corresponding to the target building. The building model data includes interior space topology data and floor layout data. The extraction module is used to extract keyframes from the multi-view video frames and extract coordinate reference information of walls and passages from the building model data to form a structured data set. The association module is used to process the keyframes in the data set to generate feature data of the indoor fixed structure, and associate the feature data of the indoor fixed structure with the coordinate reference information of the wall and the passage to generate a spatial constraint signal for limiting the spatial position of the pixel in the keyframe. The inversion module is used to receive the planar coordinates of the key frame and the spatial constraint signal, and perform inversion calculation on the planar coordinates based on the spatial constraint signal to obtain the coordinate values of the key frame in the local coordinate system of the building interior. The conversion module is used to obtain the geographic coordinate reference signal from the building model data, and combine the geographic coordinate reference signal to convert the coordinate values in the local coordinate system of the building interior into the latitude and longitude coordinate values of the digital twin model corresponding to the target building.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the method for mapping video frame planar coordinates to model latitude and longitude coordinates as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the mapping method between the planar coordinates of a video frame and the latitude and longitude coordinates of a model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual and real entity coordinate mapping method and system for building digital twinning
CN114329747A
Video stream and BIM model virtual-real mapping calibration method and device
CN119068064A