A highway multi-modal data and aerial video full-automatic fusion method and system

CN117407957BActive Publication Date: 2026-08-07SHANDONG TRAFFIC PLANNING DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG TRAFFIC PLANNING DESIGN INST
Filing Date
2023-10-17
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]然而,单一的航拍视频难以表达具体的设计意图;其次,多模态设计数据与航拍视频存在数据形态壁垒,两者难以融合应用,导致方案解释不直观,从而影响展示效果

Benefits of technology

[0038]1、本发明提出一种统筹多模态数据几何语义、属性语义与显示语义的语义信息模型,可实现多模态设计数据的语义索引机制管理,打破多模态数据与航拍视频的数据隔阂,为全自动融合提供技术保障。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117407957B_ABST
    Figure CN117407957B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of highway survey and design, and provides a highway multi-modal data and aerial video full-automatic fusion method and system, the technical scheme of which is: highway design multi-modal data is fused into unmanned aerial vehicle aerial video in the form of static or dynamic, realizing full-range and full-design-element expression of design scheme, breaking the data form barrier of the two by creating multi-modal design data semantic information model and video frame sequence three-dimensional virtual coordinate system, realizing full-automatic fusion based on semantic indexing mechanism, obtaining high-definition demonstration video expressing design intention, enriching new digital look-over line mode, and the achievement brings strong sense of presence, providing scientific decision basis for highway design review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of highway survey and design, and in particular relates to a fully automatic fusion method and system for highway multimodal data and aerial video. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In the review of highway survey and design schemes at each stage, the correctness, rationality, and standardization of the design schemes are the issues of greatest concern to review experts and owners. Highway scheme design is characterized by its high degree of specialization and interdisciplinary nature, and the design process generates a massive amount of multimodal design data. When presenting and reviewing projects, simply using slides or other presentation methods can lead to problems such as vague expression of contradictions, discontinuous scheme demonstrations, and inaccurate qualitative and quantitative analysis.

[0004] In recent years, UAV aerial surveying technology has been increasingly applied across various industries due to its advantages such as relatively low cost, high data acquisition efficiency, high data resolution, good maneuverability, and high safety performance. Currently, in the field of highway surveying and design, UAVs are mainly used to collect aerial videos of the entire design plan, thereby enabling a clear and intuitive understanding and display of the project site conditions, reducing the workload of experts in on-site surveys.

[0005] However, a single aerial video is insufficient to express specific design intent; secondly, there are data format barriers between multimodal design data and aerial videos, making it difficult to integrate and apply the two, resulting in an unintuitive explanation of the solution and thus affecting the presentation effect. Summary of the Invention

[0006] To address at least one of the technical problems mentioned above, this invention provides a fully automated fusion method and system for highway multimodal data and aerial video. This method integrates highway design multimodal data into UAV aerial video in a static or dynamic form, enabling a comprehensive and complete representation of the design scheme. This provides review experts and owners with a strong sense of presence, reduces the workload of on-site surveys, and shortens the review cycle.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The first aspect of this invention provides a fully automated fusion method for highway multimodal data and aerial video, comprising the following steps:

[0009] Acquire aerial video footage and perform accelerated frame extraction to obtain a video frame sequence;

[0010] Acquire multimodal data of highways and construct a semantic information model based on the multimodal data of highways;

[0011] By using the image point coordinates of video frame sequences, the three-dimensional coordinates of real ground points, and the perspective geometry of the camera's optical center, a flight path model corresponding to reality is established, resulting in a three-dimensional virtual coordinate system.

[0012] Highway multimodal data is loaded into a 3D virtual coordinate system. Based on the indexing mechanism of the semantic information model, the highway multimodal data is semantically parsed for attribute information, geometric information and display information. Based on the semantic information obtained from the parsing, the data is automatically located, drawn and displayed in the 3D virtual coordinate system, completing the initial fusion of highway multimodal data and video frame sequence.

[0013] Preview the initial fusion effect of highway multimodal data and video frame sequence, optimize the display position, scale and angle of highway multimodal data in video frame sequence, and output the fused video data.

[0014] Furthermore, the highway multimodal data includes labeled file data, vector line data, and BIM model data.

[0015] Furthermore, the construction of its semantic information model based on highway multimodal data includes:

[0016] Create a semantic information model object for highway multimodal data;

[0017] Define the structure, attributes, and meaning of multimodal highway data semantics;

[0018] Establish a mapping relationship between semantic information model objects and data semantics, that is, realize a one-to-one correspondence between semantic information models and highway multimodal data, thereby completing the construction of semantic information models.

[0019] Furthermore, by utilizing the image point coordinates of the video frame sequence, the three-dimensional coordinates of real ground points, and the perspective geometry of the camera's optical center, a flight path model corresponding to the actual situation is established, resulting in a three-dimensional virtual coordinate system, including:

[0020] By using the 3D coordinates of real ground points and the coordinates of image points in a video frame sequence to perform image point motion analysis, the camera exposure position is calculated, the flight path model equation is obtained, and then the ground 3D coordinates of any image point in the video frame can be calculated, thus creating a 3D virtual world coordinate system.

[0021] Furthermore, the process of loading highway multimodal data into a three-dimensional virtual coordinate system, performing attribute, geometric, and display semantic analysis on the highway multimodal data based on an indexing mechanism using a semantic information model, and automatically locating, drawing, and displaying the data in the three-dimensional virtual coordinate system based on the obtained semantic information includes:

[0022] Load all labeled spline animations and parse the semantic information one by one. Obtain the unique identifier of the attribute information to determine the positioning order. Determine the positioning coordinates by the attachment point position of the geometric information. Determine the display effect by the angle and scale of the display information. Finally, automatically locate the attachment point based on the above index information.

[0023] Load the attribute, geometric, and display information of the horizontal curve table file, automatically draw splines representing straight lines, transition curves, and circular curves that characterize the road direction, and obtain parallel markings based on spline mapping;

[0024] The unique identifier of the BIM model semantic model, the coordinates of the vertices of the minimum bounding rectangle, and the display parameters are parsed, and the model is automatically positioned in the corresponding location and the orientation is displayed correctly.

[0025] Furthermore, the display position, proportion, and angle of the multimodal data of the highway in the video frame sequence are optimized, and shots that do not meet the requirements are adjusted. After adjustment, the rendered and outputted lossless demonstration video is saved.

[0026] Furthermore, the attribute information includes the unique identifier, name, and file path of the spline animation, the horizontal curve polyline, and the model; the geometric information includes the spline animation attachment point position, the horizontal curve 3D point data, and the vertex coordinates of the model's minimum bounding rectangle; the display information includes the angle, scale, and color of the data displayed in the video frame sequence.

[0027] A second aspect of the present invention provides a fully automated fusion system for highway multimodal data and aerial video, comprising:

[0028] The video frame acquisition module is used to acquire aerial video and perform accelerated frame extraction processing to obtain a video frame sequence.

[0029] The semantic information model building module is used to acquire multimodal highway data and build its semantic information model based on the multimodal highway data.

[0030] The 3D virtual coordinate system establishment module is used to establish a flight path model that corresponds to the actual situation by using the coordinates of image points in the video frame sequence, the 3D coordinates of real ground points and the perspective geometry of the camera optical center, and thus obtain the 3D virtual coordinate system.

[0031] The automatic fusion module is used to load highway multimodal data into a three-dimensional virtual coordinate system. Based on the indexing mechanism of the semantic information model, it performs semantic parsing of the attribute information, geometric information and display information of the highway multimodal data, and automatically locates, draws and displays it in the three-dimensional virtual coordinate system based on the obtained semantic information, thus completing the initial fusion of highway multimodal data and video frame sequences.

[0032] The fusion correction module is used to preview the initial fusion effect of highway multimodal data and video frame sequences, optimize the display position, scale and angle of highway multimodal data in video frame sequences, and output the fused video data.

[0033] A third aspect of the present invention provides a computer-readable storage medium.

[0034] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the fully automated fusion method of multimodal highway data and aerial video as described above.

[0035] A fourth aspect of the present invention provides a computer device.

[0036] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the fully automated fusion method of multimodal highway data and aerial video as described above.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] 1. This invention proposes a semantic information model that integrates the geometric semantics, attribute semantics, and display semantics of multimodal data. It can realize the semantic indexing mechanism management of multimodal design data, break down the data barriers between multimodal data and aerial video, and provide technical support for fully automatic fusion.

[0039] 2. This invention performs image point motion analysis based on the principle of photogrammetry, constructs a UAV flight path model, and establishes a three-dimensional virtual coordinate system consistent with the real-world coordinates to ensure that the data is consistent with the camera motion.

[0040] 3. This invention inputs multimodal design data into a three-dimensional virtual coordinate system, analyzes its attributes, geometry and display semantics, and realizes fully automatic object attachment, drawing and positioning.

[0041] 4. Integrating multimodal design data of highways into aerial videos can more realistically reflect the virtual construction effects before, during and after construction. This can be applied to the survey, reporting and review of highway design, providing a more scientific basis for decision-making and has broad application prospects.

[0042] 5. Utilizing UAV aerial surveying technology can significantly reduce the on-site investigation by experts and the fieldwork of surveyors, reduce the intensity of fieldwork, and improve the safety factor of on-site measurements in dangerous areas. It has the advantages of operational safety and saving manpower and time costs.

[0043] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0044] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0045] Figure 1 This is a block diagram of the fully automatic fusion method of highway multimodal data and aerial video provided in the embodiments of the present invention. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0048] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] To address the technical problems mentioned in the background section, this invention proposes a fully automated fusion method and system for highway multimodal design data and UAV aerial video. The system includes aerial video acquisition, highway multimodal design data collection and processing, establishment of a three-dimensional virtual coordinate system, fully automated fusion, and optimization and output of the demonstration video. The aim is to break down the data format barriers between the two by creating a semantic information model of the multimodal design data and a three-dimensional virtual coordinate system for the video frame sequence, achieving fully automated fusion based on a semantic indexing mechanism. This results in a high-definition demonstration video that expresses the design intent, enriching new methods of digital road alignment review. The outcome provides a strong sense of presence and offers a scientific basis for highway design review.

[0050] Example 1

[0051] like Figure 1 As shown in the figure, this embodiment provides a fully automatic fusion method for highway multimodal data and aerial video, including the following steps:

[0052] Step 1: Acquire aerial video and perform accelerated frame extraction to obtain a video frame sequence;

[0053] Step 2: Collect multimodal design data for highways and construct a semantic information model based on the multimodal design data;

[0054] Step 3: Using the image point coordinates of the video frame sequence, the three-dimensional coordinates of the real ground points, and the perspective geometric relationship of the camera optical center "object-image-camera", establish a flight path model that corresponds to the actual situation, obtain a three-dimensional virtual coordinate system, and then create a three-dimensional virtual world scene;

[0055] Step 4: Load highway multimodal data in batches into the 3D virtual coordinate system, perform semantic parsing of the highway multimodal data for attribute information, geometric information and display information, and realize the positioning, drawing and display of highway multimodal data in the 3D virtual coordinate system of video frame sequence based on the semantic indexing mechanism, that is, realize the fully automatic preliminary fusion of the two.

[0056] Step 5: Preview the initial fusion effect of the highway multimodal data and video frame sequence, optimize its display parameters such as position, scale, angle and color, and output the video data after confirming that there are no errors.

[0057] In step 1, the process of acquiring aerial video is as follows: the flight path is automatically divided based on the KML data of the highway design centerline, and the UAV flies autonomously along the flight path at the same altitude and speed; GPS technology and coordinated turning mode ensure the UAV's aerial attitude, thereby acquiring high-definition, shake-free, and continuous video footage along the route.

[0058] In step 2, the highway multimodal data includes annotation file data, vector line drawing datasets, and BIM model data;

[0059] The labeled file data includes text spline animation data, image spline animation data, and short video spline animation data;

[0060] Vector line drawing datasets include horizontal curve data for highway design centerlines. They are generally in tabular form and record parameters such as the three-dimensional point coordinates, front and rear transition lengths, and radii of straight lines, transition curves, and circular curves. They come with built-in attribute and geometric information, and spline width and color are added as display information for horizontal curves.

[0061] BIM model data includes road models, bridge models, and interchange models, which are textured, realistic 3D models generated using self-developed parametric modeling software.

[0062] The attribute information includes the unique identifier, name, and file path of the spline animation, horizontal curve polyline, and model; the geometric information includes the spline animation attachment point position, horizontal curve 3D point data, and the vertex coordinates of the model's minimum bounding rectangle; and the display information includes the angle, scale, and color of the data displayed in the video frame sequence.

[0063] In this embodiment, constructing the semantic information model based on highway multimodal data includes:

[0064] First, a semantic information model object is created. Then, the semantics of highway multimodal data are defined, including the structure, attributes, and meaning of the data. The aforementioned attributes, geometry, and display information are organized according to this definition. Finally, the mapping relationship between the semantic information model and the data semantics is established, that is, the correspondence between the semantic information model and the highway multimodal data is realized.

[0065] The above technical solution enables the management of semantic indexing mechanism for multimodal design data, breaks down the data barrier between multimodal data and aerial video, and facilitates rapid indexing for subsequent data fusion applications.

[0066] In step 3, when acquiring aerial video, the coordinates of real ground points with significant features that are easily identifiable in video frames are collected along the route. Image point coordinates corresponding to the real ground points are selected in the video frame sequence. Image point motion analysis is performed using the three-dimensional coordinates of the real ground points and the image point coordinates of the video frame sequence to calculate the camera exposure position and obtain the flight path model equation. Then, the ground three-dimensional coordinates of any image point in the video frame can be calculated, thus creating a three-dimensional virtual world coordinate system.

[0067] Specifically, the step of using the three-dimensional coordinates of real ground points and the image point coordinates of video frame sequences to perform image point motion analysis and then solve the flight path model equations involves:

[0068] First, select at least 6 pairs of one-to-one corresponding 3D coordinates of ground points and image point coordinates of video frame sequences. Then, use the direct linear transformation algorithm to obtain the camera's rotation and translation matrices, realizing the coordinate transformation between the world coordinate system, camera coordinate system, and pixel coordinate system. This allows us to obtain the exposure position and attitude information of the UAV camera, thus solving for the flight path model.

[0069] In step 4, highway multimodal data is loaded into the 3D virtual coordinate system. Based on the indexing mechanism of the semantic information model, attribute, geometric, and display semantic parsing is performed on the highway multimodal data. Based on the obtained semantic information, the data is automatically located, drawn, and displayed in the 3D virtual coordinate system, completing the initial fusion of highway multimodal data and video frame sequences, including:

[0070] All annotation spline animations are stored in the same folder. When you select this folder in the system, the system will automatically import all spline animations and parse the semantic information one by one. It will obtain the unique identifier of the attribute information to determine the positioning order, determine the positioning coordinates by the location of the attachment point of the geometric information, and determine the display effect in the system by the angle and scale of the display information. Finally, it will automatically locate the attachment point based on the above index information.

[0071] Import the horizontal alignment data of the highway design centerline into the system, obtain the attribute information, geometric information and display information of the horizontal curve table file, automatically draw the spline lines representing the straight line segments, transition curve segments and circular curve segments of the highway, and obtain the parallel markings based on the spline line mapping.

[0072] Considering the robustness issues arising from the massive data volume of highway 3D BIM models, the data is divided into three folders: roads, bridges, and interchanges. Within the system, each of these folders is selected, and its unique semantic model identifier, minimum bounding rectangle vertex coordinates, and display parameters are parsed. The model automatically and accurately locates itself in the corresponding position and displays the correct orientation. The system allows users to view and edit annotation spline animations, design centerline curves, and the semantic information of the 3D BIM model.

[0073] Preview the initial fusion effect of highway multimodal data and video frame sequence. Use two methods, coarse adjustment and semantic fine adjustment, to optimize the display position, scale and angle of highway multimodal data in video frame sequence. After adjustment, save the rendered output lossless demonstration video.

[0074] Coarse adjustment refers to manual interactive adjustment, while fine adjustment involves modifying attributes, geometry, and display parameters.

[0075] Through the aforementioned technical solutions, aerial video can present a high-definition, realistic on-site environment. Computer graphics, using mathematical algorithms, provide representation, interaction, and rendering of two-dimensional and three-dimensional highway design graphics, creating a variety of eye-catching dynamic graphics and stunning visual effects. Integrating multimodal highway design data into UAV aerial video in static or dynamic form enables a comprehensive and complete expression of the design scheme, providing review experts and owners with a strong sense of presence, reducing the workload of on-site surveys, and thus shortening the review cycle.

[0076] Example 2

[0077] This embodiment provides a fully automated fusion system for highway multimodal data and aerial video, including:

[0078] The video frame acquisition module is used to acquire aerial video and perform accelerated frame extraction processing to obtain a video frame sequence.

[0079] The semantic information model building module is used to acquire multimodal highway data and build its semantic information model based on the multimodal highway data.

[0080] The 3D virtual coordinate system establishment module is used to establish a flight path model that corresponds to the actual situation by using the image point coordinates of video frame sequence, the 3D coordinates of real ground points and the perspective geometry of the camera optical center, and thus obtain a 3D virtual coordinate system.

[0081] The automatic fusion module is used to load multimodal highway data into a 3D virtual coordinate system. Based on the indexing mechanism of the semantic information model, it performs semantic parsing of the attribute information, geometric information and display information of the multimodal highway data. Based on the obtained semantic information, it automatically locates, draws and fuses the data in the 3D virtual coordinate system, completing the initial fusion of multimodal highway data and video frame sequences.

[0082] The fusion correction module is used to preview the initial fusion effect of highway multimodal data and video frame sequences, optimize display parameters such as position ratio, angle and color, and output video data after confirming that there are no errors.

[0083] Example 3

[0084] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the fully automatic fusion method for multimodal highway data and aerial video described above.

[0085] Example 4

[0086] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the fully automatic fusion method of highway multimodal data and aerial video as described above.

[0087] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0088] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0089] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0090] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0091] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0092] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A fully automated fusion method for multimodal highway data and aerial video, characterized in that, Includes the following steps: Acquire aerial video footage and perform accelerated frame extraction to obtain a video frame sequence; Acquire multimodal data of highways and construct a semantic information model based on the multimodal data of highways; By using the image point coordinates of video frame sequences, the three-dimensional coordinates of real ground points, and the perspective geometry of the camera's optical center, a flight path model corresponding to reality is established, resulting in a three-dimensional virtual coordinate system. Highway multimodal data is loaded into a 3D virtual coordinate system. Based on the indexing mechanism of the semantic information model, the highway multimodal data is semantically parsed for attribute information, geometric information and display information. Based on the semantic information obtained from the parsing, the data is automatically located, drawn and displayed in the 3D virtual coordinate system, completing the initial fusion of highway multimodal data and video frame sequence. Preview the initial fusion effect of highway multimodal data and video frame sequence, optimize the display position, scale and angle of highway multimodal data in video frame sequence, and output the fused video data; The process of loading multimodal highway data into a three-dimensional virtual coordinate system, performing attribute, geometric, and display semantic analysis on the multimodal highway data based on the indexing mechanism of the semantic information model, and automatically locating, drawing, and displaying the data in the three-dimensional virtual coordinate system based on the obtained semantic information includes: Load all labeled spline animations and parse the semantic information one by one. Obtain the unique identifier of the attribute information to determine the positioning order. Determine the positioning coordinates by the attachment point position of the geometric information. Determine the display effect by the angle and scale of the display information. Finally, automatically locate the attachment point based on the above index information. Load the attribute, geometric, and display information of the horizontal curve table file, automatically draw splines representing straight lines, transition curves, and circular curves that characterize the road direction, and obtain parallel markings based on spline mapping; The unique identifier of the semantic model of the BIM model, the coordinates of the vertices of the minimum bounding rectangle, and the display parameters are parsed, and the model is automatically positioned in the corresponding location and the orientation is displayed correctly. The attribute information includes the unique identifier, name, and file path of the spline animation, the horizontal curve polyline, and the model; the geometric information includes the spline animation attachment point position, the horizontal curve 3D point data, and the vertex coordinates of the model's minimum bounding rectangle; the display information includes the angle, scale, and color of the data displayed in the video frame sequence.

2. The fully automated fusion method for highway multimodal data and aerial video as described in claim 1, characterized in that, The highway multimodal data includes labeled file data, vector line data, and BIM model data.

3. The fully automated fusion method for highway multimodal data and aerial video as described in claim 1, characterized in that, The construction of its semantic information model based on highway multimodal data includes: Create a semantic information model object for highway multimodal data; Define the structure, attributes, and meaning of multimodal highway data semantics; Establish a mapping relationship between semantic information model objects and data semantics, that is, realize a one-to-one correspondence between semantic information models and highway multimodal data, thereby completing the construction of semantic information models.

4. The fully automated fusion method for highway multimodal data and aerial video as described in claim 1, characterized in that, By utilizing the image point coordinates of video frame sequences, the 3D coordinates of real ground points, and the perspective geometry of the camera's optical center, a flight path model corresponding to the actual situation is established, resulting in a 3D virtual coordinate system, including: By using the 3D coordinates of real ground points and the coordinates of image points in a video frame sequence to perform image point motion analysis, the camera exposure position is calculated, the flight path model equation is obtained, and then the ground 3D coordinates of any image point in the video frame are calculated, thus creating a 3D virtual world coordinate system.

5. The fully automated fusion method for highway multimodal data and aerial video as described in claim 1, characterized in that, Optimize the display position, scale, and angle of multimodal highway data in the video frame sequence, adjust shots that do not meet the requirements, and save the rendered and output lossless demonstration video after adjustment.

6. A fully automated fusion system for multimodal highway data and aerial video, characterized in that, A fully automated fusion method for highway multimodal data and aerial video as described in any one of claims 1-5 includes: The video frame acquisition module is used to acquire aerial video and perform accelerated frame extraction processing to obtain a video frame sequence. The semantic information model building module is used to acquire multimodal highway data and build its semantic information model based on the multimodal highway data. The 3D virtual coordinate system establishment module is used to establish a flight path model that corresponds to the actual situation by using the image point coordinates of video frame sequence, the 3D coordinates of real ground points and the perspective geometry of the camera optical center, and thus obtain a 3D virtual coordinate system. The automatic fusion module is used to load highway multimodal data into a three-dimensional virtual coordinate system. Based on the indexing mechanism of the semantic information model, it performs semantic parsing of the attribute information, geometric information and display information of the highway multimodal data, and automatically locates, draws and displays it in the three-dimensional virtual coordinate system based on the obtained semantic information, thus completing the initial fusion of highway multimodal data and video frame sequences. The fusion correction module is used to preview the initial fusion effect of highway multimodal data and video frame sequences, optimize the display position, scale and angle of highway multimodal data in video frame sequences, and output the fused video data.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the fully automatic fusion method of highway multimodal data and aerial video as described in any one of claims 1-5.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the fully automatic fusion method of highway multimodal data and aerial video as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Mobile-end 3D city dynamic modeling method based on geographic semantics

    CN106952330A

  • Aerial video and road three-dimensional model live-action synthesis method

    CN115797549A