Panoramic video three-dimensional splicing method and system

By building three-dimensional splicing of the site field three-dimensional model and real-time video stream, combined with the rendering of the three-dimensional GIS engine, the problem of lack of spatial and depth in the existing panoramic video stitching is solved, and a more realistic and intuitive video stitching effect is achieved.

CN120075421APending Publication Date: 2025-05-30AEROSPACE JICHUANG IOT RES INST (NANJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510106399.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing panoramic video stitching methods lack the use of three-dimensional spatial information, resulting in the lack of spatial and depth sense of video stitching results.

Method used

By constructing a three-dimensional model of the converter station, collecting real-time video streams of multiple stations, extracting key feature points of the content, matching and fusion, completing three-dimensional stitching, and fusing the panoramic video stream with the three-dimensional model, and rendering and outputting through the three-dimensional GIS engine in real time.

Benefits of technology

The spatial and depth sense of splicing results is improved, making the video splicing results more realistic and intuitive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075421A_ABST
    Figure CN120075421A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional splicing method and system for a panoramic video, and relates to the field related to computer vision, and the method comprises the steps: constructing a station yard three-dimensional model of a converter station; collecting a plurality of station yard real-time video streams of different positions and different visual angles in the converter station area; extracting a plurality of content key feature points representing video contents from the plurality of station real-time video streams; matching and fusing the plurality of content key feature points based on a station panoramic video feature processing network, and completing three-dimensional splicing of a plurality of station real-time video streams to obtain a station panoramic video stream; and fusing the panoramic video stream of the station yard with the three-dimensional model of the station yard, and rendering and outputting in real time through a three-dimensional GIS engine after fusion. The technical problem that a video stitching result lacks a space sense and a depth sense due to separation of a panoramic video and a three-dimensional model in existing panoramic video stitching is solved, and the technical effect of improving the space sense and the depth sense of the stitching result is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and particularly to a three-dimensional stitching method and system for panoramic videos. Background Art

[0002] With the rapid development of the power industry, converter stations are playing an increasingly important role in the power grid. As the core component of a high-voltage direct current (HVDC) transmission system, the safety and stability of the operating state of a converter station are directly related to the reliable operation of the entire power grid. Therefore, real-time monitoring and visual display of converter stations have become an important requirement in the power industry. Most of the existing panoramic video stitching methods are based on two-dimensional image processing technologies and do not fully utilize the three-dimensional spatial information of converter stations, resulting in the lack of a sense of space and depth in the video stitching results.

[0003] In the related technologies at the present stage, there is a technical problem in panoramic video stitching that the panoramic video is separated from the three-dimensional model, resulting in the lack of a sense of space and depth in the video stitching results. Summary of the Invention

[0004] This application provides a three-dimensional stitching method and system for panoramic videos. By constructing a three-dimensional model of the converter station yard, collecting multiple real-time video streams of the yard, extracting multiple content key feature points, performing matching and fusion, completing the stereo stitching of the real-time video streams of the yard, fusing the panoramic video stream of the yard with the three-dimensional model of the yard, and using technical means such as real-time rendering and output through a three-dimensional GIS engine, the technical effect of effectively fusing the panoramic video with the three-dimensional model to improve the sense of space and depth of the stitching result is achieved.

[0005] This application provides a three-dimensional stitching method for panoramic videos, including:

[0006] Constructing a three-dimensional model of the converter station yard; collecting multiple real-time video streams of the yard at different positions and from different perspectives within the area of the converter station; extracting multiple content key feature points representing the video content from the multiple real-time video streams of the yard; performing matching and fusion on the multiple content key feature points based on a panoramic video feature processing network for the yard to complete the stereo stitching of the multiple real-time video streams of the yard and obtain a panoramic video stream of the yard; fusing the panoramic video stream of the yard with the three-dimensional model of the yard, and performing real-time rendering and output through a three-dimensional GIS engine after fusion.

[0007] This application also provides a three-dimensional stitching system for panoramic videos, including:

[0008] A station three-dimensional model construction module, which is used to construct a three-dimensional model of the substation yard; a station real-time video stream collection module, which is used to collect multiple station real-time video streams at different positions and from different perspectives within the substation yard area; a content key feature point extraction module, which is used to extract multiple content key feature points representing the video content from the multiple station real-time video streams; a video stream three-dimensional splicing module, which is used to match and fuse the multiple content key feature points based on a station panoramic video feature processing network, complete the three-dimensional splicing of the multiple station real-time video streams, and obtain a station panoramic video stream; a panoramic video stream and three-dimensional model fusion module, which is used to fuse the station panoramic video stream with the station three-dimensional model, and after fusion, render and output in real time through a three-dimensional GIS engine.

[0009] It is intended to propose a three-dimensional splicing method and system for panoramic video through this application. First, construct a three-dimensional model of the substation yard, then collect multiple station real-time video streams at different positions and from different perspectives within the substation yard area, then extract multiple content key feature points representing the video content from the multiple station real-time video streams, and then match and fuse the multiple content key feature points based on a station panoramic video feature processing network, complete the three-dimensional splicing of the multiple station real-time video streams, and obtain a station panoramic video stream. Finally, fuse the station panoramic video stream with the station three-dimensional model, and after fusion, render and output in real time through a three-dimensional GIS engine, achieving the technical effect of improving the sense of space and depth of the splicing result by effectively fusing the panoramic video with the three-dimensional model. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations in the front or below do not necessarily need to be executed precisely in sequence. On the contrary, according to needs, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0011] Figure 1 It is a schematic flowchart of a three-dimensional splicing method for panoramic video provided by an embodiment of the present application;

[0012] Figure 2 It is a schematic structural diagram of a three-dimensional splicing system for panoramic video provided by an embodiment of the present application.

[0013] Description of the accompanying drawing reference numerals: Substation three-dimensional model construction module 10, substation real-time video stream collection module 20, content key feature point extraction module 30, video stream stereo stitching module 40, panoramic video stream and three-dimensional model fusion module 50. Detailed implementation manners

[0014] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented in accordance with the content of the description. And in order to make the above and other objects, features and advantages of the present application more obvious and understandable, the following specifically gives the detailed implementation manners of the present application.

[0015] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0016] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first" and "second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.

[0017] The embodiments of the present application provide a three-dimensional stitching method for panoramic videos, as Figure 1 shown. The method includes:

[0018] Step S100, construct a three-dimensional model of the substation yard. Specifically, collect the detailed geographical information of the substation yard (a facility in the power system for converting electrical energy between direct current and alternating current), including data such as buildings, equipment, roads, terrain, etc. Use three-dimensional modeling software (such as AutoCAD, SketchUp, Blender, etc.) to create a virtual three-dimensional model of the substation yard according to the collected data, including constructing the geometric shapes, materials and textures of each object. Make detailed adjustments and optimizations to the three-dimensional model of the substation yard to make it visually close to the real environment.

[0019] In a possible implementation, when constructing a three-dimensional model of the substation yard, step S100 further includes:

[0020] Step S110, interactively obtain the geographical information data of the area where the substation is located, the building structure information of the substation, and the equipment layout information. Specifically, collect relevant data from various sources such as satellite images, aerial photos, on-site surveying and mapping, design drawings, and equipment lists through means such as downloading, scanning, and measuring. Among them, the geographical information data refers to the data describing the surface spatial distribution, attributes, and temporal changes of the area where the substation is located, including information such as topography and road network; the building structure information refers to the internal and external structural characteristics of the substation buildings, such as the shape, size, and material of walls, columns, beams, roofs, etc.; the equipment layout information refers to the installation positions, arrangement methods, and interconnection relationships of various equipment in the substation. Step S120, based on the geographical information data, the building structure information, and the equipment layout information, construct a three-dimensional model of the substation yard through three-dimensional modeling. Specifically, use three-dimensional modeling software such as AutoCAD, SketchUp, 3dsMax, etc. According to the geographical information data, construct a basic terrain model of the area where the substation is located, including terrain undulations, road network, etc.; according to the building structure information, construct three-dimensional models of each building in the substation one by one, including drawing the outlines of the buildings, adding details such as doors, windows, and stairs, and assigning appropriate materials and textures; on the basis of the building structure model, according to the equipment layout information, place various equipment into the three-dimensional model of the substation yard according to the actual positions and postures. This implementation method ensures the high accuracy of the constructed three-dimensional model in terms of geographical spatial position, building structure details, and equipment layout by interactively obtaining the geographical information data, building structure information, and equipment layout information of the area where the substation is located, achieving the technical effect of improving the accuracy of constructing the three-dimensional model of the substation yard.

[0021] Step S200, collect multiple real-time video streams of the substation yard at different positions and from different perspectives in the substation area. Specifically, install multiple cameras in the substation area to cover all corners and important areas of the yard, and transmit the videos captured by the cameras to the processing center in real time through video transmission technologies (such as network video transmission, fiber optic transmission, etc.). Among them, the real-time video stream of the substation yard refers to a continuous sequence of video images of various elements in the substation area captured.

[0022] Step S300: Extract multiple content key feature points representing the video content from the multiple station real-time video streams. Specifically, preprocess the received station real-time video streams, such as denoising, enhancing contrast, etc., to improve the accuracy of subsequent processing. Use computer vision algorithms (such as SIFT, SURF, ORB, etc.) to extract content key feature points from the video frames of the station real-time video streams. Content key feature points are significant or unique parts in the image, such as corner points, edge intersection points, etc.

[0023] In a possible implementation, when extracting multiple content key feature points representing the video content from the multiple station real-time video streams, step S300 further includes:

[0024] Step S310, read the first frame of the station image from the multiple real-time video streams of the stations. Specifically, select any one of the multiple real-time video streams of the stations as the current processing object, decode the first frame data of the real-time video stream of the station, and extract the complete image from the decoded first frame data, that is, the first frame of the station image. Step S320, perform grayscale conversion on the first frame of the station image to obtain the first grayscale station image. Specifically, use a grayscale algorithm, such as the weighted average method, the maximum value method, the average value method, etc., to calculate the grayscale value of each pixel point of the first frame of the station image, and assign the calculated grayscale value to the corresponding pixel point to generate the first grayscale station image. Step S330, calculate the corner response value of each pixel point in the first grayscale station image to obtain the first set of corner response values of the station image. Specifically, use a corner detection algorithm, such as the Harris corner detection, the Shi-Tomasi corner detection, etc., to calculate the corner response value of each pixel point of the first grayscale station image. The corner response value is a numerical value used to represent the likelihood that the pixel point is a corner. Step S340, based on the first set of corner response values of the station image, compare the corner response values within a preset neighborhood matrix, extract the maximum value points, and obtain the second set of corner response values of the station image. Specifically, set a preset-size neighborhood matrix (such as 3x3, 5x5, etc.), traverse each element in the first set of corner response values of the station image, compare the corner response values within the neighborhood matrix with it as the center. If the corner response value of the current element is the largest within the neighborhood, then retain this element, and form the second set of corner response values of the station image by combining all the elements corresponding to the local maximum values. Step S350, set a global threshold, traverse the second set of corner response values of the station image, and extract the pixel points corresponding to the corner response values greater than the global threshold to obtain the multiple content key feature points. Specifically, set a global threshold according to actual needs for screening strong corners. Traverse each element in the second set of corner response values of the station image, compare its corner response value with the global threshold. If the corner response value is greater than the global threshold, then regard the pixel point corresponding to this element as a content key feature point, and store all the screened content key feature points to form a set of content key feature points. This implementation method uses a preset neighborhood matrix and a global threshold for screening, eliminates unstable or redundant feature points, and achieves the technical effect of improving the accuracy of content key feature point extraction.

[0025] Step S400: Based on the station panoramic video feature processing network, match and fuse the multiple content key feature points to complete the stereo stitching of the multiple real-time station video streams, and obtain the station panoramic video stream. Specifically, use a feature matching algorithm (such as FLANN, BFMatcher, etc.) to match the content key feature points from different real-time station video streams, and find the corresponding relationships between the content key feature points. According to the matching results, splice the image segments in the multiple real-time station video streams to form a continuous and seamless station panoramic video stream. Among them, the station panoramic video feature processing network is a deep learning-based network structure used to process image features and optimize the matching and stitching effects; stereo stitching refers to the process of combining image or video segments from multiple perspectives into a panoramic image or video with a broader view.

[0026] In a possible implementation manner, based on the station panoramic video feature processing network, match and fuse the multiple content key feature points to complete the stereo stitching of the multiple real-time station video streams, and obtain the station panoramic video stream. Step S400 further includes:

[0027] Step S410: Construct a content feature point vector generation sub-network, a content feature point matching sub-network, and a content feature point fusion sub-network. Sequentially connect the content feature point vector generation sub-network, the content feature point matching sub-network, and the content feature point fusion sub-network to complete the construction of the station panoramic video feature processing network. Specifically, design three sub-networks. Among them, the content feature point vector generation sub-network can be based on a deep learning network structure (such as a convolutional neural network or an autoencoder) and is used to extract the feature vectors of each content key feature point; the content feature point matching sub-network is a network structure for comparing and matching feature vectors, such as a similarity calculation layer or a distance metric learning layer; the content feature point fusion sub-network is used to fuse the matched feature points to generate corresponding positions in a unified video frame or video stream. Step S420: Generate the content feature vectors of the multiple content key feature points through the content feature point vector generation sub-network, match the multiple content feature vectors through the content feature point matching sub-network, and fuse the multiple matched content key feature points through the content feature point fusion sub-network. Specifically, input each content key feature point and its surrounding pixel information into the content feature point vector generation sub-network, and output the feature vector of each content key feature point. Use the content feature point matching sub-network to compare the feature vectors of the content feature points from different station real-time video streams and output the matching result, that is, which content key feature points are considered to correspond to each other. According to the matching result, fuse the corresponding content key feature points through the content feature point fusion sub-network. The fusion can adopt weighted average of pixel values, interpolation, or other image processing techniques. Step S430: Based on the fusion result, complete the stereo stitching of the multiple station real-time video streams to obtain a station panoramic video stream. Specifically, based on the fused content key feature points, fuse the corresponding video frames. During the fusion process, the size, rotation angle, or perspective effect of the video frames can be adjusted to ensure the coherence of the video stream. Arrange the fused video frames in sequence to form a complete station panoramic video stream. This implementation method efficiently processes the station real-time video streams from multiple cameras by constructing an integrated station panoramic video feature processing network, achieving the technical effects of improving the accuracy and efficiency of video stitching.

[0028] In a possible implementation, by using the content feature point vector generation sub-network to generate the content feature vectors of the multiple content key feature points, using the content feature point matching sub-network to match the multiple content feature vectors, and using the content feature point fusion sub-network to fuse the multiple matched content key feature points, step S420 further includes:

[0029] Step S421: Based on the multiple content key feature points, extract the first content key feature point, the second content key feature point, and the third content key feature point. Specifically, from all the detected content key feature points, according to the real-time video stream of the station yard to which the content key feature point belongs, extract any one from the multiple content key feature points corresponding to the first station yard real-time video stream as the first content key feature point, and extract any two without replacement from the multiple content key feature points corresponding to the second station yard real-time video stream as the second content key feature point and the third content key feature point. Step S422: Input the first content key feature point, the second content key feature point, and the third content key feature point into the content feature point vector generation sub-network to generate content feature vectors, obtaining the first content feature vector, the second content feature vector, and the third content feature vector. Specifically, take the first, second, and third content key feature points and the pixel information within a preset range around them as inputs, and through the content feature point vector generation sub-network (such as CNN or deep auto-encoder), perform feature extraction on the input information to generate corresponding feature vectors, namely the first content feature vector, the second content feature vector, and the third content feature vector. Step S423: Transfer the first content feature vector, the second content feature vector, and the third content feature vector to the content feature point matching sub-network, and calculate the first content similarity between the first content feature vector and the second content feature vector, and the second content similarity between the first content feature vector and the third content feature vector. Specifically, input the first, second, and third content feature vectors into the content feature point matching sub-network, calculate the similarity (such as cosine similarity, Euclidean distance, etc.) between the first content feature vector and the second content feature vector to obtain the first content similarity; similarly calculate the similarity between the first content feature vector and the third content feature vector to obtain the second content similarity. Step S424: If the ratio of the first content similarity to the second content similarity is greater than 1, then match the first content key feature point and the second content key feature point through the content feature point matching sub-network, and output the first pair of matching content key feature points. Specifically, calculate the ratio of the first content similarity to the second content similarity. If this ratio is greater than 1, it means that the similarity of the first pair of content key feature points is higher than that of the second pair, then it is considered that the first content key feature point and the second content key feature point are a pair of matching feature points. Step S425: Input the first pair of matching content key feature points into the content feature point fusion sub-network for coordinate alignment of the corresponding station yard images, and apply the Laplacian pyramid to smooth the overlapping area after coordinate alignment, and output the first pair of matching content fusion images after splicing. Specifically, align the coordinate positions of the image regions corresponding to the two feature points in the first pair of matching content key feature points to ensure that they are consistent in spatial position.Apply the Laplace pyramid to smooth the overlapping area after coordinate alignment to reduce seams or artifacts that may occur during the fusion process, and fuse the processed image area into the target panoramic image to output the first matching content fusion image after stitching. This implementation method avoids mis-matching caused by improper setting of a single similarity threshold by comparing the ratio of content similarity, achieving the technical effect of improving the accuracy of matching key content feature points.

[0030] Step S500: Fuse the station yard panoramic video stream with the station yard three-dimensional model, and output the fused result through real-time rendering by a three-dimensional GIS engine. Specifically, register and fuse the station yard panoramic video stream with the station yard three-dimensional model to make the content in the station yard panoramic video stream consistent with the positions and directions of the objects in the station yard three-dimensional model. Use a three-dimensional GIS engine (such as ArcGIS, Cesium, Unity, etc.) to perform real-time rendering on the fused data to generate a three-dimensional panoramic video that can be displayed on a computer screen. In the embodiment of the present application, by constructing a three-dimensional model of the converter station yard, collecting multiple real-time video streams of the station yard, extracting multiple key content feature points, performing matching and fusion, completing the stereoscopic stitching of the real-time video stream of the station yard, fusing the station yard panoramic video stream with the station yard three-dimensional model, and outputting the fused result through real-time rendering by a three-dimensional GIS engine and other technical means, the technical effect of improving the sense of space and depth of the stitching result is achieved by effectively fusing the panoramic video with the three-dimensional model.

[0031] In a possible implementation method, when fusing the station yard panoramic video stream with the station yard three-dimensional model and outputting the fused result through real-time rendering by a three-dimensional GIS engine, step S500 further includes:

[0032] Step S510: Generate a dynamic texture layer based on the panoramic video stream of the substation yard and a static base layer based on the 3D model of the substation yard. Specifically, extract consecutive video frames from the panoramic video stream of the substation yard and preprocess the video frames, including denoising, color correction, etc., to ensure the video quality. Use the processed video frames as textures and dynamically generate a texture layer according to the update frequency of the video frames. The dynamic texture layer contains the real-time visual information of the converter station yard. Load the 3D model data of the substation yard, and perform necessary preprocessing on the 3D model of the substation yard, such as coordinate transformation, scaling, rotation, etc., to ensure its conformity with the actual scene. Use the processed 3D model of the substation yard as the static base layer, and the static base layer provides the spatial structure and basic visual information of the converter station yard. Step S520: With the spatial coordinate consistency as the constraint, map the dynamic texture layer to the surface of the static base layer in real time to obtain a panoramic video fusion layer. Specifically, match and calibrate the feature points in the video frames with the corresponding points in the 3D model of the substation yard, and use a spatial transformation matrix to convert the coordinate system of the video frames into the same coordinate system as the 3D model of the substation yard. Map each frame of the dynamic texture layer to the corresponding surface of the static base layer in real time, that is, apply the video frames as texture maps to the corresponding patches of the 3D model of the substation yard. Step S530: Render the visual effect of the panoramic video fusion layer and then output it. Specifically, configure the rendering parameters of the 3D GIS engine, including lighting, shadows, material reflections, etc., to enhance the visual effect. Set the position, angle, and focal length of the camera as needed to obtain the best visual experience. Use the 3D GIS engine to render the panoramic video fusion layer, including merging the dynamic texture layer and the static base layer into a whole and applying various visual effects. Output the rendered video frames to a display device (such as a monitor, projector, etc.) for users to view. This implementation method provides users with a more real and intuitive visual experience of the converter station yard by fusing the panoramic video stream of the substation yard and the 3D model of the substation yard, achieving the technical effect of enhancing the sense of reality.

[0033] In a possible implementation manner, to generate a dynamic texture layer based on the panoramic video stream of the substation yard and a static base layer based on the 3D model of the substation yard, step S510 further includes:

[0034] Step S511: Configure the layer attributes for the dynamic texture layer and the static base layer. Specifically, the layer attribute configuration includes display or hiding. Specifically, the layer attribute refers to the characteristics or states of a layer, including display (visibility) or hiding (invisibility). Display / hiding is the visibility attribute of a layer, which is used to control whether the layer is displayed in the current view. Step S512: Map the layer attributes to the user interface, and perform display interaction or hiding interaction for the dynamic texture layer and the static base layer according to the mapping result. Specifically, according to the requirements of the application scenario, design a user interface with layer control functions. The user interface is an interface that includes controls such as check boxes, buttons, or sliders, allowing users to directly operate the display / hiding attributes of the layer. Write code to map the layer attributes (such as the display / hiding states of the dynamic texture layer and the static base layer) to the corresponding controls on the user interface. For example, when the user clicks a certain button, this operation will trigger an event, and this event will modify the display / hiding state of the corresponding layer and be immediately reflected in the 3D view. Provide instant feedback when the user interacts with the interface. For example, when the user changes the display state of a layer, the interface is immediately updated to reflect this change, and the 3D view is also updated accordingly. This implementation method improves the user experience by allowing users to freely control the display / hiding of layers according to their needs. By hiding some layers, the view is simplified, achieving the technical effect of making it easier for users to understand and focus on the most important information currently.

[0035] In the foregoing, with reference to Figure 1 A three-dimensional stitching method for panoramic videos according to an embodiment of the present invention is described in detail. Next, with reference to Figure 2 A three-dimensional stitching system for panoramic videos according to an embodiment of the present invention will be described.

[0036] A three-dimensional stitching system for panoramic videos according to an embodiment of the present invention is used to solve the technical problem that in the existing panoramic video stitching, the panoramic video is separated from the three-dimensional model, resulting in the lack of a sense of space and depth in the video stitching result, and achieves the technical effect of improving the sense of space and depth of the stitching result by effectively integrating the panoramic video with the three-dimensional model. A three-dimensional stitching system for panoramic videos includes: a station three-dimensional model construction module 10, a station real-time video stream collection module 20, a content key feature point extraction module 30, a video stream stereo stitching module 40, and a panoramic video stream and three-dimensional model fusion module 50.

[0037] The substation three-dimensional model construction module 10 is used to construct a three-dimensional model of the substation; the substation real-time video stream collection module 20 is used to collect multiple substation real-time video streams at different positions and from different perspectives within the substation area; the content key feature point extraction module 30 is used to extract multiple content key feature points representing the video content from the multiple substation real-time video streams; the video stream stereo stitching module 40 is used to match and fuse the multiple content key feature points based on the substation panoramic video feature processing network, complete the stereo stitching of the multiple substation real-time video streams, and obtain a substation panoramic video stream; the panoramic video stream and three-dimensional model fusion module 50 is used to fuse the substation panoramic video stream with the substation three-dimensional model, and after fusion, it is rendered and output in real time through a three-dimensional GIS engine.

[0038] Next, the specific configuration of the content key feature point extraction module 30 will be described in detail. As described above, multiple content key feature points representing the video content are extracted from the multiple substation real-time video streams. The content key feature point extraction module 30 may further include: a first-frame substation image reading unit for reading the first-frame substation images of the multiple substation real-time video streams; a grayscale conversion unit for performing grayscale conversion on the first-frame substation images to obtain first grayscale substation images; a corner response value calculation unit for calculating the corner response values of each pixel point in the first grayscale substation images to obtain a first substation image corner response value set; a maximum value point extraction unit for comparing the corner response values within a preset neighborhood matrix based on the first substation image corner response value set and extracting the maximum value points to obtain a second substation image corner response value set; a pixel point extraction unit for setting a global threshold, traversing the second substation image corner response value set, and extracting the pixel points corresponding to the corner response values greater than the global threshold to obtain the multiple content key feature points.

[0039] Next, the specific configuration of the video stream stereo stitching module 40 will be described in detail. As described above, based on the station yard panoramic video feature processing network, the multiple content key feature points are matched and fused to complete the stereo stitching of the multiple station yard real-time video streams, and a station yard panoramic video stream is obtained. The video stream stereo stitching module 40 may further include: a station yard panoramic video feature processing network construction unit for constructing a content feature point vector generation sub-network, a content feature point matching sub-network, and a content feature point fusion sub-network, and connecting the content feature point vector generation sub-network, the content feature point matching sub-network, and the content feature point fusion sub-network in sequence to complete the construction of the station yard panoramic video feature processing network; a feature point processing unit for generating content feature vectors of the multiple content key feature points through the content feature point vector generation sub-network, matching multiple content feature vectors through the content feature point matching sub-network, and fusing multiple matched content key feature points through the content feature point fusion sub-network; a station yard panoramic video stream acquisition unit for completing the stereo stitching of the multiple station yard real-time video streams based on the fusion result to obtain a station yard panoramic video stream.

[0040] Among them, the content feature vector generation of the multiple content key feature points is performed by the content feature point vector generation sub-network, the matching of multiple content feature vectors is performed by the content feature point matching sub-network, and the fusion of multiple matched content key feature points is performed by the content feature point fusion sub-network. The feature point processing unit may further include: a content key feature point extraction sub-unit for extracting a first content key feature point, a second content key feature point, and a third content key feature point based on the multiple content key feature points; a content feature vector generation sub-unit for inputting the first content key feature point, the second content key feature point, and the third content key feature point into the content feature point vector generation sub-network for content feature vector generation to obtain a first content feature vector, a second content feature vector, and a third content feature vector; a content similarity calculation sub-unit for transferring the first content feature vector, the second content feature vector, and the third content feature vector to the content feature point matching sub-network to calculate a first content similarity between the first content feature vector and the second content feature vector, and a second content similarity between the first content feature vector and the third content feature vector; a content key feature point matching sub-unit for, if the ratio of the first content similarity to the second content similarity is greater than 1, matching the first content key feature point and the second content key feature point through the content feature point matching sub-network and outputting a first matched content key feature point pair; a matched content image fusion sub-unit for inputting the first matched content key feature point pair into the content feature point fusion sub-network for coordinate alignment of the corresponding yard image, applying a Laplacian pyramid to smooth the overlapping area after coordinate alignment, and outputting a first matched content fusion image after splicing.

[0041] Next, the specific configuration of the panoramic video stream and 3D model fusion module 50 will be described in detail. As described above, the yard panoramic video stream and the yard 3D model are fused and rendered in real time through a 3D GIS engine after fusion. The panoramic video stream and 3D model fusion module 50 may further include: a layer generation unit for generating a dynamic texture layer based on the yard panoramic video stream and generating a static base layer based on the yard 3D model; a layer mapping unit for, with spatial coordinate consistency as a constraint, mapping the dynamic texture layer onto the surface of the static base layer in real time to obtain a panoramic video fusion layer; a visual effect rendering unit for performing visual effect rendering on the panoramic video fusion layer and then outputting.

[0042] Among them, after generating the dynamic texture layer based on the panoramic video stream of the station yard and generating the static base layer based on the 3D model of the station yard, the layer generation unit may further include: a layer attribute configuration subunit for configuring the layer attributes of the dynamic texture layer and the static base layer, where the layer attribute configuration includes display or hiding; an interaction subunit for mapping the layer attributes to the user interaction interface and performing display interaction or hiding interaction on the dynamic texture layer and the static base layer according to the mapping result.

[0043] Next, the specific configuration of the station yard 3D model construction module 10 will be described in detail. As described above, to construct the 3D model of the station yard of the converter station, the station yard 3D model construction module 10 may further include: a converter station related information acquisition unit for interactively acquiring the geographical information data of the area where the converter station is located, the building structure information of the converter station, and the equipment layout information; a 3D modeling unit for constructing the 3D model of the station yard of the converter station through 3D modeling based on the geographical information data, the building structure information, and the equipment layout information.

[0044] The 3D stitching system for panoramic video provided by the embodiments of the present invention can execute the 3D stitching method for panoramic video provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0045] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or the server. The various units and modules included are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0046] The above specific implementation manners do not constitute a limitation to the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application. In some cases, the actions or steps recorded in this application can be executed in a different order from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A three-dimensional stitching method for panoramic videos, characterized in that: The method comprises: Construct a three-dimensional model of the converter station; Collecting multiple real-time video streams of stations at different locations and different viewing angles within the converter station area; Extracting a plurality of key feature points representing video content from the plurality of real-time video streams of the stations; Matching and fusing the multiple content key feature points based on the station field panoramic video feature processing network, completing stereo stitching of the multiple station field real-time video streams, and obtaining a station field panoramic video stream; The panoramic video stream of the station is fused with the three-dimensional model of the station, and after fusion, it is rendered and output in real time through a three-dimensional GIS engine.

2. A three-dimensional stitching method for panoramic videos as claimed in claim 1, characterized in that: The step of extracting a plurality of key feature points representing video content from the plurality of real-time video streams of the stations includes: Reading the first frame of the station field image of the plurality of station field real-time video streams; Performing grayscale conversion on the first frame of station field image to obtain a first grayscale station field image; Calculate the corner point response value of each pixel in the first grayscale station field image to obtain a set of corner point response values ​​of the first station field image; Based on the first station field image corner point response value set, compare the corner point response values ​​in a preset neighborhood matrix, extract the maximum value point, and obtain the second station field image corner point response value set; A global threshold is set, the set of corner point response values ​​of the second station field image is traversed, and pixel points corresponding to corner point response values ​​greater than the global threshold are extracted to obtain the multiple content key feature points.

3. The three-dimensional stitching method of panoramic video according to claim 1, characterized in that: The station field panoramic video feature processing network is used to match and fuse the multiple content key feature points, complete the stereo stitching of the multiple station field real-time video streams, and obtain the station field panoramic video stream, including: Constructing a content feature point vector generation subnetwork, a content feature point matching subnetwork, and a content feature point fusion subnetwork, and sequentially connecting the content feature point vector generation subnetwork, the content feature point matching subnetwork, and the content feature point fusion subnetwork to complete the construction of the station field panoramic video feature processing network; The content feature vectors of the plurality of content key feature points are generated by the content feature point vector generation sub-network, the plurality of content feature vectors are matched by the content feature point matching sub-network, and the plurality of matched content key feature points are fused by the content feature point fusion sub-network; Based on the fusion result, the stereoscopic stitching of the multiple real-time video streams of the station is completed to obtain a panoramic video stream of the station.

4. A three-dimensional stitching method for panoramic videos as claimed in claim 3, characterized in that: The step of generating content feature vectors of the plurality of content key feature points by the content feature point vector generating sub-network, matching the plurality of content feature vectors by the content feature point matching sub-network, and fusing the plurality of matched content key feature points by the content feature point fusion sub-network includes: Based on the multiple content key feature points, extracting a first content key feature point, a second content key feature point, and a third content key feature point; Inputting the first content key feature point, the second content key feature point and the third content key feature point into the content feature point vector generation subnetwork to generate content feature vectors, thereby obtaining a first content feature vector, a second content feature vector and a third content feature vector; The first content feature vector, the second content feature vector and the third content feature vector are transferred to the content feature point matching sub-network, and a first content similarity between the first content feature vector and the second content feature vector and a second content similarity between the first content feature vector and the third content feature vector are calculated; If the ratio of the first content similarity to the second content similarity is greater than 1, matching the first content key feature point with the second content key feature point through a content feature point matching subnetwork, and outputting a first matching content key feature point pair; The first matching content key feature point pair is input into the content feature point fusion subnetwork to perform coordinate alignment of the corresponding station field image, the Laplacian pyramid is applied to smooth the overlapping area after the coordinate alignment, and the spliced ​​first matching content fusion image is output.

5. The three-dimensional stitching method of panoramic video according to claim 1, characterized in that: The step of fusing the panoramic video stream of the station field with the three-dimensional model of the station field and outputting the fusion through real-time rendering by a three-dimensional GIS engine includes: Generate a dynamic texture layer based on the panoramic video stream of the station, and generate a static base layer based on the three-dimensional model of the station; Taking spatial coordinate consistency as a constraint, mapping the dynamic texture layer to the surface of the static base layer in real time to obtain a panoramic video fusion layer; The panoramic video fusion layer is rendered with visual effects and then outputted.

6. A three-dimensional stitching method for panoramic videos as claimed in claim 5, characterized in that: After the dynamic texture layer is generated based on the panoramic video stream of the station field and the static base layer is generated based on the three-dimensional model of the station field, the method further includes: Performing layer attribute configuration on the dynamic texture layer and the static base layer, wherein the layer attribute configuration includes displaying or hiding; The layer attributes are mapped to a user interaction interface, and display interaction or hiding interaction between the dynamic texture layer and the static base layer is performed according to the mapping result.

7. The three-dimensional stitching method of panoramic video according to claim 1, characterized in that: The construction of the three-dimensional model of the converter station includes: Interactively obtain geographic information data of the area where the converter station is located, building structure information and equipment layout information of the converter station; Based on the geographic information data, the building structure information and the equipment layout information, a three-dimensional model of the converter station is constructed through three-dimensional modeling.

8. A three-dimensional stitching system for panoramic videos, characterized in that: The system is used to implement a three-dimensional stitching method of a panoramic video according to any one of claims 1 to 7, and the system comprises: A station field three-dimensional model construction module, wherein the station field three-dimensional model construction module is used to construct a station field three-dimensional model of a converter station; A station field real-time video stream collection module, the station field real-time video stream collection module is used to collect multiple station field real-time video streams at different positions and different viewing angles in the converter station area; A content key feature point extraction module, the content key feature point extraction module is used to extract a plurality of content key feature points representing video content from the plurality of station field real-time video streams; A video stream stereo stitching module, the video stream stereo stitching module is used to match and merge the multiple content key feature points based on the station field panoramic video feature processing network, complete the stereo stitching of the multiple station field real-time video streams, and obtain the station field panoramic video stream; A panoramic video stream and three-dimensional model fusion module is used to fuse the panoramic video stream of the station with the three-dimensional model of the station, and after fusion, it is rendered and output in real time through a three-dimensional GIS engine.

Citation Information

Cited By

  • Railway station three-dimensional video generation method and system based on adaptive visual angle

    CN120751105A

  • Method for rapidly processing 2D video and reconstructing three-dimensional model color fusion based on GPU

    CN122336125A