Video tag generation method based on geographic information fusion
By calibrating the camera rotation and transparent background in the GIS map three-dimensional scene, and combining with the network camera to realize video tag generation, the high cost and limitations of the existing technology are solved, and low-cost video tag generation and long-distance picture width maintenance are achieved.
Patent Information
- Application Number
- CN202311544613.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing video tag generation method is costly and has great limitations. AR video tags require high-cost equipment. Map and video fusion is affected by camera zoom, resulting in slender and narrow pictures, making it impossible to efficiently utilize GIS platform data.
By building a three-dimensional GIS map scene, combining a network camera, using webrtc-streamer.js for real-time video playback, calibrating the camera rotation, obtaining the field of view angle and focal length values, transparentizing the GIS map background, superimposing it to the camera video screen, and realizing video tag generation.
It effectively solves the problems of high cost of AR video tags and the limitations of map video fusion, realizes low-cost video tag generation, and maintains the picture width during long-distance viewing, and does not require secondary input of GIS platform data.
Smart Images

Figure CN120343180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating video tags, and in particular to a method for generating video tags based on geographic information fusion, belonging to the technical field of video tag generation. Background Art
[0002] Currently, how to make scientific and rapid correct decisions is a problem that every organization in the big data era needs to face. With the gradual maturity of high-definition video technology, data transmission, backend storage and other supporting technologies, it provides a reliable path for video data to assist in intelligent decision-making. However, simply "brutally" stacking each video frame on a large screen cannot enable decision-makers to extract effective information from the chaotic data, and data assets cannot be realized. How to build a clear and efficient intelligent decision-making system? "Real-time video + data information" may be the correct answer.
[0003] The key to its implementation lies in the integrated application of data in various dimensions. One of the difficulties is the integrated application of unstructured data such as video data. Currently, relatively mature applications include AR video tags, map video fusion, etc. However, AR video tags require the use of supporting panoramic monitoring equipment, which has disadvantages such as high cost, inability to utilize existing equipment, and the need for secondary entry of GIS platform data; while map video fusion is affected by camera zoom. When viewing distant images, the fused video will become slender and narrow, and only a fixed-focus bullet camera can be selected as the video data hardware, which has limitations. Therefore, how to directly use the existing data on the GIS platform as the data source for video tags has become the final technical direction of "real-time video + data information". Summary of the Invention
[0004] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to follow.
[0005] In view of this, in order to solve the technical problems of high cost and large limitations in the existing video tag generation method, the present invention provides a method for generating video tags based on geographic information fusion.
[0006] Solution 1. A method for generating video tags based on geographic information fusion, comprising the following steps:
[0007] S1. Build a GIS Figure 3 dimensional scene;
[0008] S2. Control access of network cameras;
[0009] S3. Call webrtc - streamer.js for real - time video playback and display;
[0010] S4. Calibrate the GIS map with the camera image and guide the camera to rotate;
[0011] S5. Select any point O in the GIS map, guide the camera to rotate to obtain the horizontal value, pitch value, and focal length value, and get ∠APB, where P is the camera position, and A and B are the illumination widths of the perspective camera;
[0012] S6. Obtain the camera field - of - view angle ∠ACB of the 3D map, where C is the camera position;
[0013] S7. Obtain the coordinates of the camera view - point when the camera looks at point O;
[0014] S8. Adjust the camera field - of - view angle in the GIS map to the view - point position and make it the same as the horizontal value and pitch value of the camera;
[0015] S9. Transparify the background of the GIS map, retain the original entity labels in the GIS map, set the rendering, and remove the atmospheric effect on the earth's surface;
[0016] S10. Overlay the transparent GIS map scene on the camera video image and fuse the GIS map labels with the video image.
[0017] Preferably, the web front - end renders and constructs a GIS three - dimensional scene through the Cesium engine Figure 3 dimensional scene.
[0018] Preferably, the control access of the network camera is based on a third - party SDK.
[0019] Preferably, based on the webrtc protocol, the browser calls webrtc - streamer.js for real - time video playback and display.
[0020] Preferably, the method for obtaining the coordinates of the camera view - point when the camera looks at point O is:
[0021]
[0022] Preferably, the method for transparifying the background of the GIS map, retaining the original entity labels in the GIS map, setting the rendering, and removing the atmospheric effect on the earth's surface is:
[0023] S91. Set the orderIndependentTranslucency property to false to disable OIT, and process the rendering order of the transparent object and the earth's surface;
[0024] S92. Specify the options for the WebGL context by passing a contextOptions object containing WebGL attributes. Specifically, set the alpha option to true to enable the transparency of the color buffer.
[0025] S93. Turn off the skyBox skybox attribute of the viewer object and set its backgroundColor property to transparent.
[0026] Solution 2: An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for generating video tags based on geographic information fusion described in Solution 1.
[0027] Solution 3: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for generating video tags based on geographic information fusion described in Solution 1.
[0028] The beneficial effects of the present invention are as follows: The method for generating video tags based on geographic information fusion in the present invention combines geographic information with a camera to realize video tag generation, effectively solving problems such as high cost and secondary entry caused by the fusion of AR video tags and map videos; at the same time, it solves the limitation that when viewing a long-distance picture affected by camera zoom, the fused video will become slender and narrow, and only a fixed-focus bullet camera can be selected as the hardware for video data. Description of the Drawings
[0029] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0030] Figure 1 It is a flowchart of a method for generating video tags based on geographic information fusion;
[0031] Figure 2 It is a schematic diagram of the control access process of an IP camera based on a third-party SDK;
[0032] Figure 3 It is a schematic diagram of the principle of a method for generating video tags based on geographic information fusion. Detailed Embodiments
[0033] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the following further details the exemplary embodiments of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0034] Embodiment 1: Refer to Figures 1-3 To illustrate this embodiment, a method for generating video tags based on geographic information fusion includes the following steps:
[0035] S1. The web front-end renders and constructs a GIS three-dimensional scene through the Cesium engine; Figure 3
[0036] S2. Based on a third-party SDK, the operation access of network cameras is carried out. Refer to Figure 2 ;
[0037] S3. Based on the webrtc protocol, the browser calls webrtc-streamer.js for real-time video playback and display;
[0038] S4. Calibrate the GIS map and the camera image, and guide the camera to rotate;
[0039] The method for calibrating the GIS map and the camera image and guiding the camera to rotate can be as follows:
[0040] Step a: Configure the high-definition satellite ground image of the monitored area;
[0041] Step b: Add the installation position and installation height of the servo turntable optical monitoring device to the high-definition satellite ground image;
[0042] Step c: Establish a 360-degree equidistant azimuth distance data matrix based on the installation position of the servo turntable optical monitoring device;
[0043] Step d: Calibrate the pitch data of the servo turntable optical monitoring device for each unit of the azimuth distance data matrix to complete the calibration;
[0044] Step e: Click on any position on the map, and calculate the azimuth, pitch, and distance parameters for guiding the servo turntable optical monitoring device through the difference calculation of the matrix to which the position point belongs, and control the servo turntable optical monitoring device to complete the device control.
[0045] S5. Select any point O on the GIS map, guide the camera to rotate to obtain the horizontal direction value, pitch value, and focal length value, and obtain ∠APB, where P is the camera position, A and B are the illumination widths of the perspective camera, and the coordinate data (lon, lat, alt) of point P is known information;
[0046] ∠APB is the angular difference between the camera looking at point B and point A; for example, when looking at B, it is 165°, and when looking at A, it is 45°, then ∠APB = 165° - 45° = 120°
[0047] S6. Obtain the camera field of view angle ∠ACB of the 3D map, where C is the camera position;
[0048] Refer to Figure 3 , where D in the figure is the installation position of the camera and P is the irradiation position of the camera; the camera field of view angle ∠ACB of the 3D map is fixed, which is the fixed viewing angle range of the GIS platform and can be directly obtained through Cesium. The irradiation width of this angle is greater than the irradiation width of the camera. Therefore, when matching the 3D map with the camera screen, the camera view point is placed in the middle of the camera P and the target point O.
[0049] S7. Obtain the coordinates of the camera view point when the camera looks at point O;
[0050]
[0051] S8. Adjust the camera field of view angle in the GIS map to the view point position and make it the same as the horizontal and pitch values of the camera;
[0052] S9. Transparify the background of the GIS map based on the CesiumJS API, retain the original entity labels in the GIS map, set the rendering, and remove the atmospheric effect on the earth's surface;
[0053] When making the earth background transparent, there may be a conflict in the rendering order between the transparent object and the earth's surface; by default, Cesium uses the Order Independent Translucency (OIT) method to handle the rendering order problem of transparent objects to obtain the correct blending effect; however, in some specific scenarios, enabling OIT may cause rendering problems on the earth's surface, such as a conflict in the rendering order when viewing underground structures through the earth's surface; therefore, to solve this problem, setting the orderIndependentTranslucency property to false can disable the OIT technology and perform a conventional rendering order process for the transparent object and the earth's surface. This can ensure that the earth's surface is rendered prior to the transparent object, thus avoiding potential rendering conflicts;
[0054] By passing a contextOptions object containing webgl properties, options for the WebGL context can be specified. Specifically, setting the alpha option to true enables transparency for the color buffer;
[0055] Turn off the skyBox sky box property of the viewer object and set its background color property backgroundColor to transparent;
[0056] S10.Overlay the transparent GIS map scene on the camera video picture and fuse the GIS map labels with the video picture.
[0057] Due to the imaging distortion of the camera picture, the accuracy of taking the center of the picture for the label is high, and the error is larger the closer to the edge. Therefore, values can be taken according to the actual situation.
[0058] Embodiment 2: The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, the steps of the above-mentioned method for generating video labels based on geographic information fusion are implemented.
[0059] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0060] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, applications required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0061] Embodiment 3: Embodiment of a computer-readable storage medium.
[0062] The computer-readable storage medium of the present invention can be any form of storage medium readable by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned method for generating video tags based on geographic information fusion can be implemented.
[0063] The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0064] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field will understand, based on the above description, that other embodiments can be conceived within the scope of the present invention thus described. In addition, it should be noted that the language used in this specification is mainly selected for the purpose of readability and teaching, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and changes are obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. A method for generating video tags based on geographic information fusion, characterized in that, It includes the following steps: S1. Build a 3D scene of the GIS map; S2. Control access to network cameras; S3. Call webrtc - streamer.js for real - time video playback and display; S4. Calibrate the GIS map with the camera image and guide the camera to rotate; S5. Select any point O in the GIS map, guide the camera to rotate to obtain the horizontal value, pitch value and focal length value, and get ∠APB, where P is the camera position, and A and B are the irradiation widths of the perspective camera; S6. Obtain the camera field - of - view angle ∠ACB of the 3D map, where C is the camera position; S7. Obtain the coordinates of the camera view point when the camera looks at point O; S8. Adjust the camera field - of - view angle in the GIS map to the view - point position and make it the same as the horizontal value and pitch value of the camera; S9. Transparify the background of the GIS map, retain the original entity labels in the GIS map, set the rendering, and remove the atmospheric effect on the earth's surface; S10. Overlay the transparent GIS map scene on the camera video image and fuse the GIS map labels with the video image.
2. The method for generating video tags based on geographic information fusion according to claim 1, wherein The web front - end builds a 3D scene of the GIS map through the Cesium engine for rendering.
3. A method for generating video tags based on geographic information fusion according to claim 1, characterized in that, Based on a third - party SDK, control access to network cameras.
4. A method for generating video tags based on geographic information fusion according to claim 1, characterized in that, Based on the webrtc protocol, the browser calls webrtc - streamer.js for real - time video playback and display.
5. A method for generating video tags based on geographic information fusion according to claim 1, wherein The method for obtaining the coordinates of the camera view point when the camera looks at point O is:
6. A method for generating video tags based on geographic information fusion according to claim 1, characterized in that, The method for transparifying the background of the GIS map, retaining the original entity labels in the GIS map, setting the rendering, and removing the atmospheric effect on the earth's surface is: S91. Set the orderIndependentTranslucency property to false to disable OIT, and process the rendering order of the transparent object and the earth's surface; S92. Specify the options of the WebGL context by passing a contextOptions object containing webgl properties; specifically, set the alpha option to true to enable the transparency of the color buffer; S93. Turn off the skyBox sky - box property of the viewer object and set its background color property backgroundColor to transparent.
7. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of a method for generating video labels based on geographic information fusion according to any one of claims 1 - 6.
8. A computer-readable storage medium, characterized in that, A computer program is stored thereon. When the computer program is executed by the processor, it implements a method for generating video labels based on geographic information fusion according to any one of claims 1 - 6.
Citation Information
Patent Citations
A construction method of a three-dimensional GIS dynamic model
CN109598794A
Stadium three-dimensional video fusion visual monitoring system and method
CN111586351A
Fusion display method and device of aerial video on digital earth
CN114494563A
Method and device for displaying panoramic videos
US20170142389A1
A real-time generation method for 360-degree VR panoramic graphic image and video
US20200128178A1