Information processing method and system based on spatialization of unmanned aerial vehicle video
By embedding camera optical parameters, pose information, and sensor type identifiers into UAV videos, a dynamic spatial mapping model is constructed, solving the problems of poor adaptability and insufficient geographic data synchronization in traditional UAV video processing methods. This achieves accurate correspondence and real-time synchronization between video and geographic data, improving application compatibility and the accuracy of geographic information.
Patent Information
- Application Number
- CN202511518632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Traditional drone video processing methods fail to effectively integrate camera optical parameters, camera pose information, and sensor type identification, resulting in poor adaptability of videos in different scenarios, making it difficult to meet diverse needs. Furthermore, the synchronous update of geographic data and video lacks real-time performance and accuracy.
The camera's optical parameters, camera pose information, and sensor type identifier are synchronously encoded into the non-display data segment of the drone video frame structure, forming an encoded video with multi-dimensional spatial information. Through scene-specific transcoding and adaptation processing, a dynamic spatial mapping model is constructed to achieve accurate association between the video and the geographical area. Combined with the ground feature recognition model, geographical data with ground feature attributes and recognition confidence are generated, and bidirectional dynamic synchronization is achieved through multi-link network transmission.
It improves the compatibility and availability of drone video in different application scenarios, ensures the consistency and real-time nature of geographic data and video content, meets diverse business needs, and enhances the quality and reliability of geographic information.
Smart Images

Figure CN121000900B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) video processing technology, and more specifically, to an information processing method and system based on UAV video spatialization. Background Technology
[0002] With the rapid development of drone technology, the application scenarios of drone video are becoming increasingly widespread, covering multiple fields such as geographic surveying, environmental monitoring, disaster assessment, and security monitoring. However, traditional drone video processing methods have many limitations.
[0003] Currently, most drone video processing focuses only on the visual content of the video itself, severely underutilizing the spatial information contained within it. Key information such as camera optical parameters, camera pose information, and sensor type identification are often treated in isolation in traditional processing methods, without being deeply integrated with or effectively utilized by the video frames.
[0004] In terms of video transmission and application, existing processing methods are insufficient to meet the diverse needs for video formats and characteristics in different scenarios. Different decoding ends have different requirements for video formats, and the image characteristics corresponding to different sensor types vary significantly. Traditional methods cannot flexibly adjust video encoding algorithms and data encapsulation structures according to these differences, resulting in poor adaptability of videos in different scenarios.
[0005] Furthermore, in terms of geographic data generation and application, traditional methods struggle to accurately correlate ground features in video with geographic extent. They fail to construct effective spatial mapping models by combining camera optical parameters and pose information, resulting in generated geographic data lacking accurate geographic benchmarks and failing to meet the precision requirements of practical applications. Moreover, regarding the synchronous updating of geographic data and video, traditional methods lack effective two-way dynamic synchronization and conflict resolution mechanisms, making it impossible to handle modifications and updates to geographic data in real time and accurately, thus affecting the consistency and usability of geographic data and video. Summary of the Invention
[0006] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide an information processing method based on UAV video spatialization, the method comprising:
[0007] Collect camera optical parameters, camera pose information and sensor type identifiers, and synchronously encode the camera optical parameters, camera pose information and sensor type identifiers into the non-display data segment of the frame structure of the UAV video to form an encoded video with multi-dimensional spatial information. Perform streaming encapsulation processing on the encoded video with multi-dimensional spatial information, add inter-frame association identifier fields, and form an encoded video stream.
[0008] The encoded video stream is subjected to scene-specific transcoding adaptation processing. Based on the format requirements of the subsequent decoding end and the image characteristics corresponding to the sensor type identifier, the quantization parameters and data encapsulation structure of the video encoding algorithm are adjusted to obtain an adapted video stream for different scenes.
[0009] The adapted video stream is subjected to layered decoding and splitting processing to separate the frame sequence of the drone live video, the camera optical parameters, camera pose information and sensor type identifier embedded in the frame sequence, and the image correction model corresponding to the camera optical parameters, camera pose information and sensor type identifier. A dynamic spatial mapping model is constructed by combining the image correction model corresponding to the camera optical parameters, camera pose information and sensor type identifier. The geographic range of each video frame in the frame sequence under different geographic references is calculated by the dynamic spatial mapping model, and the correlation between the video frame and the geographic range of multiple references is established to obtain the live video with multiple reference geographic range identifiers.
[0010] Based on live video with multi-reference geographic range identifiers, the system uses intelligent delineation operation combined with the ground feature recognition model corresponding to the sensor type identifier to perform point, line or area trajectory recording along the outline of ground features in the video frame. Based on the coordinates of the trajectory recording and the target geographic reference range corresponding to the video frame, coordinate transformation is performed to generate geographic data with ground feature attributes and recognition confidence.
[0011] Geographic data with feature attributes and identification confidence levels are transmitted to web maps, mobile terminal maps, and display interfaces via a multi-link network. At the same time, the system receives geographic data modification trajectories from web maps and mobile terminal maps. Based on the geographic data modification trajectories, the feature identification confidence levels corresponding to the modification trajectories, and the coordinate information of the corresponding features in the live video with multi-reference geographic range identifiers, the system performs hierarchical update processing to achieve bidirectional dynamic synchronization and conflict resolution between geographic data and live video.
[0012] Furthermore, embodiments of the present invention also provide an information processing system based on UAV video spatialization, characterized in that it includes:
[0013] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described information processing method based on UAV video spatialization by executing the machine-executable instructions.
[0014] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described information processing method based on UAV video spatialization.
[0015] Based on the above, by synchronously encoding the camera's optical parameters, camera pose information, and sensor type identifier into the non-display data segment of the drone video frame structure, an encoded video with multi-dimensional spatial information is formed. This video is then further processed into an encoded video stream, achieving deep fusion of spatial information and video. The encoded video stream undergoes scene-specific transcoding adaptation processing, which can flexibly adjust the quantization parameters and data encapsulation structure of the video encoding algorithm according to the format requirements of the subsequent decoding end and the image characteristics corresponding to the sensor type identifier. This results in an adapted video stream suitable for different scenarios, greatly improving the compatibility and usability of the video in different application scenarios and meeting diverse business needs. By performing layered decoding and splitting processing on the adapted video stream, and constructing a dynamic spatial mapping model in conjunction with the corresponding image correction model, the geographical range of each video frame in the frame sequence under different geographic references is calculated. The association between the video frame and the multi-reference geographic range is established, resulting in a live video with multi-reference geographic range identifiers. This achieves a precise correspondence between the video image and the geographic space. Based on the live video with multi-reference geographic range identifiers, geographic data with geographic feature attributes and recognition confidence is generated through intelligent delineation operations combined with a geographic feature recognition model. This geographic data is then transmitted to relevant maps and display interfaces through a multi-link network. Simultaneously, feedback on geographic data modification trajectories is received and layered update processing is performed. This achieves bidirectional dynamic synchronization and conflict resolution between geographic data and live video, ensuring the consistency and real-time nature of geographic data and video content, and improving the quality and reliability of geographic information. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the execution flow of the information processing method based on UAV video spatialization provided in an embodiment of the present invention.
[0017] Figure 2 This is a schematic diagram of exemplary hardware and software components of an information processing system based on UAV video spatialization provided in an embodiment of the present invention. Detailed Implementation
[0018] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an embodiment of the information processing method based on UAV video spatialization provided by the present invention. The following is a detailed description of the information processing method based on UAV video spatialization.
[0019] Step S110: Collect camera optical parameters, camera pose information and sensor type identifier, and synchronously encode the camera optical parameters, camera pose information and sensor type identifier into the non-display data segment of the frame structure of the UAV video to form an encoded video with multi-dimensional spatial information. Perform streaming encapsulation processing on the encoded video with multi-dimensional spatial information, add an inter-frame association identifier field, and form an encoded video stream.
[0020] This embodiment uses a drone operation scenario for conducting a natural resource survey of a certain area as an example for explanation. In this scenario, the drone is equipped with corresponding cameras and sensor devices to collect information on ground features within the area.
[0021] Step S111: Obtain the camera optical parameters output in real time through the data acquisition link. The camera optical parameters include focal length parameters, aperture parameters, sensor size parameters, and distortion correction coefficients.
[0022] In the aforementioned natural resource survey scenario, the data acquisition link is the crucial channel connecting the UAV camera and the data processing module. During operation, the camera outputs its current optical parameters to the data processing module in real time. Focal length reflects the camera lens's focusing capability; different focal lengths affect the size of ground features and the field of view in video frames. Aperture determines the amount of light entering the lens, thus affecting the brightness and depth of field of the video frame. Sensor size parameters, along with focal length, influence the field of view. Distortion correction coefficients are used to correct image distortion caused by lens optical characteristics, ensuring the accuracy of subsequent ground feature contour recognition. The data processing module continuously receives and stores these parameters through this link.
[0023] Step S112: Collect camera pose information through the attitude perception link. The camera pose information includes latitude and longitude parameters, altitude parameters, pitch angle parameters, yaw angle parameters, and roll angle parameters.
[0024] In natural resource survey scenarios, the attitude of the camera on a drone changes continuously during flight. The attitude perception link integrates multiple sensors, including GPS positioning devices, gyroscopes, and accelerometers, to monitor the camera's position and attitude in real time. Latitude and longitude parameters determine the camera's specific location on the Earth's surface; altitude parameters reflect the camera's height above the ground; pitch angle parameters indicate the camera's tilt angle relative to the horizontal plane; yaw angle parameters indicate the camera's horizontal rotation direction; and roll angle parameters reflect the degree of rotation of the camera around its own longitudinal axis. The combination of these parameters accurately describes the camera's attitude in three-dimensional space and is the core input data for constructing a dynamic spatial mapping model. The data processing module acquires these parameters in real time through the attitude perception link and synchronizes them with the corresponding video frames.
[0025] Step S113: Obtain the sensor type identifier through the sensor type detection link. The sensor type identifier includes a visible light sensor identifier or a thermal infrared sensor identifier.
[0026] In natural resource surveys, different types of sensors may be used depending on the specific survey requirements. The sensor type detection link identifies the sensors carried by the UAV, determines their type, and generates corresponding identifiers. If a visible light sensor is used, its identifier will be recorded as a visible light sensor identifier; this type of sensor is suitable for collecting visible light information such as the color and texture of ground features. If a thermal infrared sensor is used, a thermal infrared sensor identifier will be generated; this type of sensor can capture the temperature distribution characteristics of ground features and is often used to identify special ground features such as heat sources and water bodies. The sensor type identifier will serve as an important basis for subsequent selection of ground feature recognition models. After the data processing module obtains the identifier through this link, it will associate and store it with the corresponding video frame and optical and pose parameters.
[0027] Step S114: Determine the intra-frame embedding rules. The intra-frame embedding rules specify the starting position of the bytes, data length, parameter arrangement order, and check bit generation method of the camera optical parameters, camera pose information, and sensor type identifier in the non-display data segment of the UAV video frame, so that the camera optical parameters, camera pose information, sensor type identifier and the timestamp of the video frame form a one-to-one correspondence.
[0028] To embed the aforementioned multidimensional spatial information into the video frame structure in an orderly manner, intra-frame embedding rules need to be determined in advance. In natural resource survey scenarios, the non-display data segment of the video frame is the ideal location to store this additional information without affecting the normal display of the video. The intra-frame embedding rules first need to clarify the starting byte position of each parameter in the non-display data segment to avoid storage conflicts between different parameters. For example, the beginning of the non-display data segment can be allocated to the sensor type identifier, as its data volume is relatively small; then, the storage areas for camera optical parameters and camera pose information are allocated sequentially. Regarding data length, a fixed byte length needs to be allocated to each parameter based on its value range and accuracy requirements to ensure complete storage of parameter information. The parameter arrangement order is set according to the priority and relevance of data processing. For example, parameters directly related to location, such as latitude and longitude parameters and altitude parameters in the pose information, are arranged first for easy subsequent rapid extraction. The check bit generation method uses the CRC32 algorithm to calculate all embedded parameter data, generate a check bit, and store it in a designated location for subsequent data integrity verification. By setting these rules, it can be ensured that each set of optical parameters, pose information, and sensor type identifier can accurately correspond to a specific video frame timestamp.
[0029] Step S115: According to the intra-frame embedding rules, the camera optical parameters, camera pose information and sensor type identifier are written frame by frame into the non-display data segment of the drone video frame structure. A CRC32 check bit is generated for the data written in each frame. After writing is completed, an encoded video with multi-dimensional spatial information is formed.
[0030] During natural resource surveys, drones continuously capture video. The data processing module embeds data into each video frame according to defined intra-frame embedding rules. For each frame, the data processing module reads the camera's optical parameters, camera pose information, and sensor type identifier from the storage unit. Then, based on the byte start position and data length specified in the rules, it sequentially writes these parameters into the non-display data segment of the video frame. After writing, a CRC32 checksum is calculated for all written data, and the generated checksum is also written to a designated position in the non-display data segment. For example, for a video frame, the sensor type identifier is written first, followed by optical parameters such as focal length and aperture, then pose parameters such as latitude, longitude, and altitude, and finally the CRC32 checksum is appended. This process is repeated frame by frame, ensuring that each video frame carries corresponding multi-dimensional spatial information, ultimately forming an encoded video with multi-dimensional spatial information.
[0031] Step S116: Perform inter-frame association processing on the encoded video with multi-dimensional spatial information, extract the camera pose information difference between adjacent video frames, and add the camera pose information difference as an inter-frame association identifier field to the non-display data segment of the current video frame.
[0032] In natural resource survey scenarios, the camera pose changes between adjacent video frames during UAV flight are typically continuous. Extracting pose information differences can reflect these trends. During inter-frame correlation processing, the data processing module sequentially reads adjacent frames of the encoded video containing multi-dimensional spatial information. For the current frame and the previous frame, their camera pose information is extracted, including latitude and longitude parameters, altitude parameters, pitch angle parameters, yaw angle parameters, and roll angle parameters. The differences between these parameters are then calculated; for example, the longitude difference is obtained by subtracting the longitude of the previous frame from the current frame's longitude, and the latitude difference is obtained by subtracting the latitude of the previous frame from the current frame's latitude, and so on, to obtain the differences for all pose parameters. These differences are combined into an inter-frame correlation identifier field, which characterizes the position and attitude changes of the camera between adjacent frames. The data processing module adds this field to the non-display data segment of the current video frame and stores it along with other previously embedded information. By adding the inter-frame correlation identifier field, subsequent processing can quickly understand the spatial relationship between adjacent frames, facilitating smooth adjustments to the geographical scope and optimization of dynamic spatial mapping models.
[0033] Step S117: Perform streaming segmentation on the encoded video with added inter-frame association identifier field, divide the encoded video into a continuous sequence of data packets according to the preset data packet size, and add a sequence number identifier, timestamp identifier and sensor type identifier copy to each data packet.
[0034] To facilitate network transmission and subsequent processing of video data, the encoded video, after adding an inter-frame correlation identifier field, needs to be stream-segmented. In natural resource survey scenarios, the preset data packet size is typically determined based on network bandwidth and data processing efficiency. The data processing module segments the encoded video into consecutive data packets according to this preset size. During the segmentation process, it is necessary to ensure the integrity of each data packet to avoid video frame data being split into different data packets. For each segmented data packet, the data processing module adds a sequence number identifier to indicate the order of the data packet in the entire video stream; adds a timestamp identifier to record the acquisition time of the corresponding video frame; and adds a copy of the sensor type identifier so that the receiving end can quickly understand the sensor type even without parsing the complete data packet. The above identifier information is added to the header of the data packet to facilitate the receiving end's reassembly and verification of the data packet.
[0035] Step S118: Perform transmission protocol encapsulation on the data packet sequence with sequence number identifier, timestamp identifier and sensor type identifier copy, add data fragment identifier, transmission window size field and retransmission threshold field to form an encoded video stream that can be directly transmitted.
[0036] After packet segmentation and identifier addition, the packet sequence needs to be encapsulated using a transmission protocol to meet network transmission requirements. In natural resource survey scenarios, protocols suitable for real-time data transmission are typically used. The data processing module adds a transmission protocol header to each packet, which includes a data fragmentation identifier to indicate whether the current packet is a complete video frame or fragmented data; a transmission window size field to control network traffic, dynamically adjusting the sending window size according to network conditions to ensure data transmission stability; and a retransmission threshold field to set the maximum number of retransmissions allowed after a packet transmission failure. By adding these fields, the packet sequence conforms to the transmission protocol specifications and can be reliably transmitted over the network. The packet sequence encapsulated using the transmission protocol forms a directly transmittable encoded video stream. This encoded video stream contains video frame data and all necessary spatial and transmission control information, and can be sent over the network to a remote data processing center or display terminal.
[0037] Step S120: Perform scene-specific transcoding adaptation processing on the encoded video stream. Based on the format requirements of the subsequent decoding end and the image characteristics corresponding to the sensor type identifier, adjust the quantization parameters and data encapsulation structure of the video encoding algorithm to obtain an adapted video stream for different scenes.
[0038] In natural resource survey scenarios, encoded video streams need to be transmitted to different decoding endpoints, such as web map display systems and mobile terminal map applications. These decoding endpoints have different requirements for video format, resolution, and bitrate. Furthermore, images acquired by different types of sensors have different characteristics; visible light images and thermal infrared images also require different treatment in encoding processing. Therefore, it is necessary to perform scene-specific transcoding and adaptation processing on the encoded video streams to ensure that the decoding endpoints can decode and display the video correctly while fully preserving the key information in the images.
[0039] Step S121: Extract the transmission window size field and sensor type identifier copy from the encoded video stream, and parse the format requirement information of the subsequent decoding end. The format requirement information includes the supported encoding algorithm type, maximum resolution limit, frame rate range and acceptable bit rate fluctuation range.
[0040] Before transcoding adaptation begins, the data processing module first extracts the transmission window size field and a copy of the sensor type identifier from the header of the encoded video stream's data packets. The transmission window size field reflects the current network transmission status and can provide a reference for bitrate adjustment during transcoding; the sensor type identifier copy clarifies the sensor type of the video stream, such as visible light or thermal infrared. Next, the data processing module communicates with the subsequent decoding end to parse its format requirements. The decoding end, based on its hardware performance, display capabilities, and network conditions, provides information such as supported encoding algorithm types (e.g., H.264, H.265), maximum resolution limits (the highest image resolution the decoding end can process), frame rate range (indicating the required smoothness of video playback), and acceptable bitrate fluctuation range to avoid excessive bitrate fluctuations causing decoding stuttering. This information is crucial for transcoding adaptation; the data processing module needs to combine this information with the image characteristics corresponding to the sensor type to formulate an appropriate transcoding strategy.
[0041] Step S122: Determine the image characteristics based on the sensor type identifier copy. If it is a visible light sensor identifier, determine that the image characteristics are high dynamic range characteristics; if it is a thermal infrared sensor identifier, determine that the image characteristics are temperature gradient sensitive characteristics.
[0042] Based on the extracted sensor type identifier copy, the data processing module determines the image characteristics of the video stream. In natural resource surveys, if the sensor type is a visible light sensor, the acquired images have high dynamic range characteristics, meaning they simultaneously contain bright and dark areas with rich detail. These details need to be preserved as much as possible during transcoding to avoid overexposure in bright areas or loss of detail in dark areas. If the sensor type is a thermal infrared sensor, its image characteristics are temperature gradient sensitive. The pixel values in the image reflect the temperature information of ground features, and subtle temperature changes are manifested as gradual changes in grayscale in the image. Transcoding processing needs to focus on protecting this temperature gradient information to avoid distortion of temperature information due to encoding compression, which would affect subsequent ground feature identification and analysis.
[0043] Step S123: Adjust the quantization parameters of the video coding algorithm based on image characteristics. For images with high dynamic range characteristics, reduce the quantization step size in bright areas to preserve details; for images with temperature gradient sensitivity characteristics, reduce the quantization step size in areas with abrupt temperature changes to preserve gradient information.
[0044] Quantization parameters are key factors affecting video coding quality and bitrate. Smaller quantization steps result in higher coding quality, but also higher bitrate. Based on the determined image characteristics, the data processing module adjusts the quantization parameters of the video coding algorithm. For visible light images with high dynamic range, the data processing module identifies bright areas in the image through image analysis algorithms and then reduces the quantization step size for these areas. This allows the encoder to use a finer quantization method when encoding bright areas, thus preserving detailed information such as cloud textures and highlights on buildings. For thermal infrared images sensitive to temperature gradients, the data processing module detects areas of abrupt temperature changes in the image through temperature analysis algorithms, such as heat source boundaries and the boundary between water and land. It reduces the quantization step size for these areas to ensure that temperature gradient information is not lost during encoding, enabling accurate identification of land cover types using thermal infrared ground cover temperature feature models.
[0045] Step S124: Select the target encoding algorithm according to the encoding algorithm type in the format requirement information. If the decoding end supports multiple encoding algorithms, the target encoding algorithm with the highest compatibility with the original encoding video stream algorithm and the best fit for image characteristics should be selected first.
[0046] The format requirements of the decoding end include the types of encoding algorithms it supports. The data processing module needs to select a suitable target encoding algorithm from these. In a natural resource survey scenario, if the decoding end only supports one encoding algorithm, that algorithm is directly selected for transcoding. If the decoding end supports multiple encoding algorithms, the data processing module needs to comprehensively consider the compatibility between the original and target algorithms in the encoded video stream, as well as the adaptability of the target algorithm to the current image characteristics. For example, if the original encoding algorithm is H.264, and the decoding end supports both H.264 and H.265, and the current image is a high dynamic range visible light image, H.265 has advantages in compression efficiency and detail preservation, and also has good compatibility with H.264. Therefore, H.265 is preferred as the target encoding algorithm. By selecting a suitable target encoding algorithm, image characteristics and video quality can be preserved to the maximum extent while meeting the format requirements of the decoding end.
[0047] Step S125: Perform decapsulation processing on the data packet sequence of the encoded video stream to separate the original frame unit sequence with the inter-frame association identifier field.
[0048] To perform transcoding, the data packet sequence of the encoded video stream needs to be decapsulated to restore the original video frame data. The data processing module reassembles the data packet sequence based on the information in the transmission protocol header, removing transmission control information such as data fragmentation identifiers and transmission window size fields, resulting in a sequence of original frame units with inter-frame association identifier fields. Each original frame unit contains the pixel data of the video frame, previously embedded camera optical parameters, camera pose information, sensor type identifier, and inter-frame association identifier fields. During decapsulation, the data processing module verifies the sequence number and timestamp of the data packets to ensure the correct order and integrity of the original frame unit sequence.
[0049] Step S126: Re-encode the video pixel data in the original frame unit sequence according to the encoding rules of the target encoding algorithm and the adjusted quantization parameters, while retaining the camera optical parameters, camera pose information, sensor type identifier and inter-frame association identifier fields in the original frame unit.
[0050] After obtaining the original frame unit sequence, the data processing module re-encodes the video pixel data in the original frame unit sequence according to the encoding rules of the selected target encoding algorithm and in conjunction with the adjusted quantization parameters. During the encoding process, the encoder performs transformation, quantization, and entropy encoding operations on the video pixel data to generate video stream data that conforms to the format of the target encoding algorithm. At the same time, the data processing module ensures that the camera optical parameters, camera pose information, sensor type identifier, and inter-frame association identifier fields embedded in the original frame units are not modified or lost, and retains the above information as is in the re-encoded frame units for subsequent layered decoding and splitting processing and dynamic spatial mapping model construction.
[0051] Step S127: Reconstruct the data packet structure according to the format requirements. The new data packet structure includes an image feature identifier field, a quantization parameter configuration field, an original frame unit retention field, and a decoding priority field. Encapsulate the re-encoded video data with the retained camera optical parameters, camera pose information, sensor type identifier, and inter-frame association identifier fields according to the new data packet structure, and add an adaptation identifier field to the new data packet.
[0052] After re-encoding, the data packet structure needs to be reconstructed according to the format requirements of the decoding end. The new data packet structure retains the original information while adding new fields to adapt to transcoding needs. The image characteristic identifier field indicates the image characteristics of the video frame corresponding to the current data packet, such as high dynamic range or temperature gradient sensitivity; the quantization parameter configuration field records the quantization parameter settings used during the encoding of the current data packet, facilitating decoding optimization at the decoding end; the original frame unit retention field stores various parameter information retained from the original frame units; and the decoding priority field sets the decoding order according to the importance of the video frames, such as key frames having higher decoding priority than non-key frames. The data processing module encapsulates the re-encoded video data and the retained parameter information according to the new data packet structure and adds an adaptation identifier field to each new data packet to indicate that the data packet has undergone transcoding adaptation processing, enabling the receiving end to identify and process it.
[0053] Step S128: Perform sequence verification and bitrate fluctuation detection on the re-encapsulated data packets to ensure that the sequence number of the data packets is consistent with the sequence number of the original encoded video stream, and that the bitrate fluctuation is controlled within an acceptable range. After the verification and detection are passed, an adapted video stream suitable for different scenarios is formed.
[0054] To ensure that the transcoded data packets can correctly reconstruct the video sequence and meet the bitrate requirements of the decoding end, the data processing module performs sequence verification and bitrate fluctuation detection on the re-encapsulated data packets. Sequence verification compares the sequence number of the data packets with the sequence number of the original encoded video stream to ensure that the order of the data packets is not disordered. Bitrate fluctuation detection statistically analyzes the bitrate changes of the data packets over a certain period of time and compares them with the acceptable bitrate fluctuation range in the format requirements information. If the bitrate fluctuation exceeds the range, the data processing module dynamically adjusts the encoding parameters, such as the quantization step size and frame rate, to control the bitrate fluctuation within an acceptable range. After passing sequence verification and bitrate fluctuation detection, the re-encapsulated data packet sequence forms an adapted video stream suitable for different scenarios. This adapted video stream can meet the format requirements of different decoding ends and retains key image characteristics and spatial information.
[0055] Step S130: Perform layered decoding and splitting processing on the adapted video stream to separate the frame sequence of the drone live video, the camera optical parameters embedded in the frame sequence, the camera pose information and sensor type identifier, and combine the image correction model corresponding to the camera optical parameters, camera pose information and sensor type identifier to construct a dynamic spatial mapping model. Calculate the geographical range of each video frame in the frame sequence under different geographical references through the dynamic spatial mapping model, establish the association between the video frame and the multi-reference geographical range, and obtain the live video with multi-reference geographical range identifiers.
[0056] In natural resource survey scenarios, after the adapted video stream is transmitted to the data processing center, it needs to be processed by layered decoding and splitting to extract video frame data and related spatial information, and to build a dynamic spatial mapping model to determine the geographical range corresponding to each video frame.
[0057] Step S131: Perform unpacking processing on the adapted video stream, remove the adaptation identifier field, transmission control field and decoding priority field, and obtain the adapted frame unit sequence with image characteristic identifier field.
[0058] After receiving the adapted video stream, the data processing center first unpacks it. Each data packet in the adapted video stream contains header information such as an adaptation identifier field, a transmission control field, and a decoding priority field. This information plays a control role during transmission and adaptation processing and is no longer needed in subsequent decoding and spatial mapping processing. The unpacking module of the data processing center removes these header fields according to the data packet structure, extracting a sequence of adapted frame units containing video data and image feature identifier fields. Each adapted frame unit corresponds to one frame of video data in the original video stream, and the image feature identifier field indicates the image characteristics of that video frame.
[0059] Step S132: Perform layered decoding operation on the adapted frame unit sequence. The first layer of decoding separates the video pixel data and non-display data segments. The second layer of decoding extracts the camera optical parameters, camera pose information, sensor type identifier and inter-frame association identifier fields from the non-display data segments to obtain the frame sequence of the drone live video and the camera optical parameters, camera pose information and sensor type identifier corresponding to each frame sequence.
[0060] The layered decoding operation is performed in two layers. The first layer of decoding is completed by the video decoder, whose main task is to decode the video data in the adapted frame unit sequence, separating the original video pixel data and non-display data segments. The video pixel data is the basis for subsequent image display and ground feature recognition, while the non-display data segments contain various spatial parameter information embedded earlier. The second layer of decoding is performed by a dedicated data parsing module. This module parses the non-display data segments and extracts camera optical parameters (focal length, aperture, sensor size, and distortion correction coefficients), camera pose information (latitude and longitude, altitude, pitch, yaw, and roll angles), sensor type identifiers (visible light sensor identifier or thermal infrared sensor identifier), and inter-frame association identifier fields according to preset intra-frame embedding rules. Through the layered decoding operation, the data processing center obtains the frame sequence of the UAV live video and the spatial parameter information corresponding to each video frame.
[0061] Step S133: Call the corresponding image correction model according to the sensor type identifier. If it is a visible light sensor, call the atmospheric scattering correction model; if it is a thermal infrared sensor, call the temperature noise correction model.
[0062] The sensor type identifier determines which image correction model should be used to eliminate environmental interference. In natural resource survey scenarios, visible light images are easily affected by atmospheric scattering, leading to decreased image contrast and color distortion; thermal infrared images may be affected by temperature noise, impacting the accuracy of temperature measurements. The data processing center uses the extracted sensor type identifier to call the corresponding image correction model. When the sensor type identifier is a visible light sensor, the atmospheric scattering correction model is used. This model analyzes the atmospheric optical characteristics of the image to eliminate the influence of atmospheric scattering and restore the true color and texture of ground features. When the sensor type identifier is a thermal infrared sensor, the temperature noise correction model is used. This model removes temperature noise from the image through noise filtering algorithms, smooths the temperature distribution curve, and improves the reliability of temperature data.
[0063] Step S134: Input each video frame in the video frame sequence into the corresponding image correction model, perform image preprocessing, eliminate the influence of environmental interference factors on pixel data, and obtain the corrected video frame sequence.
[0064] After receiving video frames, the image correction model performs image preprocessing. For visible light video frames input to the atmospheric scattering correction model, the model first analyzes the image's brightness distribution and color channel information. Then, it calculates the scattering coefficient based on the atmospheric scattering model, eliminates the influence of atmospheric scattering through an inversion algorithm, and adjusts the image's contrast and color balance to make the details of ground features more clearly visible. For thermal infrared video frames input to the temperature noise correction model, the model uses an adaptive filtering algorithm to detect and filter temperature noise in the image while retaining information about areas of sudden temperature changes, avoiding excessive smoothing that could lead to the loss of temperature gradient information. After image preprocessing, environmental interference factors in the video frame sequence are effectively eliminated, resulting in a corrected video frame sequence that more accurately reflects the actual characteristics of ground features.
[0065] Step S135: Extract the pixel size information of each video frame in the corrected video frame sequence, and combine it with the focal length parameter, sensor size parameter and distortion correction coefficient in the camera optical parameters to calculate the horizontal field of view, vertical field of view and distortion correction parameters of the video frame.
[0066] Pixel size information, namely the pixel width and pixel height of the video frame, is the fundamental data for calculating the field of view. The data processing center extracts the pixel size information of each corrected video frame and combines it with the focal length parameter, sensor size parameter, and distortion correction coefficient from the camera's optical parameters to calculate the field of view and distortion correction parameters. The horizontal and vertical field of view are calculated based on the focal length and sensor size, derived through geometric optics principles, and they determine the horizontal and vertical spatial range that the video frame can cover. The distortion correction parameters are determined based on the distortion correction coefficients and are used to further correct distortion at the edges of the video frame, ensuring a linear relationship between pixel coordinates and actual spatial coordinates.
[0067] Step S136: Combine the latitude and longitude parameters, altitude parameters, pitch angle parameters, yaw angle parameters, and roll angle parameters in the camera pose information to determine the shooting origin coordinates, shooting direction vector, and attitude rotation matrix of the video frame.
[0068] The latitude, longitude, and altitude parameters in the camera pose information jointly determine the origin coordinates of the video frame, i.e., the camera's three-dimensional position on the Earth's surface. The pitch, yaw, and roll angle parameters describe the camera's attitude. Through three-dimensional coordinate transformation, these angle parameters are converted into a shooting direction vector and an attitude rotation matrix. The shooting direction vector indicates the direction of the camera's optical axis, and the attitude rotation matrix is used to convert the coordinates in the camera coordinate system to coordinates in the world coordinate system. Determining these parameters allows the data processing center to accurately describe the camera's position and attitude in three-dimensional space.
[0069] Step S137: Using the origin coordinates as the reference, and with the horizontal field of view, vertical field of view, distortion correction parameters, shooting direction vector, and attitude rotation matrix as constraints, construct a dynamic spatial mapping model. The dynamic spatial mapping model includes dual transformation formulas between pixel coordinates and the geodetic coordinate system and the local coordinate system.
[0070] The dynamic spatial mapping model is the core model connecting video frame pixel coordinates with geographic coordinates. The data processing center uses the shooting origin coordinates as the starting point of the three-dimensional space, employs the horizontal and vertical field of view angles as spatial constraints, distortion correction parameters as the basis for coordinate correction, and the shooting direction vector and attitude rotation matrix as direction and attitude constraints to construct the dynamic spatial mapping model. This model achieves bidirectional conversion from pixel coordinates to geographic coordinates by establishing mathematical transformation relationships between pixel coordinates and the geodetic and local coordinate systems. The dual transformation formulas describe how pixel coordinates are converted to coordinates in the geodetic coordinate system and how they are converted to coordinates in the local coordinate system. These formulas consider the differences in camera pose, optical characteristics, and geographic reference, ensuring the accuracy of the conversion.
[0071] Step S138: Input the pixel coordinate matrix of each corrected video frame into the dynamic spatial mapping model, and calculate the geographic coordinates of the four corner points of the video frame in the geodetic coordinate system and the geographic coordinates in the local coordinate system using the double transformation formula.
[0072] The pixel coordinates of the four corner points (top left, top right, bottom left, and bottom right) of the video frame are known: (0,0), (pixel width - 1, 0), (0, pixel height - 1), and (pixel width - 1, pixel height - 1), respectively. The data processing center inputs this pixel coordinate matrix into a dynamic spatial mapping model. The model, using a double transformation formula, calculates the geographic coordinates (longitude and latitude) of each corner point in the geodetic coordinate system and its geographic coordinates (Cartesian coordinates) in the local coordinate system. The geographic coordinates of the corner points roughly reflect the geographic coverage of the video frame and are key data points for determining the geographic range of the video frame.
[0073] Step S139: Determine the geodetic reference geographical range corresponding to the video frame using the geographical coordinates of the four corner points in the geodetic coordinate system as the boundary, and determine the local reference geographical range corresponding to the video frame using the geographical coordinates of the four corner points in the local coordinate system as the boundary.
[0074] Based on the calculated geographic coordinates of the four corner points in the geodetic coordinate system, the data processing center determines the geodetic reference geographic range. This geodetic reference geographic range is typically represented by minimum longitude, maximum longitude, minimum latitude, and maximum latitude, forming a rectangular area that covers the geographic coordinates of all pixels in the video frame in the geodetic coordinate system. Similarly, based on the geographic coordinates of the four corner points in the local coordinate system, the local reference geographic range is determined, represented by minimum X-coordinate, maximum X-coordinate, minimum Y-coordinate, and maximum Y-coordinate. The geodetic reference geographic range is suitable for large-scale geographic positioning and map stitching, while the local reference geographic range is suitable for small-scale fine-tuning and engineering applications.
[0075] Step S1310: Perform accuracy verification on the geodetic reference geographic range and local reference geographic range of each video frame. Compare the overlap between the calculated geodetic reference geographic range and the geodetic reference geographic range of adjacent video frames. Compare the overlap between the local reference geographic range and the local reference geographic range of adjacent video frames. If the overlap under any reference does not meet the preset association requirements, adjust the transformation formula parameters in the dynamic spatial mapping model and recalculate the geographic range of the corresponding reference.
[0076] To ensure the accuracy and continuity of the geographical extent of video frames, the data processing center performs accuracy verification on the geographical extent of each video frame. Overlap comparison is a crucial verification method. It involves calculating the ratio of the area of the overlapping region between the current video frame and its adjacent frames to the sum of the areas of the two geographical extents to determine if the overlap meets the preset association requirements. These preset association requirements are determined based on factors such as the UAV's flight speed and video frame rate to ensure that the geographical extent of adjacent video frames continuously covers the survey area, avoiding missed detections or excessive duplicate coverage. If the overlap does not meet the requirements, it indicates a potential deviation in the transformation formula parameters of the dynamic spatial mapping model. The data processing center then adjusts the parameters in the model, such as the elements of the attitude rotation matrix and distortion correction parameters, and recalculates the corresponding baseline geographical extent until the overlap meets the requirements.
[0077] Step S1311: Write the verified geodetic reference geographic range information and local reference geographic range information into the attribute fields of the corresponding video frame to form a frame unit with multiple reference geographic range identifiers. Arrange the frame units with multiple reference geographic range identifiers in chronological order, while retaining the adjacent frame association relationship corresponding to the inter-frame association identifier field to obtain the live video with multiple reference geographic range identifiers.
[0078] The geodetic datum and local datum geographic extent information, verified for accuracy, are written into the attribute fields of the corresponding video frames, forming frame units with multi-datum geographic extent identifiers. These attribute fields are stored along with the video frame data for easy subsequent querying and use. The data processing center arranges the frame units with multi-datum geographic extent identifiers in chronological order to restore the time sequence of the video stream. Simultaneously, the correlation between adjacent frames corresponding to the inter-frame association identifier field is preserved, enabling rapid acquisition of geographic extent information for adjacent frames during playback or processing of live video, achieving a smooth transition of geographic extent. The resulting live video with multi-datum geographic extent identifiers not only contains corrected video frame data but also carries accurate multi-datum geographic extent information.
[0079] Step S1312: Extract the distortion correction coefficients from the camera's optical parameters, combine them with the correction parameters output by the image correction model corresponding to the sensor type identifier, establish a pixel coordinate pre-correction formula, perform pre-correction processing on the original pixel coordinates of the video frame, and eliminate the influence of lens distortion on coordinate calculation.
[0080] Before constructing the dynamic spatial mapping model, the original pixel coordinates of the video frames need to be pre-corrected to eliminate the effects of lens distortion. The data processing center extracts distortion correction coefficients from the camera's optical parameters. These coefficients are calibrated at the camera's factory and describe the radial and tangential distortion characteristics of the lens. Simultaneously, the image correction model corresponding to the sensor type identifier outputs a set of correction parameters after preprocessing the video frames. These parameters reflect the initial distortion correction effect during image preprocessing. The data processing center combines the distortion correction coefficients and correction parameters to establish a pixel coordinate pre-correction formula through mathematical modeling. This formula takes the original pixel coordinates as input and outputs the pre-corrected pixel coordinates. After performing pre-correction processing on all the original pixel coordinates of the video frames, the pixel offset caused by lens distortion is eliminated, and the correspondence between pixel coordinates and actual spatial coordinates becomes more accurate.
[0081] Step S1313: Extract the latitude and longitude parameters from the camera pose information, convert the latitude and longitude parameters into Cartesian coordinates in the geodetic coordinate system, and use them as the origin coordinates of the dynamic spatial mapping model.
[0082] The latitude and longitude parameters in camera pose information are usually expressed in degrees, minutes, and seconds. To facilitate spatial coordinate calculations, they need to be converted to Cartesian coordinates in a geodetic coordinate system. The data processing center uses the Gauss-Kruger projection or other suitable map projection methods to convert the latitude and longitude parameters into Cartesian coordinates (X and Y coordinates). These Cartesian coordinates are determined as the origin coordinates of the dynamic spatial mapping model, i.e., the position of the camera in the geodetic coordinate system, and all subsequent spatial calculations will use this as the reference point.
[0083] Step S1314: Extract the altitude, pitch, and yaw parameters from the camera pose information, and construct a three-dimensional vector of the shooting direction. The direction of the three-dimensional vector is determined by the pitch and yaw parameters, and the length of the three-dimensional vector is positively correlated with the altitude parameters.
[0084] The three-dimensional vector representing the shooting direction describes the orientation of the camera's optical axis in three-dimensional space. The data processing center extracts altitude, pitch, and yaw parameters from the camera's pose information. The pitch parameter controls the vertical direction of the three-dimensional vector, while the yaw parameter controls its horizontal rotation. These two angle parameters determine the direction cosine of the three-dimensional vector. The length of the three-dimensional vector is positively correlated with the altitude parameter; the higher the altitude and the farther the camera is from the ground, the longer the three-dimensional vector, and vice versa.
[0085] Step S1315: Combine the focal length parameter and sensor size parameter in the camera's optical parameters to calculate the pixel field of view of the video frame. The pixel field of view is used to describe the spatial angle range corresponding to a single pixel.
[0086] The pixel field of view (FOP) is the angle that a single pixel can cover in space, and its calculation is closely related to the focal length and sensor size parameters. The data processing center first calculates the total horizontal and vertical FOPs of the video frame based on the focal length and sensor size parameters. Then, it divides the total horizontal FOP by the pixel width to obtain the horizontal FOP, and divides the total vertical FOP by the pixel height to obtain the vertical FOP. The smaller the pixel FOP, the higher the spatial resolution of a single pixel, and the richer the details of ground features that can be distinguished.
[0087] Step S1316: Starting from the origin coordinates, with the three-dimensional vector of the shooting direction as the central axis and the pixel field of view as the unit angle, construct a spatial ray model corresponding to the pixels of the video frame, with each pixel corresponding to a spatial ray.
[0088] The spatial ray model is fundamental for describing the correspondence between pixels and spatial points. The data processing center uses the origin coordinates of the dynamic spatial mapping model—the camera's position—as its starting point, and the three-dimensional vector of the shooting direction as its central axis. Each pixel in the video frame is considered a spatial ray emanating from the origin. The direction of the spatial ray is determined by the pixel's position within the video frame and its field of view. For the central pixel, its corresponding spatial ray coincides with the three-dimensional vector of the shooting direction; for pixels off-center, their corresponding spatial rays are deflected based on the pixel's horizontal and vertical offsets and its field of view. By constructing the spatial ray model, each pixel in the video frame corresponds to a spatial ray pointing towards a specific point on the ground surface.
[0089] Step S1317: Determine the transformation parameters between the geodetic coordinate system and the local coordinate system. The transformation parameters include translation parameters, rotation parameters, and scaling parameters, which are used to realize the mutual transformation of coordinates under the geodetic coordinate system and the local coordinate system.
[0090] Geodetic coordinates and local coordinates are two different coordinate systems. To achieve coordinate transformation between them, transformation parameters need to be determined. The data processing center obtains the coordinates of known control points in both the geodetic and local coordinate systems by setting up known control points within the survey area. Then, using parametric solution methods such as the least squares method, the translation parameters (ΔX, ΔY, ΔZ), rotation parameters (Rx, Ry, Rz), and scaling parameters (K) are calculated. These transformation parameters are stored for rapid coordinate transformation between the geodetic and local coordinate systems in subsequent calculations.
[0091] Step S1318: Perform an intersection operation between the spatial ray corresponding to each pixel and the ground reference plane to obtain the geographic coordinates of the pixel in the geodetic coordinate system, and convert the geographic coordinates to geographic coordinates in the local coordinate system through the transformation parameters.
[0092] The ground reference surface is typically assumed to be a plane, its height determined based on the average elevation of the surveyed area. The data processing center solves the equations of the spatial ray and the ground reference surface simultaneously for each pixel, obtaining the coordinates of the intersection point. These intersection coordinates are the pixel's geographic coordinates in the geodetic coordinate system (X_geodetic, Y_geodetic, Z_geodetic, where Z_geodetic is the ground reference surface height). Then, using determined transformation parameters, the geographic coordinates in the geodetic coordinate system are converted to geographic coordinates in the local coordinate system (X_local, Y_local, Z_local). Through intersection operations and coordinate transformation, each pixel corresponds to specific geographic coordinates in both the geodetic and local coordinate systems.
[0093] Step S1319: Calculate the extreme values of the geographic coordinates of all pixels in the video frame under the geodetic coordinate system to determine the geodetic reference geographic range of the video frame; calculate the extreme values of the geographic coordinates of all pixels in the video frame under the local coordinate system to determine the local reference geographic range of the video frame.
[0094] The geographic coordinate extremes are the minimum and maximum geographic coordinates corresponding to all pixels in a video frame. The data processing center statistically analyzes the X and Y geodetic coordinates of all pixels in the video frame under the geodetic coordinate system to identify the minimum X-geometry, maximum X-geometry, minimum Y-geometry, and maximum Y-geometry. These extreme values constitute the geodetic baseline geographic range of the video frame. Similarly, the X and Y local coordinates of all pixels under the local coordinate system are statistically analyzed to obtain the minimum X-locality, maximum X-locality, minimum Y-locality, and maximum Y-locality, determining the local baseline geographic range of the video frame. Compared to determining the geographic range based solely on the four corner points, the geographic range obtained by statistically analyzing the extreme values of all pixel coordinates is more accurate and can completely cover all features in the video frame.
[0095] Step S1320: Obtain the geographical range difference between adjacent video frames based on the inter-frame association identifier field, and perform smooth adjustment on the geodetic reference geographical range and local reference geographical range of the current video frame to make the geographical range transition between adjacent frames continuous.
[0096] The geographical extent of adjacent video frames should transition continuously to avoid jumps. The data processing center calculates the geographical extent difference between adjacent video frames based on the camera pose information difference recorded in the inter-frame association identifier field. Then, a smoothing filtering algorithm, such as a moving average algorithm, is used to adjust the geodetic and local reference geographical extents of the current video frame to ensure that the trend of geographical extent change is consistent with that of adjacent video frames. For example, if the geographical extent of an adjacent video frame shifts to the right in the X direction, the geographical extent of the current video frame should also be adjusted accordingly to the right to ensure a smooth transition.
[0097] Step S1320 may include:
[0098] Step S13201: Extract the inter-frame association identifier field of the current video frame and parse the difference in camera pose information between the current video frame and the previous video frame.
[0099] The inter-frame association identifier field stores the differences in camera pose information between the current video frame and the previous video frame, including differences in latitude and longitude, altitude, pitch angle, yaw angle, and roll angle. The data processing center parses this field to extract these pose information differences, which reflect the changes in the camera's position and attitude between adjacent frames.
[0100] Step S13202: Based on the geodetic reference geographical range and local reference geographical range of the previous video frame, and combined with the difference in camera pose information between the current video frame and the previous video frame, predict the initial geodetic reference geographical range and the initial local reference geographical range of the current video frame.
[0101] Based on the difference between the geographic extent and camera pose information of the previous video frame, the geographic extent of the current video frame can be predicted. The data processing center calculates the change in the camera's position in the geodetic coordinate system based on the latitude, longitude, and altitude differences in the pose information differences, and then predicts the translation direction and distance of the current video frame's geodetic reference geographic extent. Based on the pitch angle difference, yaw angle difference, and roll angle difference, it predicts the rotation and scaling changes of the geographic extent. Through these predictions, the initial geodetic reference geographic extent and the initial local reference geographic extent of the current video frame are obtained.
[0102] Step S13203: Calculate the difference between the actual geodetic reference geographic range obtained by the dynamic spatial mapping model for the current video frame and the predicted initial geodetic reference geographic range to obtain the geodetic reference geographic range deviation; calculate the difference between the actual local reference geographic range obtained by the dynamic spatial mapping model for the current video frame and the predicted initial local reference geographic range to obtain the local reference geographic range deviation.
[0103] The actual geographic extent is calculated using a dynamic spatial mapping model. The initial geographic extent is predicted based on the geographic extent and pose difference of the previous frame. The data processing center calculates the difference between the two, namely the geodetic baseline geographic extent deviation and the local baseline geographic extent deviation. The deviation can be calculated by comparing the boundary coordinate differences between the two geographic extents, such as the maximum X-coordinate difference and the minimum X-coordinate difference.
[0104] Step S13204: If the deviation of the geodetic reference geographical range is less than the preset smoothing threshold, the actual geodetic reference geographical range is directly used; if the deviation of the geodetic reference geographical range is greater than or equal to the preset smoothing threshold, a weighted average algorithm is used to fuse the actual geodetic reference geographical range with the predicted initial geodetic reference geographical range, and the weight coefficient is determined according to the difference between the camera pose information of the current video frame and the previous video frame.
[0105] A preset smoothing threshold is used to determine whether geographic range deviations require smoothing. When the deviation of the geodetic reference geographic range is less than this threshold, it indicates that the difference between the actual and predicted geographic ranges is small, and the geographic range changes relatively smoothly, so the actual geodetic reference geographic range can be directly used. When the deviation is greater than or equal to the threshold, a weighted average algorithm is used to fuse the actual and predicted geographic ranges to avoid abrupt changes in the geographic range. The weighting coefficients are related to the magnitude of the difference in camera pose information. The larger the difference in pose information, the more intense the camera movement, the lower the reliability of the predicted geographic range, and the greater the weight of the actual geographic range; conversely, the smaller the difference in pose information, the greater the weight of the predicted geographic range. Through weighted average fusion, the smoothed geodetic reference geographic range is obtained.
[0106] Step S13205: Perform the same deviation judgment and weighted fusion processing on the local baseline geographic range to obtain the smoothed local baseline geographic range.
[0107] Similar to the processing method for the geodetic baseline geographic range, the data processing center judges the deviation of the local baseline geographic range. If the deviation is less than the preset smoothing threshold, the actual local baseline geographic range is directly adopted; otherwise, a weighted average algorithm is used to merge the actual local baseline geographic range and the predicted initial local baseline geographic range to obtain the smoothed local baseline geographic range.
[0108] Step S13206: Extract the inter-frame association identifier field of the next video frame, parse the difference in camera pose information between the current video frame and the next video frame, and predict the geographical range of the next video frame.
[0109] To ensure a smooth transition in the geographical extent between the current and subsequent video frames, the data processing center also needs to consider the geographical extent of the subsequent video frame. By extracting the difference in pose information between the current video frame's inter-frame association identifier field and the subsequent video frame, a method similar to predicting the initial geographical extent of the current video frame is used to predict the geographical extent of the subsequent video frame.
[0110] Step S13207: Determine the overlap between the smoothed geographical range of the current video frame and the predicted geographical range of the next video frame. If the overlap is lower than the preset overlap threshold, adjust the weighted fusion coefficient of the current video frame and recalculate the smoothed geographical range until the overlap reaches the target.
[0111] The smoothed geographical extent of the current video frame needs to maintain a certain degree of overlap with the predicted geographical extent of the next video frame. The data processing center calculates the overlap and compares it with a preset overlap threshold. If the overlap is lower than the threshold, it indicates that the adjustment of the geographical extent of the current video frame may be excessive or insufficient. In this case, the weighted fusion coefficient of the current video frame needs to be adjusted, the smoothing adjustment needs to be performed again, and then the overlap with the geographical extent of the next video frame needs to be recalculated until the overlap meets the standard, ensuring that the geographical extent of the video frame sequence can continuously cover the survey area.
[0112] Step S13208: Perform linear interpolation on the boundary coordinates of the smoothed geodetic datum geographic range and the local datum geographic range to eliminate abrupt changes in the boundary coordinates.
[0113] Even after smoothing adjustments and overlap optimization, minor abrupt changes may still exist in the boundary coordinates of the geographic area. The data processing center employs linear interpolation to smooth these boundary coordinates. For example, for the X-direction boundary coordinate sequence of the geodetic datum geographic area, linear interpolation is used to adjust the differences between adjacent boundary coordinates, making the changes in boundary coordinates more uniform, eliminating abrupt changes, and further improving the continuity and smoothness of the geographic area.
[0114] Step S13209: Compare the adjusted geodetic reference geographic range and local reference geographic range with the preset geographic range rationality threshold. If the adjusted geodetic reference geographic range or local reference geographic range exceeds the preset geographic range rationality threshold, re-examine the inter-frame association identifier field and dynamic spatial mapping model parameters, correct them, and perform smoothing adjustment again. If the geographic range deviation of multiple consecutive frames is greater than the preset smoothing threshold, trigger the camera pose information verification process, re-acquire the camera pose information to correct the calculation basis.
[0115] The geographic range reasonableness threshold is used to determine whether the adjusted geographic range is within a reasonable spatial range, avoiding anomalies caused by calculation errors. The data processing center compares the adjusted geographic range with this threshold. If it exceeds the range, it may be due to an error in parsing the inter-frame association identifier field or improper settings of the dynamic spatial mapping model parameters. These data need to be re-checked and corrected before the geographic range is smoothed again. If the geographic range deviation in multiple consecutive frames exceeds the preset smoothing threshold, it indicates a potential large error in the camera pose information. In this case, the camera pose information verification process is triggered, re-acquiring the camera pose information through the pose perception link to correct the basic data for geographic range calculation.
[0116] Step S132010: After completing the smoothing adjustment, write the adjusted geodetic reference geographic range information, local reference geographic range information, and smoothing parameters into the attribute fields of the video frame.
[0117] After a series of calculations, adjustments, and verifications, the final geodetic reference geographic extent information and local reference geographic extent information are obtained, along with the smoothing parameters used in the smoothing adjustment process (such as weighted fusion coefficients, smoothing thresholds, etc.). The data processing center writes the above information into the attribute fields of the video frames and stores it together with the video frame data.
[0118] Step S1321: Compare the smoothed geodetic datum geographic range and local datum geographic range with the preset geographic range accuracy threshold. If the geographic range accuracy of any datum is lower than the geographic range accuracy threshold, readjust the calculation parameters of the spatial ray model, and perform the intersection operation and geographic range determination steps again until the geographic range accuracy meets the standard.
[0119] The geographic extent accuracy threshold is a standard for measuring the accuracy of geographic extent calculations, set according to the accuracy requirements of natural resource surveys. The data processing center compares the smoothed and adjusted geographic extent with this threshold. If the error of the geographic extent (such as the difference between the calculated geographic extent and the actual measured extent) is less than the accuracy threshold, the geographic extent accuracy is considered to meet the standard; otherwise, it is necessary to readjust the calculation parameters of the space ray model, such as the pixel field of view and the direction vector of the space ray, and then perform the intersection operation between the space ray and the ground reference plane again to redetermine the geographic extent, and perform smoothing adjustments and accuracy comparisons until the geographic extent accuracy meets the requirements.
[0120] Step S140: Based on the live video with multiple reference geographic range identifiers, the intelligent delineation operation is combined with the ground feature recognition model corresponding to the sensor type identifier to perform point, line or area trajectory recording along the outline of the ground features in the video frame. The coordinate transformation is performed according to the coordinates of the trajectory recording and the target geographic reference range corresponding to the video frame to generate geographic data with ground feature attributes and recognition confidence.
[0121] In natural resource surveys, operators need to delineate features on live video feeds with multiple geographic reference markers to obtain their geographic coordinates and attribute information. This process combines intelligent delineation with feature recognition models to improve the accuracy and efficiency of the delineation.
[0122] Step S141: Load the reference switching control and geographic range auxiliary lines into the playback interface of the live video with multiple reference geographic range identifiers. The geographic range auxiliary lines are generated based on the geodetic reference geographic range information or local reference geographic range information in the video frame attribute field and are used to identify the target geographic reference scale corresponding to the pixel coordinates in the video frame.
[0123] The live video playback interface serves as an interactive platform for operators to delineate geographic features. The data processing center loads a benchmark switching control onto the playback interface, allowing operators to switch between geodetic and local benchmarks. Simultaneously, based on the currently selected target geographic benchmark extent information, geographic extent auxiliary lines are generated. These auxiliary lines are displayed on the video frame as grid lines or a ruler; the grid spacing corresponds to the actual geographic distance, while the ruler indicates the geographic coordinates of the video frame edges under the target geographic benchmark. Through these auxiliary lines, operators can intuitively understand the actual geographic scale corresponding to the pixel coordinates within the video frame, aiding in accurate feature delineation.
[0124] Step S142: Receive the target geographic reference selection instruction through the reference switching control, and determine whether the target geographic reference range corresponding to the video frame is the geodetic reference geographic range or the local reference geographic range.
[0125] Operators select the target geographic benchmark by clicking the benchmark switching control, based on the needs of the survey task. After receiving the selection instruction, the data processing center determines whether the target geographic benchmark range corresponding to the video frame is a geodetic benchmark range or a local benchmark range, and extracts the corresponding geographic range information from the attribute fields of the video frame for subsequent coordinate transformation.
[0126] Step S143: Call the corresponding ground feature recognition model according to the sensor type identifier. If it is a visible light sensor identifier, call the visible light ground feature extraction model; if it is a thermal infrared sensor identifier, call the thermal infrared ground feature temperature feature model.
[0127] The sensor type identifier determines the selection of the ground feature recognition model. The data processing center calls the corresponding ground feature recognition model based on the extracted sensor type identifier. The visible light ground feature extraction model is suitable for extracting the contour and texture features of ground features from visible light video frames, and identifying ground feature types, such as buildings, roads, and vegetation, through the analysis of these features. The thermal infrared ground feature temperature feature model is suitable for extracting the temperature distribution features of ground features from thermal infrared video frames, and identifying ground features based on temperature differences, such as water bodies, heat sources, and vegetation-covered areas.
[0128] Step S144: Input the currently playing video frame into the corresponding ground feature recognition model, extract the ground feature feature vector within the video frame. The ground feature feature vector includes contour features, texture features, or temperature distribution features, and output the ground feature recognition confidence score.
[0129] The currently playing video frame is input into the invoked land cover recognition model, which extracts features from the video frame. For the visible light land cover feature extraction model, edge detection algorithms are used to extract the contour features of land covers, and methods such as gray-level co-occurrence matrix are used to extract texture features. These features are then combined into a land cover feature vector. For the thermal infrared land cover temperature feature model, temperature field analysis is used to extract the temperature distribution features of land covers, such as temperature mean, temperature variance, and temperature gradient, forming a land cover feature vector. Simultaneously, the model outputs a land cover recognition confidence score based on the degree of feature matching. This confidence score reflects the model's certainty regarding the type of land cover being identified; a higher confidence score indicates a more reliable recognition result.
[0130] Step S145: Move the operation trajectory along the contour of the ground features within the video frame through intelligent outlining operation. The ground feature recognition model matches the feature vector of the ground feature with the pixel coordinates of the operation trajectory in real time. If the matching degree is lower than the preset threshold, a trajectory correction prompt is triggered; if the matching degree is higher than the preset threshold, the operation trajectory continues to be recorded.
[0131] Operators use input devices such as mice and styluses to perform intelligent drawing operations on the playback interface, moving their trajectory along the outline of the target feature within the video frame. The feature recognition model acquires the pixel coordinates of the trajectory in real time and matches them with the extracted feature feature vectors. The matching degree is calculated based on the degree of overlap between the pixel coordinates of the trajectory and the feature outline features. If the matching degree is lower than a preset threshold, it indicates that the operator's drawing trajectory deviates from the actual feature outline, and the model triggers a trajectory correction prompt, such as displaying the correct outline guide line on the interface or issuing an audio prompt. If the matching degree is higher than the preset threshold, it indicates that the drawing trajectory conforms to the feature outline, and the model continues to record the pixel coordinates of the trajectory.
[0132] Step S146: If the target feature is a point feature, record the single-point coordinates of the operation trajectory when the operation point stops moving and the matching degree is higher than the preset threshold; if the target feature is a line feature, continuously record pixel coordinates to form a line coordinate sequence when the operation point moves and the matching degree is continuously higher than the preset threshold; if the target feature is a surface feature, record the pixel coordinate polygon of the closed trajectory when the operation point moves to form a closed trajectory and the matching degree is continuously higher than the preset threshold.
[0133] Depending on the shape and characteristics of the target feature, the recording method of the operation trajectory varies. For point features, such as isolated trees or utility poles, the operator clicks or briefly pauses at the location. When the operation point stops moving and the matching degree is higher than a preset threshold, the model records the pixel coordinates of that single point. For linear features, such as roads or rivers, the operator moves the operation point along its centerline. As long as the matching degree remains higher than a preset threshold, the model continuously records pixel coordinates at certain time intervals or distance intervals, forming a linear coordinate sequence. For areal features, such as lakes or farmland, the operator moves the operation point along its boundary. When the operation point moves to form a closed trajectory and the matching degree remains higher than a preset threshold, the model records the pixel coordinates of that closed trajectory, forming a polygon.
[0134] Step S147: Extract the geographic coordinate information corresponding to the target geographic reference range in the attribute field of the current video frame, and obtain the geographic coordinates of the upper left and lower right corners of the video frame under the target geographic reference.
[0135] To convert the recorded pixel coordinates into geographic coordinates under the target geographic reference, it is necessary to obtain the boundary geographic coordinates of the video frame under the target geographic reference. The data processing center extracts the geographic coordinate information corresponding to the target geographic reference range from the attribute fields of the current video frame, including the geographic coordinates (longitude, latitude in the geodetic coordinate system or Cartesian coordinates in the local coordinate system) of the upper left and lower right corners of the video frame under the target geographic reference.
[0136] Step S148: Calculate the horizontal and vertical geographic spans of the video frame under the target geographic reference. Combine the pixel width and pixel height of the video frame to obtain the geographic distance per unit pixel under the target geographic reference.
[0137] Geographic span refers to the geographical distance covered by a video frame in both the horizontal and vertical directions. The data processing center calculates the horizontal and vertical geographic spans by measuring the difference in geographic coordinates between the bottom right and top left corners of the video frame. For example, in a geodetic coordinate system, the horizontal geographic span is the distance corresponding to the bottom right longitude minus the top left longitude, and the vertical geographic span is the distance corresponding to the bottom right latitude minus the top left latitude. Then, the horizontal geographic span is divided by the pixel width of the video frame to obtain the geographic distance per unit pixel in the horizontal direction; the vertical geographic span is divided by the pixel height of the video frame to obtain the geographic distance per unit pixel in the vertical direction.
[0138] Step S149: Convert the geographic coordinates of the top left corner of the video frame under the target geographic reference to Cartesian coordinates.
[0139] To facilitate the conversion calculation from pixel coordinates to geographic coordinates, the geographic coordinates of the upper left corner of the video frame need to be converted to Cartesian coordinates. For latitude and longitude coordinates in the geodetic coordinate system, the same map projection method as in step S143 is used to convert them to Cartesian coordinates; for geographic coordinates in the local coordinate system, if they are already Cartesian coordinates, no conversion is needed.
[0140] Step S1410: Convert the geographic coordinates of the top left corner of the video frame under the target geographic reference to planar rectangular coordinates; for single-point pixel coordinates, calculate its horizontal and vertical offsets relative to the top left pixel, multiply the horizontal offset by the horizontal unit pixel geographic length and the vertical offset by the vertical unit pixel geographic length to obtain the planar offset, and superimpose the top left corner planar rectangular coordinates to obtain the target planar rectangular coordinates; convert the target planar rectangular coordinates back to the geographic coordinates under the target geographic reference; for linear coordinate sequences and closed coordinate polygons, repeat the above offset calculation, planar coordinate superposition and conversion steps to obtain the geographic coordinates under the target geographic reference corresponding to each coordinate.
[0141] For the recorded operation trajectory pixel coordinates, they need to be converted to geographic coordinates under the target geographic reference. Taking the Cartesian coordinates of the top-left corner of the video frame as the origin, for a single pixel coordinate, calculate its horizontal offset (X value of the pixel coordinate) and vertical offset (Y value of the pixel coordinate) relative to the top-left corner pixel. Then, multiply the horizontal offset by the geographic distance per unit pixel in the horizontal direction to obtain the horizontal plane offset; multiply the vertical offset by the geographic distance per unit pixel in the vertical direction to obtain the vertical plane offset. Superimpose these two plane offsets onto the Cartesian coordinates of the top-left corner to obtain the target Cartesian coordinates corresponding to the single pixel coordinate. Finally, convert the target Cartesian coordinates back to geographic coordinates under the target geographic reference (if the target geographic reference is a geodetic coordinate system, convert to latitude and longitude coordinates; if it is a local coordinate system, keep the plane Cartesian coordinates). For linear coordinate sequences and closed coordinate polygons, perform the above offset calculation, plane coordinate superposition, and coordinate transformation steps for each pixel coordinate to obtain the geographic coordinates under the target geographic reference corresponding to each coordinate.
[0142] Step S1411: Add feature attribute information to the converted geographic coordinates. The feature attribute information includes feature type identifier, delineation time information, timestamp information of the corresponding video frame, and feature recognition confidence score output by the feature recognition model.
[0143] The converted geographic coordinates need to be associated with feature attribute information to form complete geographic data. The data processing center adds feature type identifiers, such as "building," "road," and "water body," to the geographic coordinates based on the feature types identified by the feature recognition model. Simultaneously, it records the time information of the delineation operation, the timestamp information of the corresponding video frame, and the feature recognition confidence score output by the feature recognition model. This attribute information enriches the content of the geographic data, making it include not only spatial coordinates but also semantic and qualitative information about the features.
[0144] Step S1412: Associate and store the geographic coordinates with the added attributes and the confidence level of the ground feature recognition to generate geographic data with ground feature attributes and recognition confidence level.
[0145] The data processing center associates geographic coordinates with added feature attribute information and feature identification confidence levels, creating structured geographic data. This geographic data can be stored in the form of a database or files for easy querying, updating, and analysis. Geographic data with feature attributes and identification confidence levels is an important outcome of natural resource surveys and can be used for map creation, feature distribution analysis, and environmental change monitoring.
[0146] Furthermore, the method may also include:
[0147] Step S1413: During the playback of a live video with multiple reference geographical range identifiers, when the intelligent delineation operation is triggered, the currently playing video frame is locked, the playback of the video frame sequence is paused, and the sensor type identifier of the video frame is extracted.
[0148] When an operator triggers the intelligent delineation operation during live video playback, the data processing center immediately locks the currently playing video frame and pauses the playback of the video frame sequence to ensure that the operator can accurately delineate on a static screen. Simultaneously, it extracts the sensor type identifier of the locked video frame to prepare for calling the corresponding land feature recognition model.
[0149] Step S1414: Call the corresponding ground feature recognition model according to the sensor type identifier, and load the pre-trained ground feature feature template library of the model. The ground feature feature template library contains standard feature vectors of a variety of common ground features.
[0150] When performing feature matching, the ground feature recognition model needs to refer to standard ground feature feature vectors, which are stored in a ground feature template library. After the data processing center calls the corresponding ground feature recognition model based on the sensor type identifier, it loads the ground feature template library pre-trained by that model. The template library contains standard feature vectors of various common ground features, such as the outline and texture feature vectors of different types of buildings, the outline and texture feature vectors of different road grades, the texture and color feature vectors of different vegetation types (for visible light), or the temperature distribution feature vectors of ground features with different temperature characteristics (for thermal infrared).
[0151] Step S1415: Input the locked video frame into the land feature recognition model, traverse the pixel region of the video frame through a sliding window, and calculate the similarity between the feature vector of each window region and the standard feature vector in the land feature template library.
[0152] The feature recognition model uses a sliding window approach to scan locked video frames. The size of the sliding window is set according to the dimensions of common features, and the window moves from left to right and from top to bottom on the video frame, moving a fixed step size each time. For each window region, the model extracts its feature vector, such as contour feature vector, texture feature vector, or temperature distribution feature vector, and then calculates its similarity with all standard feature vectors in the feature template library. The similarity calculation uses methods such as cosine similarity and Euclidean distance to obtain the similarity value between each window region and the standard feature vectors of various types of features.
[0153] Step S1416: Mark the window regions with similarity higher than the preset matching threshold as candidate land cover regions, and output the boundary coordinates of the candidate land cover regions and the preliminary determination results of the corresponding land cover types.
[0154] A preset matching threshold is used to determine whether a window region is a land feature region. The model compares the similarity value of each window region with the preset matching threshold and marks window regions with similarity values higher than the threshold as candidate land feature regions. Simultaneously, based on the land feature type corresponding to the standard feature vector with the highest similarity, the model outputs a preliminary determination result of the land feature type for the candidate land feature regions. The data processing center displays the boundary coordinates of the candidate land feature regions as highlighted boxes on the playback interface, along with the preliminary determination result of the land feature type, to assist operators in the delineation operation.
[0155] Step S1417: Move the operation point within the candidate feature area through intelligent outlining operation. The feature recognition model calculates the consistency between the feature vector of the operation point location and the feature vector of the candidate feature area in real time. If the consistency is lower than a preset threshold, a vibration prompt is issued to correct the operation trajectory.
[0156] When an operator moves a control point within a candidate feature area to draw, the feature recognition model calculates the pixel feature vector of the control point's location in real time and compares it with the feature vector of the candidate feature area. Consistency reflects whether the control point's location belongs to the candidate feature area. If the consistency is below a preset threshold, it indicates that the control point has deviated from the candidate feature area. The model then uses vibration feedback or sound prompts on the playback interface to remind the operator to correct the trajectory and ensure that the drawing trajectory is within the correct feature area.
[0157] Step S1418: If the target feature is a point feature, when the operation point moves to the center of the candidate feature area and the consistency is higher than the preset threshold, record the single-point pixel coordinates of that location.
[0158] For point features, the operator moves the control point to the center of the candidate feature area. When the feature recognition model detects that the control point is in the center and the consistency is higher than a preset threshold, it automatically records the single-point pixel coordinates at that location. The center position can be determined by calculating the geometric center of the candidate feature area.
[0159] Step S1419: If the target feature is a linear feature, when the operation point moves along the center line of the candidate feature area and the consistency is consistently higher than the preset threshold, the pixel coordinates are recorded at the preset sampling interval to form a linear coordinate sequence.
[0160] Linear features have a distinct centerline characteristic. Operators move operation points along the centerline of the candidate feature area. The feature recognition model monitors the consistency of operation points in real time. When the consistency consistently exceeds a preset threshold, the pixel coordinates of the operation points are recorded at preset sampling intervals (e.g., every 10 pixels or every 0.5 seconds). These coordinates are arranged sequentially to form a linear coordinate sequence, which can accurately describe the orientation of the linear feature.
[0161] Step S1420: If the target feature is a planar feature, when the operation point moves along the boundary of the candidate feature area and the consistency is consistently higher than the preset threshold, the pixel coordinates of the boundary are recorded to form a closed coordinate polygon.
[0162] The areal features have closed boundaries. The operator moves the operation point along the boundary of the candidate feature area. When the feature recognition model detects that the consistency of the operation point is continuously higher than a preset threshold, it continuously records the pixel coordinates of the operation point. When the operation point moves back to the starting position and forms a closed trajectory, the recorded pixel coordinates form a closed coordinate polygon, which can accurately delineate the outline of the areal feature.
[0163] Step S1421: Extract the top-left and bottom-right geographic coordinates of the target geographic reference range from the attribute field of the locked video frame, and calculate the horizontal and vertical geographic lengths of the video frame.
[0164] The data processing center extracts the top-left and bottom-right geographic coordinates of the target geographic reference range from the attribute fields of the locked video frame. Then, it calculates the horizontal geographic length (i.e., the east-west distance) and vertical geographic length (i.e., the north-south distance) of the video frame based on these two coordinates. The method for calculating the geographic length depends on the type of the target geographic reference: for geodetic coordinate systems, it calculates the arc distance on the Earth's surface using latitude and longitude coordinates; for local coordinate systems, it directly calculates the difference between the plane rectangular coordinates.
[0165] Step S1422: Calculate the horizontal unit pixel geographic length based on the pixel width of the video frame, and calculate the vertical unit pixel geographic length based on the pixel height of the video frame.
[0166] The horizontal geographic length per pixel is equal to the horizontal geographic length of the video frame divided by the pixel width of the video frame, and the vertical geographic length per pixel is equal to the vertical geographic length of the video frame divided by the pixel height of the video frame. These two parameters represent the actual geographic distance of each pixel in the horizontal and vertical directions, and are the key scaling factors for converting pixel coordinates to geographic coordinates.
[0167] Step S150: Transmit geographic data with feature attributes and identification confidence levels to web maps, mobile terminal maps, and display interfaces via a multi-link network. Simultaneously, receive geographic data modification trajectories from web maps and mobile terminal maps. Based on the geographic data modification trajectories, the feature identification confidence levels corresponding to the modification trajectories, and the coordinate information of corresponding features in the live video with multi-reference geographic range identifiers, perform layered update processing to achieve bidirectional dynamic synchronization and conflict resolution between geographic data and live video.
[0168] In natural resource survey scenarios, the generated geographic data needs to be shared with different user terminals, such as web maps and mobile terminal maps. At the same time, users may modify the geographic data, and the data processing center needs to receive and process these modification trajectories to maintain the consistency between the geographic data and the live video.
[0169] Step S151: Perform data sharding processing on geographic data with feature attributes and recognition confidence. Adjust the shard size according to the transmission bandwidth limit of web map and mobile terminal map and the feature recognition confidence. The higher the feature recognition confidence, the smaller the shard is to improve transmission efficiency.
[0170] To improve the transmission efficiency and reliability of geographic data, it is necessary to perform fragmentation. The data processing center determines the basic fragment size based on the transmission bandwidth limitations of different user terminals (web maps, mobile terminal maps). Meanwhile, considering that geographic data with high feature identification confidence is relatively stable and less likely to be modified, its fragment size is set smaller to reduce the amount of data transmitted and improve transmission efficiency. For geographic data with low feature identification confidence, which may require frequent modifications and updates, the fragment size can be appropriately increased to reduce the number of fragments and management overhead.
[0171] Step S152: Add a data identifier, synchronization version number, feature identification confidence copy, and fragment check code to each fragment.
[0172] Each fragment requires the addition of necessary identification and verification information to ensure the accuracy and integrity of the transmission. The data identifier uniquely identifies each fragment, containing information such as the geographic data ID and fragment sequence number; the synchronization version number records updated versions of the geographic data, ensuring the receiving end can recognize the latest data; the feature identification confidence copy allows the receiving end to understand the reliability of feature identification before fully parsing the fragment data; the fragment checksum is generated using algorithms such as CRC32 and is used by the receiving end to verify whether errors occurred during the transmission of the fragment data.
[0173] Step S153: Establish a multi-link network transmission channel, allocate independent transmission links for web map, mobile terminal map and display interface respectively, and set a link quality monitoring field for each transmission link.
[0174] To ensure that different user terminals can simultaneously acquire geographic data, the data processing center has established a multi-link network transmission channel. Independent transmission links are allocated to web maps, mobile terminal maps, and display interfaces to avoid link congestion and data interference. Each transmission link header includes a link quality monitoring field for real-time monitoring of parameters such as transmission rate, packet loss rate, and latency.
[0175] Step S154: The fragmented geographic data is sent to the web map, mobile terminal map and display interface respectively through the multi-link network transmission channel. During the transmission process, the fragment transmission progress and link quality monitoring results are fed back in real time. The integrity of the geographic data fragments received by each end is verified by the fragment check code. If there are missing fragments, retransmission is triggered.
[0176] The data processing center transmits the fragmented geographic data to various user terminals via a multi-link network transmission channel. During transmission, it periodically reports the fragment transmission progress to the sending end, such as the number of fragments transmitted and the number of fragments remaining, while also reporting the link quality monitoring results. Upon receiving the fragmented data, the receiving end verifies the data integrity using a fragment checksum. If the verification fails, it indicates that the fragmented data is missing or incorrect. The receiving end sends a retransmission request to the sending end, and the data processing center triggers the retransmission mechanism to ensure that all fragmented data is received completely.
[0177] Step S155: Receive the geographic data modification trajectory from the web map and mobile terminal map. The geographic data modification trajectory includes the modified geographic coordinates, modification operation type, modification timestamp, and corresponding feature ID.
[0178] When users view geographic data on web maps or mobile maps, they may find data errors or information that needs updating. In this case, they can modify the geographic data through interactive operations and report the modification trajectory back to the data processing center. The geographic data modification trajectory includes the modified geographic coordinates (new location of the feature), the type of modification operation (such as adding, deleting, moving, modifying attributes, etc.), the modification timestamp (recording the time of the modification operation), and the corresponding feature's identifier ID (used to uniquely identify the modified feature).
[0179] Step S156: Based on the feature ID, query the recognition confidence of the corresponding feature. Divide the modified trajectory into a first processing group and a second processing group according to the recognition confidence of the corresponding feature. The modified trajectory with the recognition confidence of the feature corresponding to it being higher than the preset confidence threshold is assigned to the first processing group, and the modified trajectory with the recognition confidence of the feature corresponding to it being lower than or equal to the preset confidence threshold is assigned to the second processing group.
[0180] The confidence level of a feature's identification reflects the reliability of the original geographic data. The data processing center queries the geographic data containing feature attributes and their confidence levels based on the feature's ID, then divides the modification trajectories into two groups. The first group contains modification trajectories for features with confidence levels higher than a preset confidence threshold; the original data for these features is relatively reliable, and modifications require careful handling. The second group contains modification trajectories for features with confidence levels lower than or equal to the preset confidence threshold; the original data for these features has lower reliability, and modifications are more likely.
[0181] Step S157: Extract the first geographic data modification trajectory from the first processing group, and search for the corresponding original geographic coordinates, geographic feature recognition confidence and associated video frame identifier in the geographic data with geographic feature attributes and recognition confidence based on the geographic feature identifier ID.
[0182] The data processing center extracts the geographic data modification trajectories from the first processing group in a certain order (such as the order of receipt time). For each modification trajectory, the corresponding original geographic coordinates, geographic feature recognition confidence, and associated video frame identifier are searched in the geographic database based on the feature identifier ID. The video frame identifier is used to locate the original video frame corresponding to the geographic data.
[0183] Step S158: Locate the target video frame in the live video with multiple reference geographic range identifiers based on the video frame identifier, extract the target geographic reference range information of the target video frame, and calculate the pixel coordinates corresponding to the modified geographic coordinates under the target geographic reference.
[0184] Based on the associated video frame identifier, the data processing center locates the corresponding target video frame in the live video with multiple reference geographic range identifiers. Then, it extracts the target geographic reference range information (geodetic reference or local reference) of the target video frame, and calculates the pixel coordinates corresponding to the modified geographic coordinates on the target video frame based on the geographic reference range. The calculation method is similar to the geographic coordinate to pixel coordinate conversion method in step S1410.
[0185] Step S159: Determine whether the modified geographic coordinates are within the target geographic reference range of the target video frame. If the modified geographic coordinates exceed the target geographic reference range of the target video frame, send an invalid modification prompt to the feedback end and move the modification trajectory to the second processing group; if the modified geographic coordinates are within the target geographic reference range of the target video frame, retain the geographic data modification trajectory.
[0186] The modified geographic coordinates are only valid if they fall within the geographic reference range of the target video frame; otherwise, the modification operation may be erroneous. The data processing center compares the modified geographic coordinates with the target geographic reference range of the target video frame. If they exceed the range, an invalid modification message is sent to the user terminal that provided the modification trajectory, and the modified trajectory is moved from the first processing group to the second processing group. If they are within the range, the modified trajectory is retained, and subsequent processing continues.
[0187] Step S1510: Extract the modified geographic coordinates from the retained geographic data modification trajectory, update the original geographic coordinates of the corresponding geographic features in the geographic data with feature attributes and identification confidence, and update the synchronization version number at the same time.
[0188] For the retained geographic data modification trajectory, the data processing center extracts the modified geographic coordinates, finds the corresponding geographic feature record in the geographic data with feature attributes and identification confidence levels, and updates the original geographic coordinates with the new geographic coordinates. At the same time, the synchronization version number of the geographic data is incremented by 1 to indicate that the data has been updated.
[0189] Step S1511: Delete the outline corresponding to the original geographic coordinates in the target video frame, and redraw the outline according to the calculated pixel coordinates to complete the update of the object coordinate information in the target video frame.
[0190] To ensure that the outlined trajectory in the live video is consistent with the updated geographic data, the data processing center deletes the outlined trajectory graphic corresponding to the original geographic coordinates in the target video frame. Then, based on the calculated pixel coordinates corresponding to the modified geographic coordinates, the outlined trajectory is redrawn using the same style (such as color and line thickness) as the original outline, thus completing the visualization update of the object coordinate information in the target video frame.
[0191] Step S1512: Synchronize the updated target video frame and the updated geographic data with feature attributes and identification confidence to the web map, mobile terminal map and display interface, and keep the data version consistent across all terminals by synchronizing the version number.
[0192] The data processing center synchronizes the updated target video frames and updated geographic data to all user terminals (web maps, mobile terminal maps, and display interfaces) via a multi-link network transmission channel. During the synchronization process, the synchronization version number is compared to ensure that each terminal receives the latest version of the data, thus avoiding data inconsistencies.
[0193] Step S1513: After the first processing group is completed, repeat the above steps to process the geographic data modification trajectory in the second processing group. If the modification trajectory in the second processing group conflicts with the updated data, calculate the modification timestamp of the conflict trajectory and the update timestamp of the updated data, retain the modification result of the timestamp update, and discard the conflict trajectory with the oldest timestamp.
[0194] After the geographic data modification trajectories in the first processing group are processed, the data processing center processes the modification trajectories in the second processing group using the same steps. During processing, conflicts may arise between the modification trajectories and the updated data, such as when two different users modify the same feature. In this case, the data processing center compares the modification timestamp of the conflicting trajectory with the update timestamp of the updated data, retaining the modification result with the updated timestamp, as this reflects the latest modification operation, and discarding the conflicting trajectory with the older timestamp.
[0195] Step S1514: After completing all modified trajectory processing, generate a two-way synchronization completion marker and feed it back to the web map, mobile terminal map and display interface to achieve two-way dynamic synchronization and conflict resolution between geographic data and live video.
[0196] Once all geographic data modification trajectory processing is complete, the data processing center generates a two-way synchronization completion identifier. This identifier includes information such as the synchronization completion time and the amount of updated data, and sends this information back to each user terminal. Upon receiving this identifier, the user terminal displays a notification that data synchronization is complete. At this point, the two-way dynamic synchronization and conflict resolution process between geographic data and live video is complete, ensuring that all user terminals use the latest and consistent geographic data.
[0197] For example, step S1515: Establish a modified trajectory receiving buffer, set the maximum storage capacity and timeout time of the modified trajectory receiving buffer, and store the modified trajectory of the geographic data fed back by the web map and mobile terminal map into the modified trajectory receiving buffer in the order of receiving time.
[0198] To handle the potential for a large number of modified tracks to be fed back simultaneously, the data processing center has established a modified track receiving buffer. The buffer has a maximum storage capacity; when the number of modified tracks stored in the buffer reaches this capacity, the reception of new modified tracks is paused until some tracks are processed and space is freed up. A timeout period is also set; if the buffer is not filled within the timeout period, processing of already received modified tracks will be triggered. Modified tracks are stored in the buffer in the order of their reception time to ensure sequential processing.
[0199] Step S1516: When the storage capacity of the modified trajectory receiving buffer reaches the maximum storage capacity or the timeout period is exceeded, deduplication processing is performed on the modified trajectories in the modified trajectory receiving buffer. If there are multiple modified trajectories with the same feature ID, the modified trajectory with the latest modification timestamp is retained, and the remaining duplicate modified trajectories are deleted.
[0200] When the modified trajectory receiving buffer reaches the trigger condition (storage capacity reaches maximum or timeout period), the data processing center performs deduplication on the modified trajectories in the buffer. By comparing the feature IDs, multiple modified trajectories targeting the same feature are identified. The one with the latest modification timestamp is retained, while other duplicate modified trajectories are deleted to avoid duplicate processing of the same feature and improve processing efficiency.
[0201] Step S1517: Query the feature identification confidence level corresponding to each deduplicated modified trajectory based on the feature identifier ID. According to the comparison results between the feature identification confidence level and the preset confidence threshold, divide the modified trajectory into a first processing group and a second processing group. Modified trajectories with feature identification confidence levels higher than the preset confidence threshold are assigned to the first processing group, and modified trajectories with feature identification confidence levels lower than or equal to the preset confidence threshold are assigned to the second processing group.
[0202] The modified trajectories after deduplication are regrouped according to the ground feature recognition confidence level. The grouping method is the same as in step S156, dividing the modified trajectories into a first processing group and a second processing group for separate processing.
[0203] Step S1518: Process the first processing group and extract the feature ID, modified geographic coordinates, and modification timestamp from the first modified trajectory.
[0204] The data processing center extracts key information from the first modified trajectory from the first processing group, including the feature ID, the modified geographic coordinates, and the modification timestamp, in preparation for subsequent queries and processing.
[0205] Step S1519: Based on the feature identifier ID, query the corresponding original geographic coordinates, associated video frame identifiers, and target geographic reference type in the geographic data with feature attributes and identification confidence level.
[0206] Based on the feature ID, query the geographic database for the original geographic coordinates, associated video frame identifiers, and the target geographic reference type (geodetic reference or local reference) used for the modified trajectory.
[0207] Step S1520: Obtain the target video frame from the live video with multiple reference geographic range identifiers based on the video frame identifier, and extract the geographic range information corresponding to the target geographic reference type in the target video frame.
[0208] Using the associated video frame identifier, the target video frame in the live video with multiple reference geographic range identifiers is located, and the geographic range information of the corresponding target geographic reference type, such as the geographic coordinates of the upper left and lower right corners, is extracted from the attribute fields of the video frame.
[0209] Step S1521: Determine whether the modified geographic coordinates are within the geographic range of the target geographic reference type. If the modified geographic coordinates are not within the geographic range of the target geographic reference type, send a modification suggestion containing the geographic range boundary of the target geographic reference type to the feedback end, and move the modification trajectory to the second processing group. If the modified geographic coordinates are within the geographic range of the target geographic reference type, continue with subsequent processing.
[0210] Similar to step S159, the data processing center determines whether the modified geographic coordinates are within the geographic range of the target video frame. If they are not within the range, a modification suggestion is sent to the feedback terminal, indicating to the user that the modified coordinates exceed the geographic range of the video frame, and providing geographic range boundary information. At the same time, the modified trajectory is moved to the second processing group. If they are within the range, processing continues.
[0211] Step S1522: Calculate the deviation between the modified geographic coordinates and the original geographic coordinates. If the deviation exceeds the preset deviation threshold, a secondary confirmation process is triggered, and a deviation prompt is sent to the feedback terminal. Processing continues after receiving the confirmation instruction. If the deviation does not exceed the preset deviation threshold, processing continues directly.
[0212] To avoid accidental modifications, modifications to geographic data with high confidence levels in the first processing group require checking the magnitude of the modification. The data processing center calculates the deviation (e.g., Euclidean distance) between the modified geographic coordinates and the original geographic coordinates. If the deviation exceeds a preset deviation threshold, it indicates a large modification and potential error. In this case, a secondary confirmation process is triggered, sending a deviation alert to the feedback system indicating that the modification exceeds the normal range. Processing continues only after the user confirms the modification is correct. If the deviation does not exceed the preset deviation threshold, the modification is considered reasonable, and processing continues directly.
[0213] Step S1523: Based on the geographic range information of the target geographic reference type, convert the modified geographic coordinates into pixel coordinates of the target video frame.
[0214] Using a method similar to step S1410, the modified geographic coordinates are converted into pixel coordinates of the target video frame so that the drawn trajectory can be updated on the video frame.
[0215] Step S1524: Delete the outlined trajectory graphic corresponding to the original geographic coordinates in the target video frame, redraw the outlined trajectory graphic according to the converted pixel coordinates, and update the attribute information of the outlined trajectory graphic.
[0216] The data processing center deletes the original geographic coordinates-corresponding trajectory graphic from the target video frame and then redraws a new trajectory graphic based on the converted pixel coordinates. Simultaneously, it updates the attribute information of the trajectory graphic, such as modifying the timestamp and the person who modified it, to track the modification history.
[0217] Step S1525: Update the geographic coordinates and synchronization version number of the corresponding feature identifier ID in the geographic data with feature attributes and identification confidence.
[0218] In the geographic database, update the geographic coordinates of the corresponding feature ID to the modified geographic coordinates, and increment the synchronization version number by 1 to reflect the data update.
[0219] Step S1526: Synchronize the updated target video frame and the updated geographic data with feature attributes and identification confidence level to the web map, mobile terminal map and display interface through the original transmission link, and carry the synchronization version number during the synchronization process to ensure that the data on each terminal is consistent.
[0220] Through the previously established multi-link network transmission channel, the updated target video frames and geographic data are synchronized to all user terminals. During the synchronization process, the synchronization version number is used to ensure that the data on each terminal is consistent.
[0221] Step S1527: Repeat the above steps to process all modified trajectories in the first processing group. After the first processing group is completed, process the second processing group in the same way. If the modified trajectory in the second processing group conflicts with the updated geographic data, calculate the modification timestamp of the conflicting trajectory and the update timestamp of the updated data. Keep the modification result of the timestamp update and discard the conflicting trajectory with the oldest timestamp.
[0222] The data processing center processes all modified tracks in the first processing group in a loop, and then processes the second processing group according to the same process. When processing the second processing group, if a conflict occurs between a modified track and the updated data, the conflict is resolved by comparing timestamps, and the latest modification result is retained.
[0223] Step S1528: After all modified trajectories have been processed, an update report containing the processing results is generated and sent back to each terminal.
[0224] After processing is complete, the data processing center generates an update report, which includes information such as the number of modified trajectories processed, the number of successfully updated data, and the status of conflict resolution. The report is then sent to each user terminal so that users can understand the update status of the geographic data.
[0225] Throughout the aforementioned natural resource survey scenario, the collection and processing of drone video data, camera parameters, pose information, and geographic data are involved. Among these, the latitude and longitude parameters in the camera pose information are considered privacy-sensitive data; unauthorized access could potentially reveal the drone's flight path and operational area. To protect this privacy-sensitive data, during the data collection phase, encrypted transmission protocols are used to encrypt the pose information transmitted via the attitude perception link, such as SSL / TLS. During the data storage phase, the stored latitude and longitude parameters are encrypted using encryption algorithms such as AES. When data is transmitted to the user terminal, only the Cartesian coordinates related to the geographic range calculation are transmitted, rather than the raw latitude and longitude parameters directly. These technical measures ensure the protection and prevention of leakage of privacy-sensitive data.
[0226] Based on the same inventive concept, please refer to Figure 2 This paper shows a schematic block diagram of a UAV video spatialization-based information processing system 100 provided in an embodiment of this application for executing the above-described UAV video spatialization-based information processing method. The UAV video spatialization-based information processing system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0227] Alternatively, the machine-readable storage medium 120 can also be integrated into the processor 130 and can communicate and interact with external systems through the communication unit 110. The machine-readable storage medium 120 stores machine-executable instructions for executing the scheme of this application, and the processor 130 executes the machine-executable instructions stored in the machine-readable storage medium 120 to implement the information processing method based on UAV video spatialization provided in the aforementioned method embodiments.
[0228] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. An information processing method based on UAV video spatialization, characterized in that, The method includes: Collect camera optical parameters, camera pose information and sensor type identifiers, and synchronously encode the camera optical parameters, camera pose information and sensor type identifiers into the non-display data segment of the frame structure of the UAV video to form an encoded video with multi-dimensional spatial information. Perform streaming encapsulation processing on the encoded video with multi-dimensional spatial information, add inter-frame association identifier fields, and form an encoded video stream. The encoded video stream is subjected to scene-specific transcoding adaptation processing. Based on the format requirements of the subsequent decoding end and the image characteristics corresponding to the sensor type identifier, the quantization parameters and data encapsulation structure of the video encoding algorithm are adjusted to obtain an adapted video stream for different scenes. The adapted video stream is subjected to layered decoding and splitting processing to separate the frame sequence of the drone live video, the camera optical parameters, camera pose information and sensor type identifier embedded in the frame sequence, and the image correction model corresponding to the camera optical parameters, camera pose information and sensor type identifier. A dynamic spatial mapping model is constructed by combining the image correction model corresponding to the camera optical parameters, camera pose information and sensor type identifier. The geographic range of each video frame in the frame sequence under different geographic references is calculated by the dynamic spatial mapping model, and the correlation between the video frame and the geographic range of multiple references is established to obtain the live video with multiple reference geographic range identifiers. Based on live video with multi-reference geographic range identifiers, the system uses intelligent delineation operation combined with the ground feature recognition model corresponding to the sensor type identifier to perform point, line or area trajectory recording along the outline of ground features in the video frame. Based on the coordinates of the trajectory recording and the target geographic reference range corresponding to the video frame, coordinate transformation is performed to generate geographic data with ground feature attributes and recognition confidence. Geographic data with feature attributes and identification confidence levels are transmitted to web maps, mobile terminal maps, and display interfaces via a multi-link network. At the same time, the system receives geographic data modification trajectories from web maps and mobile terminal maps. Based on the geographic data modification trajectories, the feature identification confidence levels corresponding to the modification trajectories, and the coordinate information of the corresponding features in the live video with multi-reference geographic range identifiers, the system performs hierarchical update processing to achieve bidirectional dynamic synchronization and conflict resolution between geographic data and live video.
2. The information processing method based on UAV video spatialization according to claim 1, characterized in that, The process involves collecting camera optical parameters, camera pose information, and sensor type identifiers, synchronously encoding these parameters, information, and identifiers into the non-display data segment of the drone video frame structure, forming an encoded video with multi-dimensional spatial information. This encoded video is then subjected to streaming encapsulation processing, with inter-frame association identifier fields added to create the encoded video stream. This includes: The camera optical parameters, including focal length, aperture, sensor size, and distortion correction coefficient, are obtained in real time through the data acquisition link. Camera pose information is collected through an attitude perception link. The camera pose information includes latitude and longitude parameters, altitude parameters, pitch angle parameters, yaw angle parameters, and roll angle parameters. The sensor type identifier is obtained through the sensor type detection link, and the sensor type identifier includes a visible light sensor identifier or a thermal infrared sensor identifier. The intra-frame embedding rules are determined, which specify the byte start position, data length, parameter arrangement order, and check bit generation method of camera optical parameters, camera pose information, and sensor type identifier in the non-display data segment of the UAV video frame, so that the camera optical parameters, camera pose information, and sensor type identifier form a one-to-one correspondence with the timestamp of the video frame. According to the intra-frame embedding rules, the camera optical parameters, camera pose information and sensor type identifier are written frame by frame into the non-display data segment of the drone video frame structure. A CRC32 check bit is generated for the data written in each frame. After writing is completed, an encoded video with multi-dimensional spatial information is formed. Perform inter-frame association processing on encoded video with multi-dimensional spatial information, extract the difference in camera pose information between adjacent video frames, and add the difference in camera pose information as an inter-frame association identifier field to the non-display data segment of the current video frame; The encoded video with added inter-frame association identifier field is subjected to streaming segmentation processing. The encoded video is divided into a continuous sequence of data packets according to the preset data packet size, and a sequence number identifier, timestamp identifier, and sensor type identifier copy are added to each data packet. The data packet sequence with sequence number identifier, timestamp identifier and sensor type identifier copy is encapsulated using the transmission protocol, and a data fragmentation identifier, transmission window size field and retransmission threshold field are added to form an encoded video stream that can be directly transmitted.
3. The information processing method based on UAV video spatialization according to claim 1, characterized in that, The process of performing scene-specific transcoding adaptation on the encoded video stream involves adjusting the quantization parameters and data encapsulation structure of the video encoding algorithm based on the format requirements of the subsequent decoding end and the image characteristics corresponding to the sensor type identifier, to obtain an adapted video stream suitable for different scenes. This includes: Extract the transmission window size field and sensor type identifier copy from the encoded video stream, and parse the format requirement information of the subsequent decoding end. The format requirement information includes the supported encoding algorithm type, maximum resolution limit, frame rate range and acceptable bit rate fluctuation range. The image characteristics are determined based on the sensor type identifier copy. If it is a visible light sensor identifier, the image characteristics are determined to be high dynamic range characteristics; if it is a thermal infrared sensor identifier, the image characteristics are determined to be temperature gradient sensitive characteristics. Based on image characteristics, the quantization parameters of the video coding algorithm are adjusted. For images with high dynamic range characteristics, the quantization step size in bright areas is reduced to preserve details; for images with temperature gradient sensitivity characteristics, the quantization step size in areas of abrupt temperature changes is reduced to preserve gradient information. Based on the encoding algorithm type in the format requirement information, select the target encoding algorithm. If the decoding end supports multiple encoding algorithms, prioritize the target encoding algorithm that has the highest compatibility with the original encoding video stream algorithm and is adapted to the image characteristics. The data packet sequence of the encoded video stream is decapsulated to separate the original frame unit sequence with the inter-frame association identifier field; According to the encoding rules of the target encoding algorithm and the adjusted quantization parameters, the video pixel data in the original frame unit sequence is re-encoded, while retaining the camera optical parameters, camera pose information, sensor type identifier and inter-frame association identifier fields in the original frame unit. The data packet structure is reconstructed based on the format requirements. The new data packet structure includes an image feature identifier field, a quantization parameter configuration field, an original frame unit retention field, and a decoding priority field. The re-encoded video data, along with the retained camera optical parameters, camera pose information, sensor type identifier, and inter-frame association identifier fields, are encapsulated according to the new data packet structure. An adaptation identifier field is added to the new data packet. The repackaged data packets undergo sequence verification and bitrate fluctuation detection to ensure that the sequence number of the data packets is consistent with the sequence number of the original encoded video stream, and that the bitrate fluctuation is controlled within an acceptable range. After the verification and detection are passed, an adapted video stream suitable for different scenarios is formed.
4. The information processing method based on UAV video spatialization according to claim 1, characterized in that, The process involves performing layered decoding and splitting on the adapted video stream to separate the frame sequence of the drone live video, the camera optical parameters embedded in the frame sequence, the camera pose information, and the sensor type identifier. Combining the image correction model corresponding to the camera optical parameters, camera pose information, and sensor type identifier, a dynamic spatial mapping model is constructed. This model calculates the geographical range of each video frame in the frame sequence under different geographic references, establishing the association between video frames and multi-reference geographical ranges, resulting in a live video with multi-reference geographical range identifiers, including: The adapted video stream is unpacked to remove the adaptation identifier field, transmission control field, and decoding priority field, resulting in an adapted frame unit sequence with an image characteristic identifier field. Layered decoding is performed on the adapted frame unit sequence. The first layer of decoding separates the video pixel data and non-display data segments. The second layer of decoding extracts the camera optical parameters, camera pose information, sensor type identifier and inter-frame association identifier fields from the non-display data segments to obtain the frame sequence of the drone live video and the camera optical parameters, camera pose information and sensor type identifier corresponding to each frame sequence. The corresponding image correction model is called based on the sensor type identifier. If the sensor is identified as a visible light sensor, the atmospheric scattering correction model is called; if the sensor is identified as a thermal infrared sensor, the temperature noise correction model is called. Each video frame in the video frame sequence is input into the corresponding image correction model, and image preprocessing is performed to eliminate the influence of environmental interference factors on pixel data, resulting in a corrected video frame sequence. The pixel size information of each video frame in the corrected video frame sequence is extracted, and combined with the focal length parameter, sensor size parameter and distortion correction coefficient in the camera optical parameters, the horizontal field of view, vertical field of view and distortion correction parameters of the video frame are calculated. By combining the latitude and longitude parameters, altitude parameters, pitch angle parameters, yaw angle parameters, and roll angle parameters in the camera pose information, the shooting origin coordinates, shooting direction vector, and attitude rotation matrix of the video frame are determined. Based on the origin coordinates of the shooting point, and with the horizontal field of view, vertical field of view, distortion correction parameters, shooting direction vector and attitude rotation matrix as constraints, a dynamic spatial mapping model is constructed. The dynamic spatial mapping model includes dual transformation formulas between pixel coordinates and the geodetic coordinate system and the local coordinate system. The pixel coordinate matrix of each corrected video frame is input into the dynamic spatial mapping model, and the geographic coordinates of the four corner points of the video frame in the geodetic coordinate system and the geographic coordinates in the local coordinate system are calculated respectively by the double transformation formula. The geographic reference range corresponding to the video frame is determined by using the geographic coordinates of the four corner points in the geodetic coordinate system as the boundary, and the local reference range corresponding to the video frame is determined by using the geographic coordinates of the four corner points in the local coordinate system as the boundary. For each video frame, the accuracy of the geodetic reference geographic range and the local reference geographic range is verified. The overlap of the calculated geodetic reference geographic range with the geodetic reference geographic range of the adjacent video frame is compared. The overlap of the local reference geographic range with the local reference geographic range of the adjacent video frame is also compared. If the overlap of any reference does not meet the preset association requirements, the transformation formula parameters in the dynamic spatial mapping model are adjusted, and the geographic range of the corresponding reference is recalculated. The verified geodetic reference geographic range information and local reference geographic range information are written into the attribute fields of the corresponding video frames to form frame units with multiple reference geographic range identifiers. The frame units with multiple reference geographic range identifiers are arranged in chronological order, while retaining the adjacent frame association relationship corresponding to the inter-frame association identifier field, to obtain the live video with multiple reference geographic range identifiers.
5. The information processing method based on UAV video spatialization according to claim 4, characterized in that, The image correction model, which combines camera optical parameters, camera pose information, and sensor type identification, is used to construct a dynamic spatial mapping model. This model calculates the geographical extent of each video frame in the frame sequence under different geographic references, including: The distortion correction coefficients in the camera's optical parameters are extracted, and the correction parameters output by the image correction model corresponding to the sensor type identifier are combined to establish a pixel coordinate pre-correction formula. The original pixel coordinates of the video frame are pre-corrected to eliminate the influence of lens distortion on coordinate calculation. Extract latitude and longitude parameters from the camera pose information, convert the latitude and longitude parameters into Cartesian coordinates in the geodetic coordinate system, and use them as the origin coordinates of the dynamic spatial mapping model; The altitude, pitch, and yaw angle parameters are extracted from the camera pose information to construct a three-dimensional vector of the shooting direction. The direction of the three-dimensional vector is determined by the pitch and yaw angle parameters, and the length of the three-dimensional vector is positively correlated with the altitude parameter. By combining the focal length parameter and sensor size parameter in the camera's optical parameters, the pixel field of view of the video frame is calculated. The pixel field of view is used to describe the spatial angle range corresponding to a single pixel. Starting from the origin coordinates, with the three-dimensional vector of the shooting direction as the central axis and the pixel field of view as the unit angle, a spatial ray model corresponding to the pixels of the video frame is constructed, with each pixel corresponding to a spatial ray; Determine the transformation parameters between the geodetic coordinate system and the local coordinate system. The transformation parameters include translation parameters, rotation parameters, and scaling parameters, which are used to realize the mutual transformation of coordinates in the geodetic coordinate system and the local coordinate system. The spatial ray corresponding to each pixel is intersected with the ground reference plane to obtain the geographic coordinates of the pixel in the geodetic coordinate system. The geographic coordinates are then converted to geographic coordinates in the local coordinate system by the transformation parameters. The extreme values of the geographic coordinates of all pixels in the video frame under the geodetic coordinate system are counted to determine the geodetic reference geographic range of the video frame; the extreme values of the geographic coordinates of all pixels in the video frame under the local coordinate system are counted to determine the local reference geographic range of the video frame. The geographic range difference between adjacent video frames is obtained based on the inter-frame association identifier field. The geodetic reference geographic range and local reference geographic range of the current video frame are then smoothly adjusted to make the geographic range transition between adjacent frames continuous. The smoothed geodetic datum geographic range and local datum geographic range are compared with the preset geographic range accuracy threshold. If the geographic range accuracy of any datum is lower than the geographic range accuracy threshold, the calculation parameters of the space ray model are readjusted, and the intersection operation and geographic range determination steps are executed again until the geographic range accuracy meets the standard.
6. The information processing method based on UAV video spatialization according to claim 5, characterized in that, The step of obtaining the geographical range difference between adjacent video frames based on the inter-frame association identifier field, and performing smooth adjustments on the geodetic reference geographical range and local reference geographical range of the current video frame to ensure a continuous transition in the geographical range between adjacent frames includes: Extract the inter-frame association identifier field of the current video frame and parse the difference in camera pose information between the current video frame and the previous video frame; Based on the geodetic reference geographic range and local reference geographic range of the previous video frame, and combined with the difference in camera pose information between the current video frame and the previous video frame, the initial geodetic reference geographic range and initial local reference geographic range of the current video frame are predicted. The difference between the actual geodetic reference geographic range obtained by the dynamic spatial mapping model for the current video frame and the predicted initial geodetic reference geographic range is calculated to obtain the geodetic reference geographic range deviation; the difference between the actual local reference geographic range obtained by the dynamic spatial mapping model for the current video frame and the predicted initial local reference geographic range is calculated to obtain the local reference geographic range deviation. If the deviation of the geodetic reference geographical range is less than the preset smoothing threshold, the actual geodetic reference geographical range is used directly; if the deviation of the geodetic reference geographical range is greater than or equal to the preset smoothing threshold, a weighted average algorithm is used to fuse the actual geodetic reference geographical range with the predicted initial geodetic reference geographical range, and the weight coefficient is determined according to the difference between the camera pose information of the current video frame and the previous video frame. The same deviation judgment and weighted fusion process are applied to the local baseline geographic area to obtain the smoothed local baseline geographic area. Extract the inter-frame association identifier field of the next video frame, parse the difference in camera pose information between the current video frame and the next video frame, and predict the geographical range of the next video frame. Determine the overlap between the smoothed geographical range of the current video frame and the predicted geographical range of the next video frame. If the overlap is lower than the preset overlap threshold, adjust the weighted fusion coefficient of the current video frame and recalculate the smoothed geographical range until the overlap reaches the target. Linear interpolation is performed on the boundary coordinates of the smoothed geodetic datum and local datum to eliminate abrupt changes in the boundary coordinates. The adjusted geodetic reference geographic range and local reference geographic range are compared with the preset geographic range rationality threshold. If the adjusted geodetic reference geographic range or local reference geographic range exceeds the preset geographic range rationality threshold, the inter-frame association identifier field and dynamic spatial mapping model parameters are re-examined, corrected, and then smooth adjustment is performed again. If the geographical range deviation of multiple consecutive frames is greater than the preset smoothing threshold, the camera pose information verification process is triggered to reacquire the camera pose information to correct the calculation basis. After smoothing is completed, the adjusted geodetic reference geographic extent information, local reference geographic extent information, and smoothing parameters are written into the attribute fields of the video frame.
7. The information processing method based on UAV video spatialization according to claim 1, characterized in that, The live video based on multi-reference geographic range identifiers, through intelligent delineation operations combined with the ground feature recognition model corresponding to sensor type identifiers, performs point, line, or area trajectory recording along the outline of ground features within the video frame. Based on the coordinates of the trajectory recordings and the target geographic reference range corresponding to the video frame, coordinate transformation is performed to generate geographic data with ground feature attributes and recognition confidence levels, including: In the playback interface of a live video with multiple reference geographic range identifiers, a reference switching control and geographic range auxiliary lines are loaded. The geographic range auxiliary lines are generated based on the geodetic reference geographic range information or local reference geographic range information in the video frame attribute field and are used to identify the target geographic reference scale corresponding to the pixel coordinates within the video frame. The target geographic reference selection instruction is received through the reference switching control, and the target geographic reference range corresponding to the video frame is determined to be either the geodetic reference geographic range or the local reference geographic range. The corresponding ground feature identification model is called based on the sensor type identifier. If the sensor is identified as a visible light sensor, the visible light ground feature extraction model is called; if the sensor is identified as a thermal infrared sensor, the thermal infrared ground feature temperature feature model is called. Input the currently playing video frame into the corresponding ground feature recognition model, extract the ground feature feature vector within the video frame, the ground feature feature vector includes contour features, texture features or temperature distribution features, and output the ground feature recognition confidence score. The intelligent outlining operation moves the operation trajectory along the outline of the ground features within the video frame. The ground feature recognition model matches the feature feature vector with the pixel coordinates of the operation trajectory in real time. If the matching degree is lower than a preset threshold, a trajectory correction prompt is triggered; if the matching degree is higher than the preset threshold, the operation trajectory continues to be recorded. If the target feature is a point feature, the coordinates of a single point on the operation trajectory are recorded when the operation point stops moving and the matching degree is higher than the preset threshold. If the target feature is a line feature, the pixel coordinates are continuously recorded to form a line coordinate sequence as the operation point moves and the matching degree remains higher than the preset threshold. If the target feature is a surface feature, the pixel coordinate polygon of the closed trajectory is recorded when the operation point moves to form a closed trajectory and the matching degree remains higher than the preset threshold. Extract the geographic coordinate information corresponding to the target geographic reference range from the attribute field of the current video frame, and obtain the geographic coordinates of the upper left and lower right corners of the video frame under the target geographic reference. Calculate the horizontal and vertical geographic span of the video frame under the target geographic reference, and combine the pixel width and pixel height of the video frame to obtain the geographic distance corresponding to a unit pixel under the target geographic reference. Convert the geographic coordinates of the top left corner of the video frame under the target geographic reference to Cartesian coordinates. Based on the geographic distance of a unit pixel under the target geographic reference, the recorded single-point coordinates, linear coordinate sequences, or closed coordinate polygons are converted into Cartesian coordinates. During the conversion process, the Cartesian coordinates of the upper left corner of the video frame under the target geographic reference are used as the reference. The plane offset corresponding to each pixel coordinate is calculated, and the reference Cartesian coordinates are superimposed to obtain the final Cartesian coordinates. Convert the final Cartesian coordinates back to geographic coordinates under the target geographic datum; Add feature attribute information to the converted geographic coordinates. The feature attribute information includes feature type identifier, delineation time information, timestamp information of the corresponding video frame, and feature recognition confidence score output by the feature recognition model. The geographic coordinates with added attributes are associated with the confidence level of feature recognition and stored to generate geographic data with feature attributes and recognition confidence levels.
8. The information processing method based on UAV video spatialization according to claim 7, characterized in that, The process involves intelligent delineation combined with a ground feature recognition model corresponding to the sensor type identifier. Point, line, or area trajectory recording is performed along the ground feature outline within the video frame. Coordinate transformation is then performed based on the coordinates of the trajectory recordings and the target geographic reference range corresponding to the video frame, generating geographic data with ground feature attributes and recognition confidence levels, including: During the playback of live video with multiple reference geographical range identifiers, when the intelligent delineation operation is triggered, the currently playing video frame is locked, the playback of the video frame sequence is paused, and the sensor type identifier of the video frame is extracted. The corresponding ground feature recognition model is invoked according to the sensor type identifier, and the pre-trained ground feature feature template library of the model is loaded. The ground feature feature template library contains standard feature vectors of a variety of common ground features. The locked video frame is input into the land cover recognition model. The pixel region of the video frame is traversed through a sliding window, and the similarity between the feature vector of each window region and the standard feature vector in the land cover feature template library is calculated. Window regions with similarity higher than a preset matching threshold are marked as candidate land cover regions, and the boundary coordinates of the candidate land cover regions and the preliminary determination results of the corresponding land cover types are output. The operation point is moved within the candidate feature area by intelligent outlining. The feature recognition model calculates the consistency between the feature vector of the operation point and the feature vector of the candidate feature area in real time. If the consistency is lower than the preset threshold, a vibration prompt is issued to correct the operation trajectory. If the target feature is a point feature, then when the operation point moves to the center of the candidate feature area and the consistency is higher than the preset threshold, the single-point pixel coordinates of that location are recorded. If the target feature is a linear feature, then when the operation point moves along the center line of the candidate feature area and the consistency is consistently higher than the preset threshold, the pixel coordinates are recorded at the preset sampling interval to form a linear coordinate sequence. If the target feature is a planar feature, then when the operation point moves along the boundary of the candidate feature area and the consistency is consistently higher than the preset threshold, the pixel coordinates of the boundary are recorded to form a closed coordinate polygon. Extract the top-left and bottom-right geographic coordinates of the target geographic reference range from the attribute fields of the locked video frame, and calculate the horizontal and vertical geographic lengths of the video frame. The horizontal geographic length per unit pixel is calculated based on the pixel width of the video frame, and the vertical geographic length per unit pixel is calculated based on the pixel height of the video frame. Convert the geographic coordinates of the top left corner of the locked video frame under the target geographic reference to Cartesian coordinates. For a single pixel coordinate, calculate its horizontal and vertical offset relative to the top-left pixel. Multiply the horizontal offset by the horizontal unit pixel geographic length and the vertical offset by the vertical unit pixel geographic length to obtain the planar offset. Superimpose the top-left corner planar rectangular coordinates to obtain the target planar rectangular coordinates. Convert the target plane rectangular coordinates back to geographic coordinates under the target geographic datum; For linear coordinate sequences and closed coordinate polygons, repeat the above steps of offset calculation, planar coordinate superposition and transformation to obtain the geographic coordinates under the target geographic reference for each coordinate. Add feature attributes to the converted geographic coordinates. The feature attributes include feature type identifier, delineation time, video frame timestamp, and average similarity output by the feature recognition model. The geographic coordinates with added attributes are associated with and stored with the mean similarity value to generate geographic data with feature attributes and identification confidence levels.
9. The information processing method based on UAV video spatialization according to claim 1, characterized in that, The process involves transmitting geographic data with feature attributes and identification confidence levels to web maps, mobile terminal maps, and display interfaces via a multi-link network. Simultaneously, it receives geographic data modification trajectories from the web maps and mobile terminal maps. Based on the geographic data modification trajectories, the feature identification confidence levels corresponding to the modification trajectories, and the coordinate information of corresponding features in the live video with multi-reference geographic range identifiers, it performs layered update processing to achieve bidirectional dynamic synchronization and conflict resolution between geographic data and live video. This includes: Geographic data with feature attributes and recognition confidence are processed by data sharding. The size of the shards is adjusted according to the transmission bandwidth limitations of web maps and mobile terminal maps and the feature recognition confidence. The higher the feature recognition confidence, the smaller the shard is to improve transmission efficiency. Add a data identifier, synchronization version number, ground feature identification confidence copy, and fragment check code to each fragment; Establish multi-link network transmission channels, allocate independent transmission links for web maps, mobile terminal maps and display interfaces, and set link quality monitoring fields for each transmission link; The fragmented geographic data is sent to the web map, mobile terminal map and display interface through a multi-link network transmission channel. The fragmentation transmission progress and link quality monitoring results are fed back in real time during the transmission process. The fragmentation check code is used to verify the integrity of the geographic data fragments received by each end. If there are missing fragments, retransmission is triggered. Receive geographic data modification trajectory feedback from web map and mobile terminal map, wherein the geographic data modification trajectory includes the modified geographic coordinates, modification operation type, modification timestamp and corresponding feature ID; Based on the feature ID, the identification confidence of the corresponding feature is queried. The modified trajectory is divided into a first processing group and a second processing group according to the identification confidence of the corresponding feature. The feature identification confidence of the modified trajectory is higher than the preset confidence threshold and is assigned to the first processing group. The feature identification confidence of the modified trajectory is lower than or equal to the preset confidence threshold and is assigned to the second processing group. Extract the first geographic data modification trajectory from the first processing group, and search for the corresponding original geographic coordinates, geographic feature recognition confidence, and associated video frame identifier in the geographic data with geographic feature attributes and recognition confidence based on the geographic feature identifier ID; Based on the video frame identifier, locate the target video frame in the live video with multiple reference geographic range identifiers, extract the target geographic reference range information of the target video frame, and calculate the pixel coordinates corresponding to the modified geographic coordinates under the target geographic reference. Determine whether the modified geographic coordinates are within the target geographic reference range of the target video frame. If the modified geographic coordinates exceed the target geographic reference range of the target video frame, send an invalid modification prompt to the feedback end and move the modification trajectory to the second processing group; if the modified geographic coordinates are within the target geographic reference range of the target video frame, retain the geographic data modification trajectory. Extract the modified geographic coordinates from the retained geographic data modification trajectory, update the original geographic coordinates of the corresponding geographic features in the geographic data with feature attributes and identification confidence, and update the synchronization version number at the same time; In the target video frame, the outline corresponding to the original geographic coordinates is deleted, and the outline is redrawn according to the calculated pixel coordinates to complete the update of the object coordinate information in the target video frame. The updated target video frames and updated geographic data with feature attributes and identification confidence levels are synchronized to web maps, mobile terminal maps and display interfaces. The version numbers are synchronized to keep the data versions consistent across all devices. After the first processing group is completed, repeat the above steps to process the geographic data modification trajectory in the second processing group. If the modification trajectory in the second processing group conflicts with the updated data, calculate the modification timestamp of the conflict trajectory and the update timestamp of the updated data, retain the modification result of the timestamp update, and discard the conflict trajectory with the oldest timestamp. After all modified trajectories are processed, a two-way synchronization completion marker is generated and fed back to the web map, mobile terminal map, and display interface, realizing two-way dynamic synchronization and conflict resolution between geographic data and live video.
10. An information processing system based on UAV video spatialization, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the information processing method based on UAV video spatialization according to any one of claims 1 to 9 by executing the machine-executable instructions.
Citation Information
Patent Citations
Method and system for forming view field projection map based on unmanned aerial vehicle video
CN116821414A
Unmanned aerial vehicle and method based on satellite internet communication
CN119921834A