Multi-view VR video acquisition and synthesis method for light and small unmanned aerial vehicle
By establishing a unified time baseline and dynamic visibility probability field, and performing multi-camera clock fine-tuning and photometric consistency mapping, the problem of black frame skipping in occluded areas during multi-view VR video acquisition and synthesis by small and lightweight drones was solved, achieving seamless stitching and a stable immersive experience.
Patent Information
- Application Number
- CN202511458601.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-23
AI Technical Summary
Existing methods for capturing and synthesizing multi-view VR videos using lightweight drones cannot achieve smooth transitions in high-speed dynamic scenes, resulting in failed image fusion, black frame skips, or momentary interruptions, which disrupt the user's sense of spatial continuity and visual stability.
By establishing a unified time baseline, obtaining the boundary event sequence, constructing a dynamic visibility probability field, performing multi-camera clock fine-tuning, generating a seamless stitched trajectory, and using photometric consistency mapping and gamma backpropagation technology, black frame skipping in occluded areas is eliminated.
It achieves synchronous response for image acquisition with millisecond-level precision, ensuring geometric coherence and photometric continuity in occluded areas, enhancing the user's immersion and spatial perception stability in virtual reality devices, and avoiding dizziness and illusions.
Smart Images

Figure CN121397368A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video acquisition and processing, and particularly relates to a multi-view VR video acquisition method for a light small unmanned aerial vehicle. BACKGROUND
[0002] The multi-view VR video acquisition for a light small unmanned aerial vehicle refers to acquiring video data of a target scene from different angles and different directions in the air at the same time by using a light and flexible small unmanned aerial vehicle, and then performing splicing and synthesis processing on the acquired images through multi-view video registration, synchronization and fusion technology to generate virtual reality (VR) video content with immersion and stereoscopic effect. This process not only solves the problem of limited view angle of traditional fixed camera equipment by using the maneuverability of the unmanned aerial vehicle, but also realizes panoramic coverage and real restoration of the scene through the fusion of multi-view data, and finally outputs immersive video for VR devices, which is widely used in fields such as travel and tourism display, film and television shooting, emergency drill and teaching training.
[0003] The prior art has the following disadvantages: In the prior art, multi-view video splicing of a light small unmanned aerial vehicle usually relies on the boundary area between different camera positions for image fusion. However, when the shooting object is located at the overlapping boundary of multiple camera positions in a high-speed dynamic scene and enters the occlusion area at the moment, the splicing algorithm in the prior art often cannot realize smooth transition, which easily leads to failure of image fusion, manifested as black frame skipping or instantaneous flashing. Although the probability of this phenomenon is extremely low, once it is magnified in a virtual reality device, it will directly destroy the user's spatial continuity and visual stability, causing the immersive experience to collapse instantly, and even causing dizziness and cognitive confusion.
[0004] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The purpose of the present application is to provide a multi-view VR video acquisition method for a light small unmanned aerial vehicle to solve the problems in the background.
[0006] In order to achieve the above purpose, the present application provides the following technical solution: a multi-view VR video acquisition method for a light small unmanned aerial vehicle, comprising the following steps: acquiring a unified time baseline, establishing an overlapping boundary time sequence index, locking a zero-phase anchor point, and outputting a millisecond-level boundary event sequence; under the constraint of the boundary event sequence, acquiring a boundary occlusion priori, constructing a dynamic visibility probability field and labeling a potential occlusion kernel; Under the support of dynamic visibility probability field, cross-view delay gradient is obtained, potential occlusion core area is constrained, phase traction exploration is injected into the boundary event sequence, and multi-camera clock fine-tuning is completed based on the feedback result; Under the condition of clock fine-tuning completion, boundary geometry field is obtained according to cross-view delay gradient and optical flow residual, geometric restriction is set for potential occlusion core area, seamless splicing track is generated, and critical transition window is preset on unified time baseline; In the critical transition window, based on the boundary geometry field, the luminosity consistency mapping is obtained, the color bidirectional mapping and gamma inverse strategy are adopted, and the local color compensation is performed on the potential occlusion core area, and the luminosity continuity condition is established; Under the support of luminosity continuity condition, time reversal phase gate control instruction is obtained, programmable polarization super surface is driven to inject phase conjugate frame envelope, reverse energy channel is constructed on unified time baseline, black frame extinguishing is triggered for potential occlusion core area, and calibration factor is written back to unified time baseline.
[0007] Preferably, the boundary event sequence output step is as follows: The main control unmanned aerial vehicle integrated with high-stability constant temperature crystal oscillator chip and global navigation satellite receiving module is selected as time reference source, and satellite time pulse is received for calibrating local clock signal; The rest of the unmanned aerial vehicles receive time synchronization information through short wave frequency hopping communication link, perform smoothing correction operation after receiving multiple continuous synchronization signals, and unify the time reference of all unmanned aerial vehicles to unified time baseline; When each unmanned aerial vehicle collects image frames, the flight speed, heading angle, pitch angle and three-axis attitude angle information are recorded synchronously, and the boundary time sequence index is established combined with image frame timestamp; According to the boundary time sequence index, edge feature extraction and similarity matching are performed on the synchronous image frames, the image frames with stable motion vector changes in the image frame group are selected as zero-phase anchor points, and the boundary event sequence is constructed as the input of dynamic occlusion modeling.
[0008] Preferably, the dynamic visibility probability field construction step is as follows: According to the boundary event sequence, the time and space information of target entering, occlusion and leaving in the image frame is extracted, the image sequence in the time range before and after each event node is analyzed, the occlusion trend of the target object is obtained, and the occlusion prior information is output; Based on the occlusion prior information, a two-dimensional grid probability map is established, the grid elements on the target motion path are subjected to probability attenuation, and the dynamic visibility probability field is formed by real-time updating; The region with probability value lower than the threshold and meeting the shape stability is identified as potential occlusion core area, and the occlusion persistence is verified through comparative frame luminosity difference analysis, and the effective occlusion core is confirmed; The effective occlusion kernel is mapped back to the image frame as the basis for subsequent differential processing of geometric modeling and photometric correction.
[0009] Preferably, the multi-camera clock fine-tuning step based on the feedback result is as follows: According to the position offset of the occlusion kernel region in different view images, a cross-view delay gradient matrix is calculated, and the offset is converted into a time delay value combined with the target speed mapping coefficient; According to the occlusion probability and area of the occlusion kernel region, a weighted influence factor is calculated to give priority to calibration in the delay gradient matrix; According to the delay gradient matrix, a plurality of phase traction probes are injected into the boundary event sequence, and the optimal time offset is selected to generate a target time correction table through image fusion feedback; According to the target time correction table, the clock of the multi-rotor unmanned aerial vehicle is fine-tuned, and whether the frame-level synchronization accuracy meets the splicing requirements is verified through two rounds of fusion tests.
[0010] Preferably, the critical transition window presetting step is as follows: According to the delay gradient and optical flow residual, edge pixel points are extracted, and a boundary geometric field containing spatial coordinates, time offset and residual intensity is established; Based on the outline coordinates of the occlusion kernel region, an occlusion restriction body is constructed to limit the splicing trajectory from crossing the occlusion kernel region and control the curvature variation range; According to the minimum optical flow residual and delay path, a trajectory point is selected, combined with smoothing filtering to generate a continuous splicing trajectory that does not cross the occlusion kernel, and a critical transition window is marked around it for subsequent photometric and texture processing.
[0011] Preferably, the photometric continuity condition establishment step is as follows: The pixel regions on both sides of the splicing trajectory are extracted, the three-channel gray scale distribution is counted, and a color mapping index matrix is constructed; Based on the mapping matrix, bidirectional color correction is performed, combined with the exposure characteristics of the image acquisition device for gamma inverse processing to ensure the brightness continuity of the splicing area; In the occlusion kernel region, classification color compensation is performed according to the color statistical characteristics of the outer ring pixels, and a ring-shaped mixed transition zone is constructed to realize color smooth transition and photometric consistency.
[0012] Preferably, in the color compensation of the occlusion kernel region, the average brightness difference of the ring-shaped edge region is used to perform pixel brightness enhancement, and the red, green and blue channel proportion of the pixels in the occlusion kernel is adjusted through a color temperature rotation matrix, so that the compensation result is consistent with the surrounding image in brightness and color temperature.
[0013] Preferably, under the support of photometric continuity conditions, the time reversal phase gating instruction is obtained to drive the programmable polarization super surface to inject the phase conjugate frame envelope, to construct the reverse energy channel on the unified time baseline, to trigger the blackout of the blocked nuclear region black frame skipping, and to write back the calibration factor to the unified time baseline step as follows: Under the condition of photometric continuity, by monitoring the brightness change and color channel change rate of the image frame, the frame skipping or black frame early warning area is identified, and the gating instruction is output according to the unified time baseline, indicating the compensation time interval, phase direction and region position; Under the control of the gating instruction, the polarization control unit generates a phase conjugate frame envelope, collects the pixel values of the image frames before and after the target region, constructs the target image structure and applies a reverse phase rotation to form a compensation image; Draw the reverse frame timing channel on the unified time baseline, allocate the compensation frame dedicated time slot, mark the frame envelope content in reverse order, and write it into the channel frame by frame to form a continuous compensation sequence; Read the frame envelope compensation image, perform brightness interpolation and texture redrawing on the frame skipping area, embed the compensation image into the original frame through weighting, and synchronously adjust the color temperature difference and gamma factor to ensure natural visual transition; Write the time stamp, brightness correction coefficient, color mapping matrix and interpolation parameter involved in the compensation process back to the unified time baseline to complete the image repair information binding and dynamic control link closure.
[0014] In the above technical solution, the technical effects and advantages provided by the present application are as follows: The present application constructs fine timing control through unified time baseline to ensure that multi-camera image acquisition realizes synchronous response under millisecond level precision; through constructing dynamic visibility probability field and delay gradient model, accurate identification and dynamic tracking of the blocked area are realized; through multi-camera clock fine tuning and boundary geometry constraint, seamless splicing track is generated, so that the image still maintains geometric continuity in the blocked nuclear region; through the introduction of photometric consistency mapping, color compensation and gamma inverse technology, color smooth transition in the boundary region is ensured; finally, through the programmable polarization super surface injection phase conjugate frame, the reverse energy channel is constructed, the black frame skipping is eliminated in a directional manner, and the dynamic control link is closed, which fundamentally eliminates the frame blackout and flashing phenomenon caused by the blocked area. This scheme ensures the continuity of the picture while significantly enhancing the user's immersion and spatial perception stability in the virtual reality device, avoiding dizziness, illusion or cognitive disorder caused by visual mutation. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description only represent some embodiments of the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0016] Figure 1 A method flowchart of the light small unmanned aerial vehicle multi-view VR video acquisition and assembly method of the present application. DETAILED DESCRIPTION
[0017] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art.
[0018] The present application provides a light small unmanned aerial vehicle multi-view VR video acquisition and assembly method as shown in Figure 1 The method comprises the following steps: A unified time baseline is obtained, an overlapping boundary time sequence index is established, a zero-phase anchor point is locked on the unified time baseline, and a millisecond-level interface event sequence is output, which is used to provide a basic condition for boundary occlusion modeling; To realize high-precision splicing and occlusion identification processing of multi-view images by multiple light small unmanned aerial vehicles in a high-speed dynamic scene, a unified time baseline needs to be established and an interface event sequence for modeling needs to be output, and the specific steps are as follows: A master unmanned aerial vehicle in a network control center is selected as a time reference source, which is internally integrated with a high-stability constant-temperature crystal oscillator chip and a GNSS global navigation satellite receiving module. The module receives a satellite time pulse once per second, which is used to calibrate the local clock signal. The slave unmanned aerial vehicles performing shooting tasks receive time synchronization information broadcast by the master unmanned aerial vehicle through a short-wave frequency hopping communication link every 50 milliseconds. The information carries a millisecond-level timestamp, a frequency deviation correction value and a synchronization reference flag in the form of 32-bit binary. After receiving, each unmanned aerial vehicle immediately suspends its internal timer, corrects the difference through a counter, and performs smooth correction after receiving three consecutive synchronization signals to avoid clock instability caused by instantaneous jumps. Finally, the time references of all unmanned aerial vehicles are unified to the same timeline, with a maximum deviation of less than 0.5 milliseconds. Once the unified time baseline is established, it will serve as a global time anchoring standard for all image frame acquisition, spatial modeling and event extraction.
[0019] Each UAV records the flight speed, heading angle, pitch angle and three-axis attitude angle information through its inertial navigation unit while collecting image frames, and binds them one by one with each image frame. Using a three-dimensional coordinate transformation model, the shooting field of view of the image frame is projected onto the ground surface to form a quadrilateral projection area, and the boundary overlapping area between adjacent UAVs is identified by comparing the vertex space coincidence rate and the central angle. After detecting an area where the image boundary overlaps by more than 50%, mark this time as the overlapping boundary key frame. Subsequently, according to the time stamp of each UAV image frame, combined with the image boundary taken at the same time by adjacent UAVs, a bidirectional queryable boundary time sequence mapping table is established. The mapping table contains information such as time stamp, view number, overlapping area pixel position range, space angle, coincidence rate, etc., with a precision of one time slice every 50 milliseconds, and is serialized and sorted with time as the main index field. This boundary time sequence index will be used to quickly retrieve the overlapping relationship of the view fields of multiple UAVs at the spatial boundary at any time.
[0020] Based on the established boundary time sequence index, select the image frame combination with the most sufficient spatial overlap and the smallest time error in each group of synchronized frames as the candidate image set. Use the joint algorithm based on SIFT (Scale Invariant Feature Transform) and SURF (Speeded Up Robust Features) to extract edge feature points in each image boundary area. Calculate the image similarity score by the number of feature matches, spatial residual distance and angle difference. Then, the image gray gradient direction distribution map is convolved with each other to judge the structural consistency between the boundary images, and the image frame group with the maximum matching degree and the smallest frame jitter at a specific time stamp is selected. This time stamp corresponds to a specific millisecond value on the unified time baseline, and is defined as the zero-phase anchor point of the boundary area. To ensure the stability of the anchor point, the system needs to compare three time periods before and after the point, ensuring that the motion vector changes between the image frames before and after the anchor point are in a steady state, and there is no image blur or shooting angle mutation phenomenon.
[0021] With the zero-phase anchor point as the reference frame, image frames within ±500 milliseconds before and after the anchor point are selected as the analysis window, and frame-by-frame difference processing is performed on the image sequence within the time window. By comparing the image gray value change pixel by pixel, the edge profile of the moving object is extracted and a binary mask image is generated. Then, combined with the changes in the flight attitude and the boundary projection overlap area of each frame, the specific frame sequence of the target object entering the boundary, being blocked at the boundary, and leaving the boundary is identified. For each event frame, the millisecond timestamp under the unified time baseline, the event type (entering, blocking, leaving), the pixel position of the target in the image, the corresponding angle range of the boundary, the target speed vector and displacement direction, and other attributes are recorded. All events are arranged in chronological order into a structured boundary event sequence and output as a structured data list, serving as the basic input for subsequent dynamic occlusion modeling. To ensure the reliability of the data, the event sequence also needs to undergo at least one redundancy detection to eliminate image frames with image blur, defocus, sensor overexposure, or severe field-of-view occlusion, ensuring that the final boundary event sequence has integrity, continuity, and high-resolution dynamic description capabilities.
[0022] Under the constraint of the boundary event sequence output on the unified time baseline, the boundary occlusion prior is obtained, and a dynamic visibility probability field is constructed, in which potential occlusion kernels are labeled for differential processing of occlusion risk areas in subsequent analysis. After completing the establishment of the unified time baseline and the output of the boundary event sequence, the occlusion trend needs to be further identified based on the sequence and the visibility model needs to be built, thereby providing targeted prediction basis for stitching stability control and subsequent processing. The specific steps are as follows: Based on the output boundary event sequence, the specific time and corresponding spatial position of each target entering, blocking, or leaving the image boundary region are included. Based on this time sequence, the image frames within ±400 milliseconds before and after each event node are taken as the occlusion judgment window, and the image sequence within the window is analyzed continuously. Each image frame is accompanied by flight attitude parameters after being captured by the unmanned aerial vehicle acquisition device, including three-axis rotation angle, inertial velocity, geographic position coordinates, and overhead field-of-view direction. By mapping the imaging area of each image frame to the ground projection coordinates, combined with the pixel position, edge profile, and speed vector of the target object in the image, its real physical path trajectory and motion direction can be calculated. Based on this, the trajectory continuity of the moving target within the image boundary is analyzed, the inclination angle, displacement speed change rate, and image center of mass movement pattern of the trajectory are extracted, and the moving path with stable occlusion trend is identified. At the same time, the acceleration characteristics of the target before entering the boundary region are analyzed, and the occlusion probability trend is calculated. The occlusion prior information is finally output in table form, including the unique number of the target object, the starting occlusion time, the predicted occlusion end time, the moving path trajectory, the speed trend graph, and the corresponding boundary position block.
[0023] Based on the image resolution, the boundary region is divided into a two-dimensional square grid unit composed of 10x10 pixels, and each grid unit represents a spatial sampling unit in the image. On each grid, a variable probability value P is established, initially set to 0.95, indicating complete visibility. Under the guidance of occlusion prior information, the predicted occlusion path of each target object is fitted, and the P value is attenuated on the grid unit through which its trajectory passes, with the lowest value being 0.15, indicating a high occlusion risk. During the image frame sequence advancement process, the grid probability state is updated every 40 milliseconds. The update algorithm is based on the deviation between the actual observation results of the previous frame and the occlusion prediction: if the target actually appears in a certain grid unit, the P value of that unit is adjusted up by 0.1 in the next round of update, otherwise it is adjusted down by 0.1. Through continuous dynamic correction, a two-dimensional dynamic visibility probability map with time memory and trajectory inertia correction capability is finally formed. The probability map fully covers the image boundary region and reflects the visibility of the target object at different times and different positions, which is the basis for subsequent occlusion kernel judgment.
[0024] The constructed dynamic visibility probability map is analyzed, and the P value of a region with more than 200 pixels and a region with a P value less than 0.3 is identified as a preliminary occlusion abnormal region. Further boundary shape analysis is performed on each candidate region: if its edge presents a closed or approximately closed structure, and its center is in a stable motion trend in the past 500 milliseconds (i.e. the centroid change is less than 5 pixels), it is determined as a potential occlusion kernel region.
[0025] To ensure that the labeled occlusion kernel is representative, the region also needs to be verified by comparing frames. Select an image frame 100 milliseconds apart from the current time frame and perform photometric difference analysis on its occlusion kernel position. If the brightness change of the region exceeds 30 gray levels, it indicates that the occlusion degree has stability and continuity in time sequence, and the region is finally confirmed as an effective occlusion kernel. For each confirmed occlusion kernel region, record its start and end frame time, horizontal and vertical coordinate range, centroid position, shape contour and corresponding target number. The occlusion kernel information will provide clear position reference in the stages of stitching, geometric fitting, color correction and frame compensation processing.
[0026] The mapping result is mapped back to the original image frame to ensure that the spatial position of the occlusion core corresponds to the time frame number, and a differentiated subsequent processing strategy is formulated accordingly. In the optical flow calculation stage, the traditional flow field fitting method based on the mean value of the entire frame is disabled in the occlusion core area, and a local pixel block iterative estimation method is used instead to enhance the accuracy of edge transitions. In the geometric field construction stage, the contour boundary of the occlusion core is used as a hard constraint condition for geometric field fitting to prevent incorrect splicing connections. In the photometric mapping process, the local color statistics histogram smoothing method is used in the occlusion core area instead of global mapping curves to avoid local brightness jumps caused by occlusion. In addition, before image fusion, the occlusion core area will be reconstructed into a neutral transition frame in an interpolation synthesis manner to ensure that the final splicing result has visual continuity.
[0027] Under the support of dynamic visibility probability field, cross-view delay gradient is obtained, and weighted constraint is applied to potential occlusion core area, phase traction exploration is injected into the boundary event sequence output on the unified time baseline, and based on the feedback result, multi-camera clock fine tuning is completed, so as to ensure the timing consistency of the occlusion area and the non-occlusion area; To realize the timing accurate alignment of the occlusion area and the non-occlusion area in multi-view image acquisition, delay gradient calculation, clock difference weighting, boundary event exploration adjustment and multi-camera clock fine tuning are needed under the support of dynamic visibility probability field, which includes the following steps: In the last step, the dynamic visibility probability field containing occlusion trend prediction has been constructed, and the potential occlusion core area has been labeled at the pixel level. Based on the probability field, first, select the image frames of the same occlusion core area taken by different unmanned aerial vehicles under the same time baseline, extract the gradient direction distribution graph, gray variation trend graph and texture frequency feature graph of the edge contour in the image. Then, compare these feature graphs with the previous frame and the next frame, calculate the target position offset value of the occlusion core area in different view images through pixel-level alignment. After image scaling uniform processing, the offset value is converted into pixel displacement distance, and then divided by the image mapping coefficient of the target motion speed, and finally the time delay value in milliseconds is obtained. Repeat the above calculation process for the image frames of the occlusion core area taken by each pair of cameras to form a two-dimensional matrix, each element of the matrix corresponds to the time offset value of the occlusion core area between the two unmanned aerial vehicles, and the matrix is the cross-view delay gradient matrix.
[0028] After obtaining the cross-view delay gradient matrix, the control priority of different regions in time fine-tuning is determined according to the occlusion probability value of each occlusion kernel region in the dynamic visibility probability field. The specific method is: for each occlusion kernel region marked in the image, the center point coordinates, the area pixel number and the occlusion probability value of the region are extracted; the probability value is normalized to a weight coefficient, and multiplied by the area of the region to obtain its comprehensive influence factor. This factor will be used to adjust the sensitivity of the corresponding element in the delay gradient matrix. For example, if the area of a certain occlusion kernel region in the image is 400 pixels and the occlusion probability is 0.85, the influence factor is 340, which will be used to increase the weight of the region in the clock adjustment calculation, and give higher calibration priority to the corresponding element in the matrix. In all occlusion kernel regions, the top 10% of the regions in the influence factor ranking will be preferentially sampled in the subsequent boundary event adjustment process to construct the phase traction probe sequence.
[0029] In order to ensure that the time alignment of the occlusion kernel region has practical splicing improvement significance, phase disturbance probes need to be performed on the boundary event sequence on the unified time baseline. In this step, several UAV pairs with the largest offset value in the delay gradient matrix are selected, and the corresponding image frames of the occlusion kernel regions are combined to perform micro phase disturbance under the unified time baseline. The disturbance method is: based on the current event node timestamp, ±0.2ms, ±0.5ms and ±0.8ms are respectively applied to construct six groups of image frame versions. Each group of versions is used for image fusion simulation test, and the improvement degree of fusion quality is evaluated by calculating the pixel-level gradient continuity error, gray residual cumulative value and photometric consistency deviation in the occlusion kernel boundary region. The evaluation scores are summarized to form a feedback curve, and the optimal time offset is determined according to the extreme value position of the curve. The optimal offset of all UAV corresponding occlusion kernel regions is filled into the delay matrix to complete the target time correction table for the occlusion kernel region.
[0030] After the generation of the target time correction table, a clock adjustment operation is performed for each UAV. The adjustment process relies on the internal capacitor frequency adjustment mechanism of each on-board timer. Each fine adjustment has a minimum step size of ±0.1 milliseconds, and the image acquisition trigger time point is changed by controlling the slight drift of the crystal oscillator frequency. To avoid triggering new timing mismatches in the non-occluded areas, the adjustment process needs to consider the average value of the global synchronization deviation of the non-occluded areas at the same time, and the average smoothing adjustment method is used for the non-occluded areas to ensure that the time error does not exceed ±0.3 milliseconds. After the initial clock adjustment is completed, a new set of image frames is collected, and a fusion test is performed again on the occluded core areas. If the fusion score does not reach the set threshold (such as a gray error average of less than 12 and a boundary overlap degree of greater than 92%), a second fine adjustment is performed. After completing two rounds of clock fine adjustment, if all the occluded core areas meet the fusion quality standards, all the UAVs can be considered to have frame-level synchronization accuracy under the current time baseline, and can enter the next stage of the fitting and splicing trajectory generation process of the boundary geometric field.
[0031] Under the condition that the multi-camera clock fine adjustment is completed, the boundary geometric field is obtained according to the cross-view delay gradient and the optical flow residual, and the geometric restriction is set in the potential occluded core area to generate a seamless splicing trajectory, and a critical transition window is preset on the unified time baseline to ensure that the potential occluded core does not trigger geometric misplacement and provide stable support for photometric mapping; Under the premise that the multi-camera clock fine adjustment is completed and the image time alignment establishes frame-level synchronization, in order to ensure the geometric continuity of the image splicing boundary in the occluded core area and the stable support of photometric mapping in the splicing process, the boundary geometric field needs to be constructed according to the delay gradient and the optical flow residual, the occlusion restriction needs to be applied, and the splicing trajectory needs to be generated. The specific steps are as follows: In the last step, a unified time baseline has been established based on multi-camera image acquisition, and frame-level alignment has been achieved through time correction of the occluded core area. In order to further map the time consistency to the image geometric space, two key inputs need to be used to construct the boundary geometric field: one is the delay gradient matrix between multi-view images, and the other is the optical flow residual graph between adjacent image frames.
[0032] During the execution, first, the edge pixel points of the same target area of each pair of adjacent cameras are extracted from the image frames under the unified time baseline, and their corresponding time offset values and spatial displacement vectors are recorded. Then, all the edge points are spatially registered in the image space, and the point set is labeled with a weight according to the time delay value. This weight represents the relative offset of the time of appearance of the edge point in the image.
[0033] In parallel, the optical flow method is used to calculate the motion vector of each pixel in the adjacent frame image, and the vector residual between the actual motion and the expected motion is obtained by subtracting the theoretical translation trajectory, and the optical flow residual field is obtained. Superimpose the optical flow residual values on the edge point set to establish a three-dimensional boundary point data set, each point containing spatial coordinates, time offset value and residual intensity. Based on the point set, a boundary geometry field is constructed to form a geometric expression structure with spatial curvature and time response characteristics, providing a stable reference for generating a stitching trajectory.
[0034] Because the occlusion core region is often accompanied by image structure loss, local distortion enhancement and brightness discontinuity, if the stitching path passes through the center of the occlusion core, it is easy to cause boundary mismatch, ghosting, jumping, and abnormality of the stitching seam, so a clear geometric restriction must be set in this area. The specific operation is as follows: according to the occlusion core coordinate data output in the previous step, the outer edge of the corresponding occlusion region in the image is extracted frame by frame to form a closed polygon occlusion area profile. Map the profile to the boundary geometry field coordinate system to establish an occlusion restriction body in the form of a three-dimensional volume. The restriction body has a height restriction, that is, it is prohibited for the stitching trajectory to pass through its internal core region, and at the same time, a minimum curvature change constraint is added within the envelope radius range of its edge to avoid excessive twisting of the trajectory.
[0035] In addition, define a 10-pixel range outside the edge of the occlusion restriction body as a boundary transition zone. In this transition zone, the stitching trajectory needs to meet two additional conditions: first, the trajectory point set is not allowed to deviate more than 0.2 milliseconds in the time dimension, and second, the spatial connection curvature cannot exceed 1.5 times the curvature change amplitude of the previous point. The trajectory points in this region will preferentially use the pixel points with lower delay gradient and minimum optical flow residual to enhance the boundary stability of the stitching result.
[0036] The stitching trajectory generation process is based on all edge points in the boundary geometry field to construct an image connection path. Starting from the minimum distortion point in the image overlap area, the pixel-level point selection is performed along the path with the minimum optical flow residual, and the fitting trajectory is connected in turn, and the trajectory needs to pass through the point set of the minimum delay path to form a stitching path with both time consistency and geometric continuity. When passing through the edge transition zone of the occlusion restriction body, the trajectory is automatically subjected to second-order smoothing filter processing to remove high-frequency jitter points on the trajectory, and the trajectory connection direction change angle is calculated in real time, and if it exceeds the limit, the previous trajectory point is replaced until a complete, continuous, curvature stable and non-crossing occlusion core region path is generated.
[0037] After the stitching trajectory is determined, its time sequence under the unified time baseline is taken as a reference, and the image frame data within a range of 200 milliseconds on both sides is extracted correspondingly, and the corresponding pixel coordinate region in the image space is calibrated to form a critical transition window. This window is a special structure area reserved for photometric correction and texture consistency processing, which contains the image edge area of 20 pixels on both sides of the stitching trajectory.
[0038] In this window, the subsequent lightness mapping stage will perform a local brightness equalization process, that is, by comparing the histogram distribution of the two side images, the pixel value distribution is linearly aligned within the gray channel, and a gamma adjustment factor is used to generate a brightness gradient buffer layer on the trajectory path, so that the brightness difference between the left and right images gradually reduces to within 5 gray units from the edge to the center, ensuring visual continuity.
[0039] In addition, this window is also used for texture reconstruction processing of image structure edges. Through edge enhancement algorithm and local texture compensation mechanism, the details of the image at the splicing place are redrawn, the texture density consistency and structure continuity are maintained, and the texture breaking or sawtooth effect caused by the splicing trajectory is avoided.
[0040] In the critical transition window, based on the boundary geometry field, the lightness consistency mapping is obtained, the color bidirectional mapping and gamma inverse strategy are used, and the local color compensation is performed for the potential occlusion core area, the color correction of the boundary area is completed, which is used to establish the lightness continuity condition for time reversal control; On the basis of having completed the boundary geometry field construction and the critical transition window setting, in order to ensure that the image splicing area has continuity in the lightness dimension, it is necessary to obtain the lightness consistency mapping, perform color correction and occlusion core compensation, which includes the following steps: On the premise that the splicing trajectory has been derived from the boundary geometry field and its direction along the image edge has been determined, a region with a width of 20 pixels is extracted from both sides of the splicing trajectory as an analysis window. Among them, the left window corresponds to the pixel band of the left view image near the splicing boundary, and the right window comes from the right view image at the same position. For the two pixel band regions, the gray distribution data of the red, green and blue three channels are counted respectively, and the color histogram with more than 1000 pixels is constructed. On the basis of the histogram data, the mean, standard deviation, maximum and minimum of the color distribution are calculated, and the cumulative distribution function curve of each channel is generated. Next, taking the color distribution of the left view as the lightness reference benchmark, each color value in the right view is mapped to the color space of the left view, and by constructing a set of color correspondence table, the mapping relationship at the pixel level is realized. Each entry records a pair of RGB values and offset of left and right color values, forming a three-dimensional color mapping index matrix. This matrix covers both sides of the splicing trajectory, which is used for color transition and brightness equalization calculation in the subsequent stage, and is the core basis for realizing pixel-level lightness consistency control.
[0041] To prevent the color block fault or brightness jump problem of the splicing boundary, color bidirectional mapping and gamma inverse processing are performed in the junction area between the left and right images. First, using the color mapping index matrix described above, 200 key pixel points are extracted on both sides of the splicing track, and the numerical difference of each pair of points in three channels is calculated. For each channel, an incremental mapping function is constructed, so that the right image color value is close to the corresponding value of the left image after weighted correction, and vice versa, finally forming a color correction path of bidirectional convergence. Then, the gray value of the corrected pixel points is subjected to gamma inverse operation. With reference to the exposure characteristic curve of the image acquisition device, the pixel brightness value of the splicing track area is applied to the exponential inverse transformation, so that the brightness change is more natural in vision. For example, if the average gamma value of the left view splicing edge is 2.1 and the right view is 1.9, a smooth gamma interpolation curve from 2.1 to 1.9 is established at the center line of the track, and is applied to the overlapping area of the left and right images. On the basis of the gamma interpolation result, the splicing area is subjected to pixel-by-pixel remapping, and each pixel value is adjusted by bidirectional mapping and gamma inverse operation, to ensure the continuity of the boundary brightness curve. This processing method can significantly reduce the visual jump problem caused by photometric mutation, improve the naturalness of the boundary transition, and establish a visual uniform basis for subsequent occlusion area processing.
[0042] Due to the lack of structural information, exposure abnormalities and other reasons, the occlusion core area is often accompanied by color distortion, brightness drift and even image fault phenomenon, so color compensation operation needs to be performed locally to ensure the overall picture consistency. First, the edge closed contour line of the occlusion core area is extracted from the dynamic visibility probability field, forming a specific coordinate closed area in the image. Then, the annular edge area within 5 to 15 pixels outside the contour is collected, and the average value of the three-channel color, brightness distribution and color temperature trend are counted. The annular area is regarded as a reference sample band for color compensation, and its color statistical characteristics represent the color environment that the occlusion core should present in the normal state. Subsequently, classification compensation strategies are performed according to the color loss types of the image pixels inside the occlusion core. For example, when the gray value of the occlusion area is overall dark, the average brightness difference of the reference sample band is added to each pixel to realize the brightness lifting; when the image saturation is abnormally reduced, nonlinear gain amplification is performed according to the maximum saturation value of the sample band to restore the color to a natural density; when the color temperature appears a blue or yellow trend, the RGB channel ratio of the pixels in the occlusion core is adjusted through the color temperature rotation matrix to make it consistent with the color temperature direction of the sample band. In addition, to avoid the hard boundary of the compensation area, a radial progressive color diffusion processing is performed on the edge of the occlusion core, i.e. the compensation reference value is gradually decreased to the original value pixel by pixel, to build a ring-shaped mixed transition zone. This mixing method avoids edge breakage and color splitting in the color repair process, so that the occlusion core area and its surrounding environment maintain a natural transition in color, brightness and texture.
[0043] Under the support of photometric continuity condition, the time reversal phase gating instruction is obtained to drive the programmable polarization super surface to inject phase conjugate frame envelope, to construct reverse energy channel on the unified time baseline, to trigger black frame extinguishing preferentially to potential occlusion core area, and to write back the calibration factor to the unified time baseline, thereby closing the dynamic regulation link; On the basis of having completed the construction of photometric continuity of the boundary area, in order to solve the problems of black frame, frame skipping and instantaneous flash in multi-view stitching of potential occlusion core area, a dynamic image repair mechanism based on time reversal is introduced, a frame-level phase response structure is constructed to realize compensation and extinguishing control of the key area, and the specific implementation is as follows: Under the premise that photometric mapping has been completed and color compensation of the occlusion core area has ensured image brightness and color continuity, by monitoring the brightness fluctuation trajectory and color channel change rate of each frame in the image sequence, it is identified whether there is a sudden visual abnormal phenomenon. For example, for the area where the brightness decreases by more than 25% between consecutive frames, or the color channel change rate significantly deviates from the average value by more than a set threshold, the system will judge it as a potential frame skipping or black frame warning area. On this basis, combined with the unified time baseline that has been constructed, the key data such as the accurate timestamp of the abnormal frame, the stitching trajectory area coordinates, and the current frame phase offset information are extracted to generate a gating instruction for activating the time reversal compensation mechanism. The gating instruction clearly marks the image range to be compensated, the phase direction adjustment amount to be applied, and the compensation start and end time, providing accurate control conditions for subsequent frame envelope injection.
[0044] Under the control of the gating instruction, a group of image frame envelopes with phase conjugate characteristics are generated for the marked area through the physically controllable polarization surface array unit. The specific method is as follows: first, extract the pixel matrix of the area in the stitching trajectory of the previous frame and the next frame, and analyze the spatial texture distribution, brightness curve trend, and color saturation change trend of the pixels in the area. On this basis, the ideal image state that the current frame should have but lacks is constructed as the target image structure. Next, by controlling each independent regulation unit in the polarization surface to apply reverse phase rotation to the incident image signal, the wavefront and brightness response of each pixel in the output image are changed towards the reverse trajectory of the previous frame image. Such images visually present a "time reversal" effect, that is, the compensation frame restores its image state at the previous time in the missing area of the current frame. Finally, a complete frame envelope structure is formed, which fills in the black frame area pixel by pixel in the spatial image and reproduces the lost information in the time logic.
[0045] To ensure the frame envelope structure has a clear timing position and execution priority in the image processing chain, a reverse frame timing channel needs to be inserted on the unified time baseline. First, according to the time range provided in the gating instruction, a temporary frame time slot of 3-5 milliseconds is allocated in the original time sequence for inserting the compensation frame. Each time in the time slot is allocated a specific time tag for mapping the frame envelope pixel value, thereby forming a set of continuously controllable frame time nodes. Then, each frame in the compensation frame envelope is numbered in reverse order of time and written into the reverse energy channel one by one. The reverse energy channel is not an intervention on the original image time stream, but a temporary interpolation reference, whose only function is to temporarily superimpose phase conjugate information on the current image frame for subsequent image fusion links to call, thereby achieving visual transition compensation.
[0046] After the frame envelope data is written, it is first positioned in all occlusion core regions and scanned for the frame skipping section. Based on the known occlusion core coordinates, the pixel matrix in the current frame image is redrawn, and the pixel values prepared in the frame envelope are used to replace the black frame, empty frame or frame skipping pixel points. To avoid color and texture inconsistency, the dimming process adopts a brightness gradient control method, that is, the compensation pixel value is weighted and interpolated with the pixel values before and after it in the original frame to ensure that the final output image brightness and texture transition is natural and continuous. At the same time, color temperature correction and gamma curve adjustment are performed under the cooperation of the luminosity mapping matrix, so that the compensated image segment is completely embedded in the original frame sequence in the color space, achieving visual consistency and completing high-fidelity restoration of the frame skipping section.
[0047] After frame compensation is completed, all dynamic control parameters involved in the compensation process need to be recorded to the unified time baseline. These parameters include: the timestamp corresponding to each frame in the phase conjugate frame envelope, the brightness correction coefficient, the color mapping difference matrix, the gamma adjustment factor, the interpolation weight coefficient, etc. Through the write interface of the unified time baseline, these parameters are bound to the corresponding positions in the original image frame index, ensuring that the same compensation logic can be reused at any time node in the subsequent image processing process. When the multi-view data stream has synchronization problems or the image state is unstable, the image sequence integrity can be quickly restored by reading the written calibration factors, and the image processing strategy of other regions can be further adjusted, truly realizing adaptive control of the image repair closed loop. At this point, the dynamic regulation link is effectively closed, and the occlusion area timing, luminosity and phase compensation process in the entire multi-view video splicing process are completed.
[0048] The application constructs fine timing control by unifying the time baseline, ensures that multi-camera image acquisition is synchronized in response under millisecond level precision; through the construction of dynamic visibility probability field and delay gradient model, the accurate identification and dynamic tracking of the occluded area are realized; through the multi-camera clock fine tuning and boundary geometry constraint, the seamless splicing track is generated, so that the image still maintains geometric continuity in the occluded core area; through the introduction of luminosity consistency mapping, color compensation and gamma backstepping technology, the color smooth transition of the boundary area is ensured; finally, through the programmable polarization super surface injection phase conjugate frame, the reverse energy channel is constructed, the black frame skipping is eliminated and the dynamic regulation link is closed, which fundamentally eliminates the frame extinguishing and flashing phenomenon caused by the occluded area. The scheme ensures the continuity of the picture, significantly enhances the immersion and spatial perception stability of the user in the virtual reality device, and avoids the problems of dizziness, illusion or cognitive disorder caused by visual mutation.
[0049] The above only describes certain exemplary embodiments of the application by way of illustration, and it is needless to say that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the application. Therefore, the above drawings and descriptions are illustrative in nature and should not be understood as limiting the scope of protection of the claims of the application.
Claims
1. A method for multi-view VR video acquisition and assembly of light small unmanned aerial vehicles, characterized in that, The method comprises the following steps: acquiring a unified time baseline, establishing an overlapping boundary timing index, locking a zero-phase anchor point, and outputting a millisecond-level interface event sequence; under the constraint of the interface event sequence, acquiring boundary occlusion priori, constructing a dynamic visibility probability field and labeling potential occlusion kernels; under the support of the dynamic visibility probability field, acquiring a cross-view delay gradient, imposing constraints on the potential occlusion kernel region, injecting phase traction probes into the interface event sequence, and completing multi-camera clock fine-tuning based on feedback results; under the condition that the clock fine-tuning is completed, acquiring a boundary geometry field according to the cross-view delay gradient and the optical flow residual, setting a geometric limit for the potential occlusion kernel region, generating a seamless splicing track, and presetting a critical transition window on the unified time baseline; within the critical transition window, acquiring photometric consistency mapping based on the boundary geometry field, adopting color bidirectional mapping and gamma inverse strategy, and performing local color compensation on the potential occlusion kernel region to establish photometric continuity conditions; under the support of the photometric continuity conditions, acquiring time reversal phase gating instructions, driving programmable polarization metasurfaces to inject phase conjugate frame envelopes, constructing reverse energy channels on the unified time baseline, triggering black frame blanking for the potential occlusion kernel region, and writing the calibration factor back to the unified time baseline. 2.The method of claim 1, wherein, The interface event sequence output step is as follows: selecting a main control unmanned aerial vehicle integrated with a high-stability constant-temperature crystal oscillator chip and a global navigation satellite receiving module as a time reference source, receiving satellite time pulses for calibrating a local clock signal; the remaining unmanned aerial vehicles receive time synchronization information through a short-wave frequency hopping communication link, perform smoothing correction operations after receiving multiple continuous synchronization signals, and unify the time references of all the unmanned aerial vehicles to the unified time baseline; each unmanned aerial vehicle synchronously records flight speed, heading angle, pitch angle and three-axis attitude angle information when collecting image frames, and establishes a boundary timing index in combination with image frame time stamps; according to the boundary timing index, edge feature extraction and similarity matching are performed on the synchronous image frames, image frames with stable motion vector changes in the image frame group are selected as zero-phase anchor points, and an interface event sequence is constructed as an input for dynamic occlusion modeling. 3.The method of claim 2, wherein, The dynamic visibility probability field construction step is as follows: extracting the time and space information of target entry, occlusion and exit in the image frames according to the interface event sequence, performing trajectory analysis on the image sequence within the time range before and after each event node, acquiring the occlusion trend of the target object and outputting the occlusion priori information; based on the occlusion priori information, a two-dimensional grid probability map is established, the grid elements on the target motion path are subjected to probability attenuation and real-time updating to form a dynamic visibility probability field; regions with probability values below a threshold and meeting the shape stability are identified as potential occlusion kernel regions, and the occlusion persistence is verified through comparative frame luminosity difference analysis to confirm effective occlusion kernels; the effective occlusion kernels are mapped back to the image frames as the basis for subsequent geometric modeling and luminosity correction differential processing. 4.The method of claim 3, wherein, The multi-camera clock fine-tuning step based on feedback results is as follows: a cross-view delay gradient matrix is calculated according to the position offset of the occlusion kernel region in different view images, and the offset is converted into a time delay value in combination with the target speed mapping coefficient; A weighted influence factor is calculated according to the occlusion probability and area of the occlusion kernel region, and each term in the delay gradient matrix is given a calibration priority; According to the delay gradient matrix, a plurality of phase traction probes are injected into the boundary event sequence, and the optimal time offset is selected through image fusion feedback to generate a target time correction table; According to the target time correction table, the clock of the multi-rotor unmanned aerial vehicle is fine-tuned, and whether the frame-level synchronization accuracy meets the splicing requirements is verified through two rounds of fusion tests. 5.The method of claim 4, wherein, The presetting steps of the critical transition window are as follows: According to the delay gradient and the optical flow residual, edge pixel points are extracted, and a boundary geometry field is established; Based on the contour coordinates of the occlusion kernel region, an occlusion restriction body is constructed to restrict the occlusion kernel region from being crossed by the splicing trajectory and control the curvature variation range; According to the minimum optical flow residual and the delay path, a trajectory point is selected, a continuous and non-crossing occlusion kernel splicing trajectory is generated by combining a smoothing filter, and a critical transition window is calibrated around the trajectory for subsequent photometric and texture processing. 6.The method of claim 5, wherein, The steps for establishing the photometric continuity condition are as follows: Extract the pixel regions on both sides of the splicing trajectory, count the three-channel gray scale distribution, and construct a color mapping index matrix; Based on the mapping matrix, bidirectional color correction is performed, gamma reverse processing is performed in combination with the exposure characteristics of the image acquisition device, and the brightness continuity of the splicing area is ensured; In the occlusion kernel region, classified color compensation is performed according to the color statistical characteristics of the outer ring pixels, and a ring-shaped mixed transition zone is constructed to achieve smooth color transition and photometric consistency. 7.The method of claim 6, wherein, In the color compensation of the occlusion kernel region, the average brightness difference of the ring-shaped edge region is used to perform pixel brightness enhancement, and the red, green and blue channel proportion of the pixels in the occlusion kernel is adjusted through a color temperature rotation matrix, so that the compensation result is consistent with the surrounding image in brightness and color temperature. 8.The method of claim 6, wherein, Under the support of the photometric continuity condition, the time reversal phase gate instruction is obtained, the programmable polarization super surface is driven to inject the phase conjugate frame envelope, the reverse energy channel is constructed on the unified time baseline, the black frame in the occlusion kernel region is triggered to extinguish, and the calibration factor is written back to the unified time baseline as follows: Under the condition of photometric continuity, the brightness change and color channel change rate of the image frame are monitored to identify the frame skipping or black frame warning area, and the compensation time interval, phase direction and region position are marked according to the unified time baseline output gate instruction; Under the control of the gate instruction, the polarization control unit generates a phase conjugate frame envelope, the pixel values of the target region before and after the image frame are collected, the target image structure is constructed, and the compensation image is formed by applying a reverse phase rotation; On the unified time baseline, the reverse frame time sequence channel is demarcated, the compensation frame is allocated a dedicated time slot, the frame envelope content is labeled in reverse order, and the continuous compensation sequence is formed by writing it into the channel frame by frame; Read the frame envelope compensation image, perform brightness interpolation and texture redraw on the frame skipping area, embed the compensation image into the original frame through a weighted method, and synchronously adjust the color temperature difference and gamma factor to ensure natural visual transition; The time stamp, brightness correction coefficient, color mapping matrix and interpolation parameter involved in the compensation process are written back to the unified time baseline to complete the image repair information binding and dynamic control link closure.