A method, system, and computing device for processing video data

By cropping and distortion correction of the region of interest in the panoramic video, combining tracking algorithms to analyze key moments, and aligning and compositing the video clips with the video clips from the second shooting device on the timeline, the problem of automatic selection and switching between panoramic and local images was solved, achieving high-quality video fusion and enhanced viewing experience.

CN121239897BActive Publication Date: 2026-02-03SUZHOU DEEPSIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511820021.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-02-03
Estimated Expiration
2045-12-04

AI Technical Summary

Technical Problem

Existing video processing methods struggle to automatically select and switch between panoramic and partial views, resulting in visual discontinuity during scene transitions and an inability to simultaneously cover both scene coverage and detail presentation. This is especially problematic when moving quickly or when multiple targets enter a critical scene, making it easy to miss crucial events.

Method used

By cropping the region of interest and correcting distortion in the video from the panoramic shooting device, combining it with tracking algorithms to analyze key moments, and then aligning and compositing it with the video clips from the second shooting device on the timeline, high-quality target shooting footage is generated.

Benefits of technology

It achieves seamless integration of panoramic and partial images, improves image quality and viewing experience, reduces the cost of manual intervention, ensures logical consistency and subject prominence in image switching, and provides a high-quality visual experience that combines macro-narrative and micro-detail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121239897B_ABST
    Figure CN121239897B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of processing method, system and computing device of video data, the object of the method processed includes the first original video collected by panoramic shooting device and the second original video collected by second shooting device.First, the first original video is analyzed, the region of interest is determined based on tracking algorithm, and target shooting picture is cropped according to preset composition strategy;Subsequently, the first video segment belonging to highlight moment is identified in the first original video, and the second video segment corresponding to its time axis is intercepted in the second original video.If the second video segment also belongs to highlight moment, it is synthesized with target shooting picture, and the final processing picture is generated by replacing or picture-in-picture etc..The scheme of the present application has the dual advantages of wide field of view of panoramic picture and high local close-up definition, can improve the presentation effect of highlight picture in automatic editing process, and reduce the workload of manual screening and editing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video analysis and editing, and more particularly to a method, system, and computing device for processing video data. Background Technology

[0002] With the development of image processing technology, artificial intelligence algorithms, and intelligent shooting equipment, more and more sporting events, performances, and classrooms are relying on multi-camera video capture to improve content quality. Panoramic shooting devices can cover a wider field of view, ensuring the integrity of the image information, while tracking or fixed shooting devices provide clearer details within specific areas. How to effectively integrate these video resources from different devices in post-processing, so that the final image not only possesses complete panoramic information but also highlights key actions or exciting moments, is a common technical challenge facing the industry.

[0003] Existing video processing methods typically rely on manual editing or single-camera-based video analysis, making it difficult to automatically select and switch between panoramic and close-up shots. When there is fast movement, changes in perspective, or multiple targets entering a critical scene simultaneously, relying on a single video source often fails to simultaneously cover both scene coverage and detail. Furthermore, another common approach is to record close-up shots from fixed or tracking camera positions, but these shots often don't strictly correspond to the highlights of the panoramic video, causing automated editing to miss crucial events or create visual discontinuities during shot transitions.

[0004] Therefore, a technical solution is needed that can establish a timeline correlation between multi-source video data, automatically identify key moments by combining image analysis and tracking algorithms, and perform precise cropping and compositing based on panoramic images. By intelligently fusing high-quality local images from a second shooting device with panoramic images, the image quality and viewing experience of key moments can be effectively improved, the cost of manual intervention can be reduced, and the logical consistency of scene transitions and the prominence of the subject can be guaranteed. This invention is an improved solution proposed based on this technical requirement. Summary of the Invention

[0005] In view of the above, the purpose of this application is to overcome the shortcomings of the prior art. In a first aspect of this application, a method for processing video data is provided, wherein the video data includes a first original video captured by a panoramic shooting device and a second original video captured by a second shooting device, characterized in that the method includes:

[0006] The first original video is analyzed, and the frame containing the region of interest is cropped out as the target shooting frame;

[0007] Based on the first original video, analyze the first video segment that belongs to the highlight moment;

[0008] Extract the second video segment from the second original video that corresponds to the timeline of the first video segment;

[0009] If the second video clip is a highlight, then the second video clip is combined with the target shot to generate the processed shot.

[0010] In one embodiment, the analysis of the first original video includes:

[0011] The first original video is analyzed based on a tracking algorithm, and the obtained tracking points are taken as the region of interest.

[0012] In one embodiment, the second shooting device is a tracking shooting device, which tracks and shoots the scene based on the tracking algorithm to obtain a second original video.

[0013] In one embodiment, the second shooting device is a fixed shooting device, which shoots at a preset area in the scene to obtain a second original video.

[0014] In one embodiment, the process of combining the second video segment with the target captured image includes:

[0015] Replace the corresponding first video segment in the target shooting scene with the second video segment.

[0016] In one embodiment, the process of combining the second video segment with the target captured image includes:

[0017] The second video clip is used as a picture-in-picture of the corresponding first video clip in the target shooting frame.

[0018] In one embodiment, the step of cropping the image containing the region of interest as the target image includes:

[0019] Based on the location of the region of interest in the first original video and the preset composition strategy, the corresponding scene is cropped out as the target shooting scene.

[0020] In a second aspect of this application, a video data processing system is provided, the video data including a first raw video captured by a panoramic shooting device and a second raw video captured by a second shooting device, characterized in that it includes:

[0021] The cropping module is adapted to analyze the first original video and crop out the image containing the region of interest as the target shooting image;

[0022] The analysis module is adapted to analyze the highlights of the first original video and use them as the first video segment.

[0023] The extraction module is adapted to extract a second video segment from the second original video that corresponds to the timeline of the first video segment.

[0024] The compositing module is adapted to determine whether the second video segment is a highlight based on the analysis module. If so, the second video segment is combined with the target shooting image to generate the processed shooting image.

[0025] A third aspect of this application provides a computing device, characterized in that the computing device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method of the first aspect of this application.

[0026] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when run on a processor, executes the method of the first aspect of this application.

[0027] This invention effectively solves the technical challenge of simultaneously capturing global information and key details from a single perspective by integrating the wide-angle view of a panoramic shooting device with the high-definition local details of a second shooting device. The system establishes precise timeline correlations between multi-source video data, automatically identifying and locating key moments using image analysis and tracking algorithms, thereby achieving intelligent filtering and synthesis from panoramic views to high-definition close-ups. This automated processing significantly reduces the tedium and cost of manual editing, avoiding the omission of key moments that might occur during manual operation. Furthermore, it seamlessly integrates low-distortion, high-resolution close-ups into the target image through various synthesis methods such as picture-in-picture or image replacement. Ultimately, this solution significantly improves the clarity and viewing experience of key shots in sports events or activity recordings while ensuring the integrity of the video content, providing users with a high-quality visual experience that combines macro-narrative and micro-detail. Attached Figure Description

[0028] Figure 1 This is a flowchart of a video data processing method according to an embodiment of this application;

[0029] Figure 2 This is a schematic diagram illustrating a picture-in-picture method for displaying video in an embodiment of this application;

[0030] Figure 3 This is a schematic diagram of the data flow in a video data processing system according to an embodiment of this application;

[0031] Figure 4 This is a schematic diagram of the structure of the terminal device or server in the embodiments of this application. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] In the field of computer vision analysis, the origin of the coordinate system is usually located at the upper left corner of the screen by default. In all embodiments of this application, unless otherwise stated, the upper left corner is used as the origin of the screen coordinate system. Those skilled in the art should know that such a coordinate system setting is not absolutely fixed. When the origin of the coordinate system is set at any position inside or outside the screen, the corresponding technical solutions that can be obtained by simple adjustments to this solution without creative effort are all within the protection scope of this application.

[0034] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0035] The event video editing method provided in this application can run on various computing devices, including laptops, smartphones, tablets, or other smart devices with data acquisition capabilities, such as smart cameras and wearable devices. It can also run in the cloud and provide editing services through network communication.

[0036] In one embodiment, this application provides a method for processing video data. The video data involved in this application embodiment is multi-channel video acquisition data, including data acquired by at least two different shooting devices. In one embodiment, the video data includes a first original video acquired by a panoramic shooting device and a second original video acquired by a second shooting device. The panoramic shooting device is a device capable of capturing the entire scene, and can be implemented as a wide-angle lens or a combination of multiple lenses, or as a 360° panoramic video acquisition device. The panoramic shooting device has a high degree of distortion, achieving scene-free shooting at the expense of resolution and image flatness. The second shooting device is a supplementary shooting device, used to acquire high-quality images of specific areas or hotspots in the scene. Compared to the first shooting device's approach of sacrificing resolution and image flatness to achieve large-area coverage, the second shooting device uses a smaller field of view imaging method, providing higher resolution, lower distortion, and clearer image details in local areas, thereby specifically capturing key content in the scene and compensating for the shortcomings of the first shooting device in image quality. The second shooting device can be a device that takes pictures of a designated key area from a fixed perspective, or it can be a device that takes pictures of a real-time hot spot area on the field from a variable perspective; this application does not limit the specific type of the second shooting device or the method of image acquisition.

[0037] The first and second shooting devices capture the same scene. They can be set in the same location or in different locations. Their relative positions are irrelevant to the technical solutions in this application embodiment.

[0038] Furthermore, there is a direct or indirect communication connection between the first and second shooting devices. Specifically, the first and second shooting devices can agree to start shooting simultaneously via the communication connection, or one shooting device can send an electrical signal to notify the other shooting device when it starts shooting, so that the first and second shooting devices can align their timelines. The communication connection can be direct, such as through Bluetooth networking; or indirect, such as through a signal repeater like WLAN.

[0039] like Figure 1 As shown, the video data processing method in this embodiment includes:

[0040] S1: Analyze the first original video and crop out the frame containing the region of interest as the target shooting frame.

[0041] In this embodiment, the first original video is a panoramic video captured by the first shooting device, with a field of view of 180° to 360°, capable of completely recording all or most of the spatial information of the shooting scene. Because panoramic videos are captured using multi-frame stitching, wide-angle, or fisheye lenses, their original images typically exhibit significant barrel or spherical distortion characteristics, and the resolution within a unit field of view is relatively low in order to achieve large field-of-view coverage. When processing this panoramic video, the location information of the region of interest (ROI) must first be determined. This ROI can be a preset fixed area, such as the goal area in a sports field or the three-point line area in a basketball court, or it can be a dynamically determined area based on target detection or tracking algorithms, such as an area containing athletes, balls, or other moving targets. After determining the ROI, the system extracts the corresponding local area from the complete panoramic image based on the position parameters of this area in the panoramic video coordinate system. This cropping process usually requires the use of distortion correction algorithms to convert the panoramic image with spherical distortion into a planar image that conforms to normal perspective, thereby providing a better viewing experience for the cropped target image. Specifically, panoramic video can be reprojected using methods such as equidistant cylindrical projection, cube projection, or perspective projection. During this process, the pixel region corresponding to the target image is calculated based on the center point coordinates of the region of interest and the desired field of view. A clear cropped image is then generated using image interpolation algorithms. After cropping, the target image not only focuses on key content in the scene but also achieves a more natural image effect through distortion correction. Furthermore, since the cropped area is relatively small compared to the panoramic image size, the data volume can be effectively reduced during subsequent video encoding and transmission.

[0042] In one embodiment, the principle of distortion correction is to calculate the corrected coordinates based on the distortion model using pre-calibrated distortion parameters, such as radial distortion coefficients k1 and k2, and tangential distortion coefficients p1 and p2.

[0043] One example of a distortion model is as follows:

[0044] Where r is the distance from the pixel to the center of the panoramic image from the wide-angle camera, and C x and C y Let x be the coordinates of the center point. real y real These are the coordinates of the pixel before distortion correction, x corrected y corrected These are the coordinates of the pixel after distortion correction, where x is the x-coordinate of the pixel. real relative to the x-coordinate C of the center point x The offset, where y is the x-coordinate of the pixel. real relative to the x-coordinate C of the center point yThe offset.

[0045] In one embodiment, for ultra-wide-angle cameras with high distortion (field of view greater than 140°), global correction using an equirectangular projection model is supported. This method projects the original image onto a spherical coordinate system and then reprojects it according to the field of view of the tracking point to generate a distortion-free image. This spherical reprojection is particularly suitable for fisheye lens scenes, ensuring that when the sub-image cropping angle reaches 30° to 40°, there will be no stretching or distortion in the edge areas.

[0046] In one embodiment, to improve real-time performance, the present invention employs a pre-computed lookup table (LUT) method. During the initialization phase, a coordinate mapping table with the same resolution as the image is generated using calibration parameters, and the original image sampling coordinates corresponding to each tracking point pixel are pre-stored. During video execution, this LUT can be used to complete distortion correction within a single frame, avoiding the computational overhead caused by real-time pixel-by-pixel calculation.

[0047] In this embodiment, to further improve image stability, a geometric consistency check is performed after distortion correction. When significant curvature is detected in the straight structures at the four corners of the image (such as ground edges or building edges), the system adaptively updates the fine-tuning parameters of the distortion coefficients to dynamically compensate for optical deviations caused by temperature changes or lens aging. This adjustment can be achieved by combining edge detection algorithms (such as Canny edge detection) with straight line fitting, with a real-time correction error of less than 0.1°.

[0048] In one embodiment, the distortion correction module is typically located before tracking point detection and cropping calculation in the cropping process. That is, the input frame is first processed through distortion-free mapping to generate a distortion-free image, and then input into the tracking point detection algorithm to obtain accurate tracking point coordinates. In performance-constrained scenarios, a "regional distortion correction" scheme can also be used, where fast distortion correction is performed only on a certain range (e.g., 500×500 pixels) of the tracking point's neighborhood to reduce computational burden. In this case, the system dynamically generates a local mapping matrix based on the tracking point's location and only resampling is performed on the sub-image region to be cropped.

[0049] Furthermore, the image is selected based on a composition strategy. The composition strategy involves placing the region of interest (ROI) in the final image position. In different implementations, the composition strategy keeps the ROI at the center of the real-time image, at a specific golden ratio point, or at any arbitrary position depending on the specific circumstances of the scene.

[0050] S2: Based on the first original video, analyze the first video segment that belongs to the highlight moment.

[0051] Steps S2 and S1 are both based on the analysis and processing of the first original video, and there is no necessary order between them; they can be processed in parallel or sequentially. The specific method for analyzing highlights in a scene is not the subject of this invention; it can be achieved through multi-dimensional video content analysis in existing technologies. From an audio perspective, highlights are often accompanied by cheers, applause, or sudden changes in the volume and speed of the commentator. Therefore, abnormal peaks in these acoustic features can be captured through audio energy detection, spectrum analysis, or speech emotion recognition. When a significant energy increase or concentrated burst of a specific frequency band is detected in the audio signal within a short period, it usually means that an event has occurred that triggered a strong reaction from the audience. From a visual content analysis perspective, key events on the field can be determined through object detection and behavior recognition technologies. For example, in a basketball game, by identifying the trajectory of the basketball, when the ball passes through the basket area quickly at a specific angle, combined with the deformation of the basket net or the rebound pattern of the basketball after landing, it can be determined that a basket has been scored. Meanwhile, changes in player posture can also provide important clues, such as a sudden increase in jump height, multiple players moving rapidly in the same direction within a short period, and the appearance of physical contact. These dramatic changes in movement patterns often correspond to exciting moments such as steals, blocks, and fast breaks. Furthermore, interface elements in the game footage are also important indicators. Changes in the scoreboard numbers directly indicate the occurrence of a scoring event, while slow-motion replay markers, special effects animations, or the director switching to close-up shots are usually active annotations of exciting moments by the broadcaster. Using OCR technology to identify changes in the scoreboard can efficiently locate exciting moments. Furthermore, deep learning models can be built to fuse and train the above multimodal features, enabling the model to learn the comprehensive feature patterns of different types of exciting moments. This allows for end-to-end automatic recognition and classification and annotation according to different exciting types (scoring, steals, blocks, etc.), ultimately outputting a list of exciting moments with timestamps and event tags.

[0052] In another possible implementation, the monitoring of highlights is achieved through a large model training and invocation method. After determining the specific type of highlight, those skilled in the art can use common existing technologies to collect highlight material to train the large model, enabling the large model to distinguish the corresponding highlights.

[0053] S3: Extract the second video segment from the second original video that corresponds to the timeline of the first video segment.

[0054] As mentioned above, in this embodiment, there is a direct or indirect communication connection between the first shooting device and the second shooting device. This communication connection can be direct, such as establishing a point-to-point communication link via Bluetooth, Wi-Fi Direct, or wired connections (e.g., USB, HDMI, Ethernet cable); or it can be indirect, such as the two shooting devices being connected to the same WLAN network and exchanging data through signal repeaters like routers or switches, or communicating via a cloud server as a relay node. Through this communication connection, the first and second shooting devices can exchange information, including the transmission of control commands, the transfer of time synchronization information, and notifications of shooting status.

[0055] Specifically, the first and second shooting devices can agree to start shooting simultaneously via a communication connection. In one embodiment, after the user issues a shooting command on the control terminal, the command is simultaneously transmitted to both the first and second shooting devices. Upon receiving the command, both shooting devices synchronously start their video capture functions. In another embodiment, one shooting device acts as the master device, and the other as the slave device. When the master device starts shooting, it sends a trigger signal to the slave device via the communication connection. Upon receiving the trigger signal, the slave device immediately starts shooting, thereby achieving data sharing between the two shooting devices. This sharing mechanism ensures that the first and second original videos can know each other's start time, thus determining the timeline offset between the two devices and laying the foundation for subsequent timeline alignment.

[0056] To further ensure precise timeline alignment, the first and second shooting devices synchronize their times during shooting. This time synchronization can be achieved in various ways. In a preferred embodiment, after establishing a communication connection, the two shooting devices first perform clock calibration, for example, using Network Time Protocol (NTP) or Precision Time Protocol (PTP) to synchronize their clocks, ensuring that the system clocks of the two devices remain consistent. During shooting, each video frame is precisely timestamped, generated based on the synchronized system clock, thus ensuring that the video frames captured by the two shooting devices have comparable time stamps. In another embodiment, the two shooting devices can maintain time alignment by periodically exchanging synchronization signals. For example, at regular time intervals (e.g., 1 second, 5 seconds, etc.), the master device sends a time synchronization signal to the slave device. Upon receiving the signal, the slave device adjusts its time base to compensate for any potential clock drift.

[0057] After identifying the first video segment representing a highlight moment from the first original video, a second video segment corresponding to the timeline of the first video segment needs to be extracted from the second original video. This extraction process is based on the timeline correspondence. Specifically, each first video segment has a defined time range, defined by a start time point and an end time point. For example, if the first video segment is from the 30th second to the 45th second of the first original video, then the time range of this video segment is 15 seconds, with the start time point corresponding to the 30th second after the start of the first original video and the end time point corresponding to the 45th second after the start of the first original video.

[0058] Since the first and second shooting devices are time-axis aligned, corresponding time points also exist in the second original video. The system locates a video frame with the same timestamp in the second original video as the starting point of the first video segment; and locates a video frame with the same timestamp in the second original video as the ending point of the first video segment. Then, the video content from the starting point to the ending point is extracted to form the second video segment. In this way, the second video segment and the first video segment maintain a strict time-axis correspondence; that is, the two video segments record the scene content that occurs within the same time period, only differing in shooting angle and shooting range.

[0059] In a specific application scenario, assume that both the first and second original videos were captured starting at time T0, with the first original video having a frame rate of 30fps and the second original video having a frame rate of 60fps. Analysis determines that the segment from time T0+30 seconds to T0+45 seconds in the first original video is the first video segment representing the highlight. The system first calculates the corresponding frames 900 (30 seconds × 30fps) to 1350 (45 seconds × 30fps) in the first original video. Since the timelines of the two shooting devices are aligned, the system locates the video frame corresponding to time T0+30 seconds in the second original video, which is frame 1800 (30 seconds × 60fps), and the video frame corresponding to time T0+45 seconds, which is frame 2700 (45 seconds × 60fps). Subsequently, the system extracts the content from frames 1800 to 2700 in the second original video as the second video segment. In this way, although the two video segments have different frame rates and frame counts, the correspondence of timestamps ensures that the first video segment and the second video segment record the exact same time period.

[0060] In practical applications, considering minor time deviations that may be caused by factors such as network latency and device response time, the system can also set a certain time tolerance range. For example, when extracting the second video segment, the calculated start time point can be extended by several milliseconds (e.g., 50 milliseconds, 100 milliseconds), and the calculated end time point can also be extended accordingly. Then, the video segment that best matches the first video segment is selected from the extended time range as the second video segment. This tolerance mechanism improves the system's robustness to time deviations, ensuring that even with slight time asynchrony, the corresponding video content can still be accurately extracted.

[0061] Furthermore, in some embodiments, the first and second shooting devices may not start shooting strictly simultaneously, but rather with a certain time difference. In this case, the system needs to first determine the time offset between the two shooting devices. This time offset can be obtained in various ways, such as by having the two shooting devices record their respective start timestamps at the start of shooting and exchange these timestamps via a communication connection to calculate the time offset; or by analyzing synchronous events in the two videos (such as simultaneous flashes, sudden sound changes, etc.) to determine the time offset. After obtaining the time offset, the system will take this offset into account when extracting the second video segment, ensuring that the extracted second video segment corresponds to the first video segment in actual time.

[0062] S4: If the second video clip is a highlight moment, then the second video clip is combined with the target shooting screen to generate the processed shooting screen.

[0063] The determination of whether the second video clip belongs to a highlight moment is consistent with the method described in S2 above, which analyzes the first video clip to determine if it belongs to a highlight moment based on the first original video. Therefore, it will not be repeated here. If the second video clip belongs to a highlight moment, it is logically considered to be the same content as the corresponding first video clip. Since the second video clip has high resolution, low distortion, and may contain zoomed-in footage, it is composited with the target captured image to provide an additional highlight image, generating the processed captured image.

[0064] In step S4, if the second video clip does not belong to a highlight moment, it indicates that the content of the second video clip is inconsistent with that of the first video clip. There could be various factors that could cause this, such as screen obstruction by the second shooting device or tracking errors.

[0065] In a preferred embodiment, the step of analyzing the first original video includes analyzing the first original video based on a tracking algorithm and using the obtained tracking points as the region of interest.

[0066] Understandably, the basic logic of a tracking algorithm is to determine whether the shooting angle should be adjusted to continuously track and capture key targets in real time by observing the movement of targets in consecutive frames captured by the shooting device. For each new input target detection result in each frame, the tracking algorithm may calculate a new tracking strategy. In the embodiments of this application, the tracking algorithm determines which position in the captured image should be designated as the tracking point based on the target identification situation of the current frame. For example, the tracking algorithm can be implemented as an algorithm that identifies and tracks a basketball or soccer ball in real time, or as an algorithm that identifies and tracks a specified human target in real time, or as an algorithm that performs comprehensive position tracking on a specified group in the image. Similarly, it can also be a complex tracking algorithm that integrates multiple tracking logics, where the tracking point is constantly changing and could be a ball target in the image, or the center point of a human target or multiple human targets. This invention does not limit the specific tracking point calculation logic of the tracking algorithm. Specifically, the tracking algorithm can be implemented as CLNet (Compact Latent Network), OpenCV, optical flow, or any existing or proprietary tracking algorithm that conforms to the above description. After obtaining the calculated tracking points, the tracking system adjusts the gimbal angle to track and capture the scene based on the position of the tracking points in the real-time image and compositional factors. In different implementations, the tracking system keeps the tracking points at the center of the real-time image or at a specific golden ratio composition point.

[0067] Furthermore, in another preferred embodiment, the second shooting device is a tracking shooting device. This device tracks and shoots the scene based on the tracking algorithm to obtain a second original video. The tracking shooting device is a device that tracks and shoots hotspot areas in a scene based on a tracking algorithm. It includes at least a video acquisition module and a rotation module. The video acquisition module is responsible for capturing real-time images. It can use a high-resolution image sensor to ensure rich detail and high clarity in the acquired video images, while also possessing a low-distortion optical lens to reduce distortion at the edges of the image and improve image quality. The rotation module is used to drive the video acquisition module to rotate at multiple angles. It typically consists of a motor-driven gimbal structure, capable of 360-degree horizontal rotation and pitch adjustment within a certain angle range in the vertical direction, thus ensuring that the tracking shooting device can flexibly track the movement trajectory of the target. In actual operation, the tracking shooting device receives tracking point information output by the tracking algorithm and adjusts the angle of the rotation module in real time to ensure that the tracking point is always within the effective shooting range of the video acquisition module and, as far as possible, at a key compositional position in the image, such as the center or the golden ratio point, to obtain the best visual effect. For example, in a basketball game, when the tracking algorithm identifies the basketball player as a key target and outputs its real-time position as the tracking point, the rotation module of the tracking and shooting device will respond quickly, driving the video acquisition module to rotate, ensuring that the athlete always appears clearly in the picture. Even if the athlete moves quickly or changes direction, stable tracking can be achieved through continuous angle adjustment, thereby generating a second original video with a focused perspective and clear details.

[0068] In this embodiment, the tracking algorithm running in the tracking shooting device is consistent with the tracking algorithm that analyzes the first original video to obtain the tracking points. This ensures that the real-time tracking points of the tracking shooting device are consistent with the region of interest of the first original video, that is, it ensures that the area where the tracking shooting screen is located is consistent with the area where the target shooting screen is located.

[0069] In another possible embodiment, the second shooting device is configured as a fixed shooting device. A fixed shooting device refers to an image acquisition device whose physical position and shooting angle remain relatively still during the shooting process. Specifically, the second shooting device can be placed in a specific location at the shooting site using physical support structures such as tripods, pan-tilt heads, brackets, or suction cup bases, or it can be installed on fixed infrastructure such as walls, ceilings, lampposts, or monitoring poles. This fixed setup allows the second shooting device to provide stable, shake-free images, providing a baseline perspective for the entire shooting scene. For example, in a live sports event scenario, the second shooting device could be a fixed-position camera mounted on a bracket behind the goal; in a meeting recording scenario, the second shooting device could be a fixed camera placed at one end of the conference table and facing the speaker.

[0070] The fixed shooting device performs point-to-point shooting of a preset area in the scene. The preset area is a spatial range in the shooting scene that is predefined as the focus of attention. In one embodiment, the preset area is physically selected by the user during the shooting preparation stage by manually adjusting the lens orientation and focal length of the second shooting device. For example, the user points the lens of the second shooting device at the center of the stage, the blackboard area, or the finish line of a game, and confirms the coverage area through the viewfinder. Once confirmed, the mechanical structure of the device is locked, thereby establishing the preset area. In another embodiment, the preset area can be set through interactive software interface. The user views the real-time preview screen of the second shooting device on the control terminal and draws the Region of Interest (ROI) on the screen. The second shooting device (if it is a fixed device with a gimbal function) automatically adjusts the gimbal angle and zoom magnification according to the drawn ROI to ensure that the center of the image is aligned with the preset area and locked. In addition, the preset area can also be automatically determined based on recognition algorithms. For example, the system automatically identifies densely populated areas of faces, specific landmarks (such as a podium, a net), etc., in the scene, and sets the area containing the identified object as the preset area.

[0071] After determining the preset area, the fixed shooting device performs a fixed-point shooting task. During video acquisition, the optical axis direction and field of view (FOV) of the second shooting device remain unchanged (or only make necessary autofocus adjustments within a preset small range), continuously acquiring image data of the preset area. The data generated by this shooting method has the characteristic of a relatively static background, which is beneficial for subsequent data processing (such as background modeling, moving target detection, etc.). Through the above fixed-point shooting process, the second raw data is obtained. Since the second shooting device is fixed and shoots towards the preset area, the second raw data often records environmental information or the continuous state of key positions in the scene. For example, in a football match, the target shooting frame cropped from the first raw video is the area determined by the position of the football on the field, capturing the most intense area of ​​competition on the field, while the second shooting device, as a fixed shooting device, can be fixedly aimed at the goal and shoot. The second raw video it generates records all events that occur within the goal area and can perform dynamic optical zoom according to the location of the events. This setup allows the second raw data to serve as an effective supplement to the first raw data. When a key event is detected in the first shooting device, the system can call up the corresponding footage from the second raw data at any time to supplement the information, or use the specialized perspective provided by the second raw data to assist in generating rich video effects such as picture-in-picture and multi-view splicing.

[0072] Preferably, the second shooting device also performs zoom shooting of the scene based on a zoom algorithm. In one possible embodiment, the step of compositing the second video clip with the target shooting image in step S4 can be implemented as a direct image replacement. That is, when the second video clip is determined to be a highlight moment, the system directly replaces the first video clip in the target shooting image with the second video clip. This replacement method can quickly integrate high-resolution, low-distortion highlight images into the target image. For example, in the target shooting image of a basketball game, when the first video clip identifies the highlight moment of a player completing a dunk, if the corresponding second video clip is a close-up of the dunk taken from a closer and clearer angle, the system will directly replace the corresponding time period content in the target image with this close-up clip, allowing the audience to intuitively see more impactful highlight details. In addition, the size and proportion of the second video clip can be adaptively adjusted during the replacement process to ensure that it is consistent with the overall display effect of the target shooting image, avoiding problems such as abrupt image or disproportionate image. Preferably, a gradual transition effect can also be added.

[0073] In another possible embodiment, the step of compositing the second video clip with the target captured image in step S4 can be implemented as an addition of images. That is, when the second video clip is determined to be a highlight, the system inserts the second video clip after the first video clip in the target captured image has finished playing. This addition method can provide the audience with additional exciting perspectives without affecting the main narrative of the target captured image. For example, in the target captured image of a concert, if the first video clip records a panoramic view of the singer's climax, and the corresponding second video clip is a close-up of the singer's face and emotional expression details captured from the side of the stage, the system will automatically insert this close-up clip after the panoramic view finishes playing, allowing the audience to understand the overall atmosphere of the performance and deeply appreciate the details of the singer's performance, enriching the viewing experience. During the addition process, reasonable insertion positions and transition effects, such as fade-in / fade-out and sliding transitions, can be set according to the length and content characteristics of the second video clip to ensure a natural and smooth transition of the image and avoid interfering with the audience's viewing rhythm. In addition, the system can also impose duration limits or priority sorting on the added second video clips according to user needs or preset rules, ensuring that the most valuable content is displayed first when multiple exciting clips exist.

[0074] In another possible embodiment, the step of combining the second video segment with the target captured image in step S4 can be implemented as a picture-in-picture display. For example... Figure 2As shown, the first video clip is continuously played in the main video screen 200, while a picture-in-picture (PiP) screen 210 is overlaid in a small window in a specific area of ​​the main video screen 200 (such as the upper right corner, lower right corner, or other preset positions). This display method allows viewers to simultaneously view exciting content from another perspective while watching the main screen, achieving a parallel presentation of multi-dimensional information. For example, in a live broadcast of a football match, the main window plays a panoramic view of the first video clip, clearly showing the players' overall movement and tactical coordination, while the PiP window simultaneously plays the second video clip, focusing on close-ups of players dribbling past defenders or details of goalkeeper saves, allowing viewers to grasp the overall game while not missing key moments. Furthermore, to avoid the PiP window obscuring the main screen content, semi-transparency or intelligent avoidance algorithms can be used. When a key target in the main screen moves into the PiP window area, the position of the PiP window is automatically adjusted or its size is temporarily reduced, ensuring the continuity and integrity of the viewing experience.

[0075] Based on the same inventive concept described above, this application also provides a video data processing system. The system processes video data including first raw video data captured by a panoramic shooting device and second raw video data captured by a second shooting device. In its specific operational logic, the system mainly includes a cropping module, an analysis module, a cropping module, and a compositing module. Figure 3 A schematic diagram of the data flow in a video data processing system according to an embodiment of this application is shown. For ease of understanding, in... Figure 3 Data is represented by blue boxes, and system modules are represented by gray boxes.

[0076] The cropping module is configured to perform intelligent analysis on the first original video, identify regions of interest (ROIs) by recognizing salient features or specific objects in the image, and perform distortion correction or reconstruction on the panoramic view to crop out the image containing the ROIs as the target shot. Simultaneously, the analysis module is adapted to perform temporal feature extraction based on the first original video data, identifying key moments by analyzing changes in image content or audio features, and defining these as the first video segment.

[0077] To utilize multi-view footage, the system's extraction module is adapted to locate and extract the corresponding time segment of video stream from the second original video data based on the start and end points of the first video segment on the timeline, thus obtaining the second video segment and achieving synchronization of footage from different camera positions in the time dimension. Further, the compositing module is responsible for generating the final image. This module is adapted to perform logical processing based on the judgment results of the analysis module: specifically, the analysis module further determines whether the aforementioned extracted second video segment meets the criteria for a highlight moment (e.g., through image quality assessment or content relevance analysis); if so, it triggers the compositing module to composite the second video segment with the target shooting image output by the cropping module, for example, using picture-in-picture, split-screen splicing, or fusion display, ultimately generating the processed shooting image to output a video file containing highlights from multiple perspectives.

[0078] like Figure 4 As shown, based on the same inventive concept described above, this application also provides a computing device, which can be implemented as a terminal device or a server. The computing device includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the terminal device or server. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0079] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.

[0080] Specifically, according to embodiments of this application, the above method flow steps can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the system of this application.

[0081] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, or any suitable combination thereof.

[0082] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0083] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be located in a processor. The names of these units or modules do not, in certain circumstances, constitute a limitation on the unit or module itself.

[0084] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that are used by one or more processors to execute the methods described in this application.

[0085] In another aspect, embodiments of this application also provide a computer program product that, when executed by a processor, implements the methods of any of the above embodiments.

[0086] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions claimed in this application.

Claims

1. A method for processing video data, wherein the video data includes a first original video captured by a panoramic shooting device and a second original video captured by a second shooting device, characterized in that, The method includes: The first original video is analyzed, and the frame containing the region of interest is cropped out as the target shooting frame; Based on the first original video, analyze the first video segment that belongs to the highlight moment; Extract the second video segment from the second original video that corresponds to the timeline of the first video segment; If the second video clip is a highlight, then the second video clip is combined with the target shot to generate the processed shot.

2. The method as described in claim 1, characterized in that, The analysis of the first original video includes: The first original video is analyzed based on a tracking algorithm, and the obtained tracking points are taken as the region of interest.

3. The method as described in claim 2, characterized in that, The second shooting device is a tracking shooting device. The second shooting device tracks and shoots the scene based on the tracking algorithm to obtain a second original video.

4. The method as described in claim 1 or 2, characterized in that, The second shooting device is a fixed shooting device, which shoots at a preset area in the scene to obtain a second original video.

5. The method as described in claim 1, characterized in that, The process of combining the second video segment with the target captured image includes: Replace the corresponding first video segment in the target shooting scene with the second video segment.

6. The method as described in claim 1, characterized in that, The process of combining the second video segment with the target captured image includes: The second video clip is used as a picture-in-picture of the corresponding first video clip in the target shooting frame.

7. The method according to any one of claims 1-3, characterized in that, The step of cropping out the image containing the region of interest as the target image includes: Based on the location of the region of interest in the first original video and the preset composition strategy, the corresponding scene is cropped out as the target shooting scene.

8. A video data processing system, wherein the video data includes a first raw video captured by a panoramic shooting device and a second raw video captured by a second shooting device, characterized in that, include: The cropping module is adapted to analyze the first original video and crop out the image containing the region of interest as the target shooting image; The analysis module is adapted to analyze the highlights of the first original video and use them as the first video segment. The extraction module is adapted to extract a second video segment from the second original video that corresponds to the timeline of the first video segment. The compositing module is adapted to determine whether the second video segment is a highlight based on the analysis module. If so, the second video segment is combined with the target shooting image to generate the processed shooting image.

9. A computing device, characterized in that, The computing device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a processor, performs the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Panoramic image shooting method, device and storage medium

    CN110430360A

  • Live broadcast picture output method and device, computer equipment and readable storage medium

    CN117979044A