Video generation method, electronic equipment, computer readable storage medium and computer program product

By generating a bullet-time center array and high-resolution video in the human-computer interaction interface, the problem of individual users being able to independently select the objects and time points of interest in sports videos is solved, realizing a personalized video experience with a 360-degree free viewpoint and improving user viewing satisfaction.

CN121218005APending Publication Date: 2025-12-26MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511525452.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult for individual users to independently select the focus and time points in 360-degree video presentation of exciting moments from events, hindering regular viewing and simple, quick personalized creation. Furthermore, the full-screen video interaction mode relies on manual perspective switching, resulting in an insufficiently outstanding video presentation effect.

Method used

By presenting video content in a human-computer interaction interface, responding to user operations to obtain video frame data, generating a bullet time center array, calculating the frame dwell time, performing recognition, cropping, and image quality enhancement, and generating high-resolution free-viewpoint videos, users can select the focus person for 360-degree viewing anytime, anywhere.

Benefits of technology

It enables users to trigger 360-degree free-view viewing anytime, anywhere by clicking or pausing, generating short videos with reasonable layout, smooth rotation, and highlighting key points, providing a personalized video experience and solving the pain points of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121218005A_ABST
    Figure CN121218005A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: presenting video content in a human-computer interaction interface, wherein the video content comprises a focus character; obtaining video frame data corresponding to the current timestamp from a video stream of the current video content in response to a free view angle video generation operation of the user for the current video content; generating a bullet time center array based on the video frame data and the focus figure; calculating the residence time of each frame based on the bullet time center array, and obtaining a first close-range bullet time video according to the residence time of each frame; performing identification cutting and image quality enhancement processing on the first close-range bullet time video to obtain a second close-range bullet time video; generating a free view angle video of the user for the current video content based on the second close-range bullet time video clip; displaying the free view angle video on the human-computer interaction interface; therefore, the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and video processing, and in particular to a video generation method, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Currently, several technical solutions exist for presenting 360-degree video of exciting moments from sporting events. One is the 3D reconstruction solution, which uses 3D reconstruction technology to digitally recreate the event scene after the event and generates multi-view virtual short videos using a 3D engine. Another is the full-fidelity video solution, which involves deploying numerous high-definition cameras around the venue and simultaneously capturing footage, stitching the video streams together into high-resolution content, allowing users to view the content from various angles through interactive operations. A third is the highlight moment generation solution, where a media server extracts frames from multiple video streams at the same moment from different perspectives and sends them to the user's device upon request. A fourth is the bullet-time video generation solution, where the server extracts frame sequences at specific times from various perspective streams to generate a surround-view video.

[0003] However, current technologies generally rely on professional teams for centralized production, making it difficult for individual users to regularly and freely choose what to watch and when to create content. Furthermore, current interaction modes primarily depend on users manually switching perspectives, lacking automatically generated engaging short video content. Summary of the Invention

[0004] To address the aforementioned technical problems, embodiments of this application provide a video generation method, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] The video generation method provided in this application includes: The video content, including the featured person, is presented in the human-computer interaction interface. In response to a user's free-viewpoint video generation operation for the current video content, video frame data corresponding to the current timestamp is obtained from the video stream of the current video content; wherein, the video frame data includes all frame images in the video stream; Based on the video frame data and the focus character, generate a bullet time center array; The dwell time of each frame is calculated based on the bullet time center array, and the first close-up bullet time video is obtained based on the dwell time of each frame. The first close-up bullet time video is subjected to recognition, cropping, and image quality enhancement processing to obtain a second close-up bullet time video; wherein, the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video; The user's free-view video of the current video content is generated based on the second close-up bullet-time video clip; The free-viewpoint video is displayed on the human-computer interaction interface.

[0006] This application provides an electronic device, the electronic device comprising: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the video generation method provided in the embodiments of this application.

[0007] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the video generation method provided in this application.

[0008] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the video generation method provided in this application.

[0009] In the technical solution of this application embodiment, video content, including a focal figure, is presented in a human-computer interaction interface; in response to the user's free-viewpoint video generation operation for the current video content, video frame data corresponding to the current timestamp is obtained from the video stream of the current video content; wherein, the video frame data includes frame images from all perspectives in the video stream; a bullet time center array is generated based on the video frame data and the focal figure; the dwell time of each frame is calculated based on the bullet time center array, and a first close-up bullet time video is obtained based on the dwell time of each frame; the first close-up bullet time video is subjected to recognition, cropping, and image quality enhancement processing to obtain a second close-up bullet time video; wherein, the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video; based on the second The system generates close-up bullet-time video clips that allow users to view the current video content from a free-viewpoint perspective. These free-viewpoint videos are then displayed on the human-computer interaction interface. This provides a simple and effective interactive experience, allowing users to trigger a 360-degree free-viewpoint viewing mode anytime, anywhere by tapping the screen or pausing while watching a live broadcast. Users can select any player they like to access a 360-degree short video pop-up of the current highlight. Furthermore, by analyzing the positions of key players on the field, the system provides specialized camera speed choreography and close-up shot production solutions. It stitches together all frames from the current perspective into a well-structured, smoothly rotating, and focused short video, addressing the pain points of current full-fledged video viewing and providing users with a better personalized experience. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the structure of the video generation system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the human-computer interaction interface provided in an embodiment of this application; Figure 5 This is a schematic diagram of a frame of video stream data provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the principle of bullet time generation provided in the embodiments of this application; Figure 7 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 8 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 9 This is a schematic diagram illustrating the principle of bullet time generation provided in the embodiments of this application; Figure 10 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 11 This is a schematic diagram illustrating the principle of the image cropping method provided in the embodiments of this application; Figure 12 This is a schematic diagram illustrating the principle of the image cropping method provided in the embodiments of this application; Figure 13 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 14 This is a schematic diagram of the human-computer interaction interface provided in an embodiment of this application; Figure 15 This is a flowchart illustrating the video generation method provided in an embodiment of this application; Figure 16 This is a schematic diagram of the human-computer interaction interface provided in an embodiment of this application; Figure 17 This is a schematic diagram of the human-computer interaction interface provided in an embodiment of this application; Figure 18 This is an overall schematic diagram of the video generation method provided in the embodiments of this application; Figure 19 This is a schematic diagram illustrating the effect of the generated video provided in the embodiments of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0013] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0014] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0015] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0016] 1) Bullet time is a visual effect created through technical means. This technique uses the synthesis of multiple cameras or computer-generated images to make viewers see an event happening on a very slow time scale. This effect allows viewers to see every detail and angle of the event, thus creating a feeling as if time has been slowed down.

[0017] 2) Full-length video is achieved by deploying a circle of high-definition cameras (approximately 36 units) around the venue to simultaneously collect video signals from around the venue. These signals are then stitched together using algorithms to create an 8K video stream, which is then sent to the terminal.

[0018] In the field of 360-degree viewing and production of exciting moments from sporting events, there are currently four main solutions: The first is a 3D reconstruction solution. This solution requires reconstructing the event scene using 3D reconstruction technology after the event and then rendering it using a 3D engine to generate 360-degree multi-view content to create virtual short videos. While professional, this method is time-consuming. The second is a full-motion video solution. This involves deploying approximately 36 high-definition cameras in a circular array at the venue to simultaneously capture video signals from around the field. These signals are then stitched together using algorithms to create an 8K video stream, which is transmitted to mobile devices. Users can then swipe their screens to view the video in real-time from a 360-degree free-view perspective. However... The interactive viewing method has limitations on the model and performance of personal terminals, resulting in some users being unable to obtain the corresponding viewing experience; the third is the highlight moment generation scheme, which obtains multi-view video streams from the media server, determines the target video frames of different perspectives at the same playback moment based on the current video frame identification information in the highlight moment generation request, and sends them to the user terminal according to the preset delivery method; the fourth is the bullet time video generation scheme, in which the media server selects the target video frames corresponding to each bullet time point from the video streams of each perspective based on the current video frame in the bullet time video generation request, and then assembles a bullet time video with clockwise or counterclockwise perspective rotation. However, existing solutions rely entirely on professional creators for content production. Individual users cannot choose which players and time points to focus on while watching the game, making it difficult to achieve regular viewing and simple, quick personalized production. Users often rely on manually switching perspectives left and right, failing to create engaging video content that can be played automatically. Furthermore, the generated videos are mostly played at a uniform clockwise or counterclockwise speed based on the user's viewing perspective, resulting in an unremarkable video presentation. They do not fully take into account the characteristics of bullet time, nor can they specifically highlight the user's focus (such as the player with the ball) at different moments on the field.

[0019] To address the aforementioned issues, embodiments of this application provide a video generation method, an electronic device, a computer-readable storage medium, and a computer program product that enable users to trigger a 360-degree viewing mode anytime, anywhere. Users can select any player they like to receive a 360-degree free-view video pop-up of the current exciting moment, providing a better personalized experience.

[0020] Taking the application of this application's embodiments to a scenario where a user watches sports videos as an example, see... Figure 1 , Figure 1 This is a schematic diagram of the video generation system architecture provided in the embodiments of this application, exemplified. Figure 1 The system involves server 100, terminal 200, and network 300. Terminal 200 is connected to server 100 through network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.

[0021] In some embodiments, the video generation system provided in this application can be implemented collaboratively by a server and a terminal. For example, in response to a user's free-viewpoint video generation operation for the current video content, the terminal 200 sends a video generation request to the server 100. After receiving the video generation request, the server 100, based on the video generation method provided in this application, sends the generated free-viewpoint video of the current video content to the terminal 200. The terminal 200 receives the free-viewpoint video of the current video content sent by the server 100 and displays the free-viewpoint video on the human-computer interaction interface.

[0022] Here, server 100 can be a single server. In this case, the video generation method provided in this application embodiment can be implemented by the same server. Server 100 can also be a cluster of servers. In the case that server 100 is a cluster of servers, the video generation method provided in this application embodiment can be implemented by different servers, and this application embodiment does not impose any limitations.

[0023] In some embodiments, the terminal or server can implement the video generation method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as game applications or live streaming applications; or they can be applets embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0024] In some embodiments, server 100 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminals.

[0025] In some embodiments, multiple servers can form a blockchain, and server 100 is a node on the blockchain. Information connections can exist between each node in the blockchain, and information can be transmitted between nodes through these connections. The data related to the video generation method provided in this application embodiment can be stored on the blockchain.

[0026] The embodiments of this application can be implemented using artificial intelligence (AI) technology. AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0027] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0028] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The electronic device 400 shown can be either the server 100 or the terminal device 200 mentioned above. Figure 2 The illustrated electronic device 400 includes at least one processor 410, a memory 430, and at least one network interface 420. The various components of the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0029] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0030] The memory 430 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 430 may optionally include one or more storage devices physically located away from the processor 410.

[0031] The memory 430 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 430 described in this application embodiment is intended to include any suitable type of memory.

[0032] In some embodiments, memory 430 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0033] Operating system 431 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 432 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A video generation apparatus 433 stored in memory 430 is shown. This apparatus can be software in the form of programs and plugins, and includes the following software modules: a data acquisition module 4331 and a data processing module 4332. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0034] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the multimodal face driving method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0035] See Figure 3 , Figure 3 This is a flowchart illustrating the video generation method provided in the embodiments of this application, which will be combined with... Figure 3 The steps shown are explained.

[0036] Step 301: Present the video content in the human-computer interaction interface. The video content includes the focus person.

[0037] In some implementations, the human-computer interaction interface can refer to video content viewed from a specific perspective on a user's playback interface, where the video content includes a focal figure. For example, when a user is watching a regular sporting event such as a 2D basketball broadcast, the focal figure could be a player holding a basketball, or any player the user likes or supports.

[0038] Step 302: In response to the user's free-viewpoint video generation operation for the current video content, obtain the video frame data corresponding to the current timestamp from the video stream of the current video content; wherein, the video frame data includes all the frames in the video stream.

[0039] In some implementations, users can trigger free-viewpoint video generation on the human-computer interaction interface in a preset manner. For example, users can trigger free-viewpoint video generation by tapping the screen or pausing the screen, so that when the terminal or mobile device detects the user's triggering of the free-viewpoint video generation operation, it obtains a free-viewpoint video generation request based on the currently displayed video content.

[0040] In some implementations, users can trigger the free-view video generation operation on the human-computer interaction interface and select a focus figure according to a preset method. For example, users can select any player they like as the focus figure by tapping the screen or pausing the screen, so that a free-view video can be generated around that focus figure later.

[0041] For example, refer to Figure 4 , Figure 4 This means that when a user enters a video app to watch a match they are interested in, during the match, by monitoring events on the playback device, a pop-up window is triggered to play a 360-degree video. When the user finds a part exciting, they will subconsciously tap the screen to pause and see the details, so the system can listen for user actions such as tapping the screen more than twice within 1 second, or tapping the pause button on the player. In response to the user's action of generating a free-view short video of the current video content, the terminal displays a prompt in the human-computer interaction interface, allowing the user to select a player to play their current highlight moment in a 360-degree video, i.e., a free-view short video.

[0042] In some implementations, the terminal responds to a user's free-view video generation operation for the current video content. Specifically, after the user selects a player as the focus, the terminal sends a free-view video generation request to the server, which then generates the free-view video. The server archives the generated free-view video stream and returns it to the terminal for display.

[0043] In some implementations, when a user performs a free-viewpoint video generation operation on the current video content, the terminal retrieves the video frame data corresponding to the current timestamp from the video stream of the current video content. Specifically, the terminal obtains the video content currently being viewed by the user, determines the timestamp, frame, and the user's potential focus character from the current video content (here, the user may not select a focus character, in which case the terminal selects the focus character). The terminal then sends the determined frame information to the server for processing.

[0044] In some implementations, when a user performs a free-viewpoint video generation operation on the current video content, the server receives a timestamp, frame images, and information about the focus person that the user may select from the terminal. The server retrieves all frames corresponding to the current timestamp from the R-channel (R>=36) video streams on site. It then matches and obtains the viewer's current viewing angle to get the starting frame for rotation.

[0045] Here, refer to Figure 5 , Figure 5This represents the total number of frames (R) from all angles at the same moment in the entire video stream. Specifically, the server obtains all frames corresponding to the current timestamp from R (R>=36) video streams from the venue. This acquisition can be done through full-scale signal acquisition. For example, approximately 36 high-definition cameras are deployed around the venue, simultaneously capturing video signals from around the entire area, which are then stitched together into an 8K video stream using an algorithm. In this way, the server can obtain frame data from a 360-degree free-view perspective at the current timestamp.

[0046] Step 303: Generate a bullet time center array based on video frame data and the focus character.

[0047] In some implementations, when the server-side produces free-viewpoint video, to achieve a more comfortable, smooth, and engaging viewing experience, it arranges all frames to produce a two-round 360-degree effect. The first round creates a uniformly moving long shot, while the second round creates a close-up, bullet-time effect. Bullet time is a visual effect created through technical means. This technique uses the synthesis of images from multiple cameras or computer-generated images to make viewers perceive an event occurring on a very slow timescale. This effect allows viewers to see every detail and angle of the event, creating the feeling that time has been slowed down.

[0048] In some implementations, refer to Figure 6 For close-up bullet-time videos, a bullet-time center array is generated to achieve three smooth slow-motion bullet-time effects, thus simultaneously improving both visual appeal and ease of movement. The bullet-time center array includes three slowest stopping points at the bullet center (N1, N2, N3). N1 represents the first-person perspective, i.e., the direct viewpoint, facing the focal character and key objects. N2 represents the second-person perspective, and N3 represents the third-person perspective. N2 and N3 are respectively rotated 120 degrees clockwise and 240 degrees clockwise from N1, ensuring the bullet-time center array maximizes detail viewing without obstructing the central character and is evenly distributed in 120-degree intervals.

[0049] In some implementations, the video content also includes key objects; see [link to relevant documentation]. Figure 7 , Figure 3 Step 303 shown can be implemented through steps 701 to 704, which are explained in detail below.

[0050] Step 701: Obtain the position coordinates of the focus person and key objects in each frame of the video frame data.

[0051] Step 702: Determine the first-view direction based on the position coordinates of the focus character and the key object; the first-view direction is the view direction facing the focus character and the key object.

[0052] Step 703: Rotate the first view direction clockwise based on the first preset threshold to obtain the second view direction, and rotate the first view direction clockwise based on the first preset threshold to obtain the third view direction.

[0053] Step 704: Generate a bullet time center array based on the first-view direction, the second-view direction, and the third-view direction.

[0054] In some implementations, positional recognition technology is used to identify the coordinates of the focal figures and key objects in the video frame data. For example, in a sports broadcast, the direction of the line connecting the player and the basket in each frame can be calculated based on existing positional recognition technology, and the viewing direction facing the player and the basket is determined as the base angle N1. Then, starting from the base angle N1, two other evenly distributed viewing directions N1, N2, and N3 are obtained by rotating clockwise by 120 degrees and 240 degrees, ensuring that the three bullet-time center points surround the player and the viewing angles are evenly distributed, avoiding occlusion and presenting the best composition.

[0055] Step 304: Calculate the dwell time of each frame based on the bullet time center array, and obtain the first close-up bullet time video based on the dwell time of each frame.

[0056] In some implementations, bullet time processing typically involves lengthening the duration of a particular frame during a uniformly played video to create the effect of a still frame. The duration of each frame is calculated using a symmetrical, smooth Gaussian distribution curve algorithm on both sides of the bullet time center array. Then, video generation techniques are used to generate a first close-up bullet time video.

[0057] In some implementations, see Figure 8 , Figure 3 Step 304 shown can be implemented through the following steps 801 to 803, which are explained in detail below.

[0058] Step 801: Determine the peak position and peak amplitude of each frame; where the peak amplitude represents the duration of the still frame corresponding to the bullet time center array, and the peak position is the bullet time center array.

[0059] Step 802: Construct a dwell time function for each frame based on the Gaussian distribution curve. Based on the dwell time function for each frame, the peak position and peak amplitude of each frame, determine the dwell time value of each frame.

[0060] Step 803: Based on the dwell time value of each frame, re-encode and re-time the video frame data to obtain the first close-up bullet time video.

[0061] Here, the dwell time function for each frame can be expressed as: (1) Among them, A i This is the amplitude of the i-th peak, which is the still frame duration corresponding to the bullet time center array (N1, N2, N3). For example, it's typically 0.5 seconds for video creation. i σ is the i-th peak position, which is the bullet time center array (N1, N2, N3); σ is the standard deviation of the Gaussian distribution, which determines the width of the peak; C is the dwell time of a normal speed frame, for example, the value corresponding to a 25FPS video is 0.04s.

[0062] For example, different peak positions, amplitudes, and standard deviations will produce different results. (See reference...) Figure 9 When calculating the dwell time of 36 cameras, A i =(0.5,0.5,0.5), x i =(6,18,30), σ=1, C=0.04. The dwell times of frames 1-36 are calculated to be [0.04, 0.042, 0.062, 0.165, 0.393, 0.54, 0.393, 0.165, 0.062, 0.042, 0.04, 0.04, 0.04, 0.042, 0.062, 0.165, 0.393, 0.54, 0.393, 0.165, 0.062, 0.042, 0.04, 0.04, 0.04, 0.042, 0.062, 0.165, 0.393, 0.54, 0.393, 0.165, 0.062, 0.042, 0.04, 0.04] (seconds).

[0063] Step 305: Perform recognition, cropping, and image quality enhancement processing on the first close-up bullet time video to obtain the second close-up bullet time video; wherein, the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video.

[0064] In some implementations, when playing close-up shots, the image needs to be magnified. Based on the selected focal person, the depth distance of the focal person in each frame is calculated, and all frames are magnified and cropped to a suitable distance. The cropped frames are then subjected to super-resolution processing to improve the image resolution to the level of the original video.

[0065] In some implementations, see Figure 10 , Figure 3Step 305 shown can be implemented through steps 1001 to 1006, as explained in detail below.

[0066] Step 1001: Use the target detection model to identify the focal person in each frame of the first close-up bullet time video and obtain the pixel distance data of each frame.

[0067] In some implementations, refer to Figure 11 The bounding box of the focal person is identified using an object detection model. For each frame of the first close-up bullet-time video, the object detection model is used to identify and process the data, and the pixel distance data (a, b, c, d) of the focal person's body bounding box from the top, bottom, left, and right edges of the screen are obtained.

[0068] Step 1002: Determine the width and height pixel values ​​and aspect ratio of the first close-up bullet time video.

[0069] In some implementations, refer to Figure 11 The image parameters are identified and processed using an object detection model. For each frame of the first close-up bullet-time video, the object detection model is used to identify and process the image parameters, obtaining the width and height pixel values ​​(m, n) and aspect ratio k of the source frame, where k = n / m.

[0070] Step 1003: Calculate and process the pixel distance data of each frame based on the aspect ratio and the first preset threshold to obtain the first ratio value; the first ratio value is used to perform global uniform magnification of each frame.

[0071] In some implementations, a minimum safe magnification ratio is calculated for each frame. To prevent the body parts of the focus figure from being incomplete or too close to the edge of the image after cropping, a safe magnification ratio needs to be calculated for each frame. Combining the aspect ratio k of the frame and the preset edge safety constant C, formula (2) is used to calculate and process each frame to obtain a minimum magnification ratio W that ensures the layout of the spheres meets the preset standard. i .

[0072] (2) Where (a, b, c, d) represent the pixel distance data of the bounding box of the focused character's body from the top, bottom, left, and right edges of the screen.

[0073] Next, a globally unified magnification ratio determination process is needed to obtain the first ratio value w. To ensure a consistent viewing angle after cropping all images, a single-frame safe magnification ratio W is calculated for all frames. i The sequence is taken to find the minimum value, and a global minimum value is calculated to obtain the final uniform magnification ratio applied to all frames, i.e., the first ratio value w=MIN(w1, w2, ...).

[0074] Step 1004: Based on the first ratio value, perform pixel cropping processing on the pixel distance data of each frame to obtain the actual cropped pixel value of each frame.

[0075] In some implementations, refer to Figure 12 Based on the first ratio value, the cropping pixel calculation of each frame is performed. According to the determined uniform magnification ratio w, the original boundary distance data (a, b, c, d) of each frame is scaled and calculated to obtain the actual cropping pixel values ​​(a1, b1, c1, d1) in the four directions of top, bottom, left and right of each frame.

[0076] Step 1005: Perform pixel cropping on each frame of the first close-up bullet time video according to the actual cropping pixel values ​​to obtain the first close-up bullet time video after pixel cropping.

[0077] In some implementations, dynamic image cropping is performed. Based on the calculated cropping pixels (a1, b1, c1, d1) of each frame, pixel cropping is performed on each frame of the first close-up bullet time video to obtain a set of intermediate resolution frame sequences with uniform shot size, reasonable spherical layout and conforming to preset composition standards.

[0078] Step 1006: Perform image quality enhancement processing on the first close-up bullet time video after pixel cropping to obtain the second close-up bullet time video.

[0079] In some implementations, image super-resolution enhancement processing is performed. The intermediate resolution frame sequence, after cropping and reducing its resolution, is enhanced using a video super-resolution model to obtain a set of high-definition close-up frame sequences with resolution increased to the same level as the original video, which are then used for final video compositing.

[0080] Step 306: Generate a free-view video of the current video content based on the second close-up bullet time video clip.

[0081] In some implementations, execution Figure 3 Before step 306 shown, the video frame data can also be processed by playing the frame sequence at a constant speed to obtain a distant view video.

[0082] In free-viewpoint video production, excessively fast camera movement can cause visual discomfort; therefore, a preset short video frame rate of 25 FPS can be implemented. To deliver a more comfortable, smooth, and engaging video viewing experience, a set of automatic preview playback or short video creation and sharing techniques is provided, consisting of two loops. This means the arrangement produces two loops of 360-degree effects, with the first loop creating a uniformly moving distant view. For example, the first loop effect is a uniformly moving playback of all frames of the distant view, with a single frame dwell time of 0.04 seconds. The total video duration of the first loop is t1 = 0.04 × R, where R represents the total number of frames in one loop. This application does not limit the specific implementation method of processing the video frame data into a uniformly moving frame sequence.

[0083] In some implementations, a long-range view video and a second close-up bullet-time video are stitched together to generate a free-view video of the user's favorite moments from the current video content.

[0084] Here, a dual-circle video concatenation and synthesis process is performed. The uniform-speed distant view video generated in the first circle is spliced ​​with the ultra-high-definition close-up bullet time video obtained in the second circle, i.e., the second close-up bullet time video, to finally obtain a 360-degree free-view video with a reasonable layout, smooth perspective rotation, and a focus on close-up of the main characters.

[0085] Step 307: Display the free-view video on the human-computer interaction interface.

[0086] In some implementations, the server sends a short, free-view video of the user capturing the highlights of the current video content to the terminal app. After receiving the video stream from the server, the app can play the video in a pop-up video playback window in the lower right corner of the event video.

[0087] As can be seen from the above, the video generation method provided in this application presents video content, including a focal person, in a human-computer interaction interface; in response to a user's free-viewpoint short video generation operation for the current video content, it obtains video frame data corresponding to the current timestamp from the video stream of the current video content; wherein, the video frame data includes all frames in the video stream; based on the video frame data and the focal person, it generates a bullet time center array; based on the bullet time center array, it calculates the dwell time of each frame, and obtains a first close-up bullet time video based on the dwell time of each frame; it performs recognition, cropping, and image quality enhancement processing on the first close-up bullet time video to obtain a second close-up bullet time video; wherein, the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video. Based on second-close-up bullet-time video clips, a free-viewpoint short video is generated for the user based on the current video content. This free-viewpoint short video is then displayed on the human-computer interaction interface. In this way, on the one hand, a simple and effective interactive experience is proposed. While watching a live sports broadcast, users can trigger a 360-degree free-viewpoint viewing mode anytime, anywhere by clicking the screen or pausing. Users can select any player they like to get a 360-degree short video pop-up window showing the current highlight. On the other hand, by analyzing the position of key players on the field, a specialized camera speed arrangement and close-up shot production scheme is provided. All frames from the current perspective are stitched together to create a short video with a reasonable layout, smooth rotation, and prominent focus, solving the pain points of current full-fledged video viewing and providing users with a better personalized experience.

[0088] See Figure 13 , Figure 13 This is a flowchart illustrating the video generation method provided in the embodiments of this application, which will be combined with... Figure 13 The steps shown are explained.

[0089] Step 1301: In response to the user's swipe to watch the video from a free-viewpoint perspective, obtain the swipe playback speed controlled by the user.

[0090] Step 1302: Display the free-view video on the human-computer interaction interface according to the user-controlled sliding playback speed.

[0091] Here, refer to Figure 14 Users can control the viewing angle of the event by swiping left and right on the generated short video interface, or browse directly according to the speed and rhythm of the bullet time video produced by the server. In this way, when watching regular sports broadcasts, users can trigger the 360-degree viewing mode anytime, anywhere by clicking the screen or pausing. Users can select any player they like to get a 360-degree short video pop-up window of the current highlight moment, which will play automatically. Users can also freely drag and rotate the view to view technical details.

[0092] See Figure 15 , Figure 15 This is a flowchart illustrating the video generation method provided in the embodiments of this application, which will be combined with... Figure 15 The steps shown are explained.

[0093] Step 1501: Count the number of free-view videos shown to users. The number of free-view videos is used to send prompts to users to create and share videos.

[0094] Here, refer to Figure 16 Users can share videos from a free-viewpoint perspective.

[0095] Step 1502: In response to the user's selection of the prompt to create and share the video, call the editing service to edit and process multiple free-viewpoint videos to obtain a composite free-viewpoint video.

[0096] Here, refer to Figure 17 After the event ends, the number of highlights watched by users is counted, and a prompt appears asking if users want to create and share short videos. If users select yes, they can use the editing service to edit multiple free-viewpoint short videos into a single video.

[0097] Below, we will combine Figure 18 and Figure 10 This application illustrates an exemplary application of the embodiments of this application in a 2D sports video broadcasting scenario. Figure 18 This is a schematic diagram of the system framework for the video generation method provided in this application embodiment. Figure 19 This is a 360-degree effect illustration of the video generation method provided in this application embodiment. A user is watching a 2D sports broadcast video on a mobile app and selects a player in a free-viewpoint short video generation operation for the current video content. The terminal responds to the user's free-viewpoint short video generation operation by sending the current timestamp information and player information to the server. The server responds to the user's free-viewpoint short video generation operation by retrieving the video frame data corresponding to the current timestamp from the video stream of the current video content, i.e., retrieving 36 video feeds from the same viewpoint at the same time; at this time, the video frame data includes the frame images from the 36 viewpoints corresponding to the current timestamp.

[0098] When the server generates the video, it focuses on appropriate camera speed and reasonable close-up positions of the main characters. Specifically, after acquiring frame data, the server first processes the camera speed. To provide a smoother and more engaging video viewing experience, it offers users an automatic preview playback or short video creation and sharing scheme, consisting of two rounds of 360-degree effects. The first round creates a uniformly moving long shot, and the second round creates a close-up bullet-time effect. To achieve both simultaneous effects and comfortable movement, this round of video includes three smooth slow-motion bullet-time effects. After camera speed processing, close-up processing is performed. Based on the selected spherical figures, the depth distance of the spherical figures in each frame is intelligently calculated, and all frames are proportionally enlarged and cropped to an appropriate distance. The cropped frames undergo super-resolution processing, increasing the resolution to the original video level. Afterward, a 360-degree short video pop-up window of the current highlight moment is generated, playing automatically. Users can freely drag and rotate the view to view technical details. For favorite moments, users can automatically create 360-degree short videos for distribution or setting as ringtones.

[0099] As described above, this application provides a video generation method and proposes a simple and effective interactive experience. When users are watching a live broadcast of an exciting sporting event, they can trigger a 360-degree viewing mode anytime, anywhere by clicking the screen or pausing. Selecting any player they like will generate a 360-degree short video pop-up of the current exciting moment. Users can freely drag and rotate the view to view technical details. For favorite moments, 360-degree short videos can be automatically created for distribution or as ringtones. This solves the pain points of current full-screen video viewing and provides users with a better experience. On the other hand, by analyzing the positions of the focus player and key on-court landmarks (such as the basketball hoop) and calculating the camera distance of the focus player, a specialized camera speed arrangement algorithm and close-up shot production scheme are provided. This stitches together all frames from the current perspective into a short video with a reasonable layout, smooth rotation, and prominent focus. This also solves the pain points of current full-screen video viewing and provides users with a better personalized experience.

[0100] The following description continues to illustrate the exemplary structure of the video generation device 433 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software module stored in the video generation device 433 of the memory 430 may include: The data acquisition module 4331 is used to respond to the user's free-viewpoint short video generation operation for the current video content, and to acquire video frame data corresponding to the current timestamp from the video stream of the current video content; wherein, the video frame data includes frame images of all viewpoints corresponding to the current timestamp in the video stream.

[0101] The data processing module 4332 is used to generate a bullet time center array based on the video frame data and the focus character; calculate the dwell time of each frame based on the bullet time center array, and obtain a first close-up bullet time video based on the dwell time of each frame; perform recognition, cropping and image quality enhancement processing on the first close-up bullet time video to obtain a second close-up bullet time video; wherein the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video; and generate a short video of the user's free-viewpoint perspective on the current video content based on the second close-up bullet time video clip.

[0102] In some embodiments, the data acquisition module 4331 is further configured to acquire the position coordinates of the focal person and the position coordinates of the key object in each frame of the video frame data. The data processing module 4332 is further configured to determine a first viewing direction based on the position coordinates of the focal person and the position coordinates of the key object; the first viewing direction is a viewing direction facing the focal person and the key object; the first viewing direction is rotated clockwise based on a first preset threshold to obtain a second viewing direction, and the first viewing direction is rotated clockwise based on the first preset threshold to obtain a third viewing direction; the bullet time center array is generated based on the first viewing direction, the second viewing direction, and the third viewing direction.

[0103] In some implementations, the data processing module 4332 is further configured to determine the peak position and peak amplitude of each frame; wherein the peak amplitude represents the duration of the still frame corresponding to the bullet time center array, and the peak position is the bullet time center array; construct a dwell time function for each frame based on a Gaussian distribution curve, and determine the dwell time value of each frame based on the dwell time function, the peak position and peak amplitude of each frame; and re-encode and re-time the video frame data according to the dwell time value of each frame to obtain the first close-up bullet time video.

[0104] In some embodiments, the data processing module 4332 is further configured to: identify the focal person in each frame of the first close-up bullet time video using a target detection model to obtain pixel distance data for each frame; determine the width and height pixel values ​​and aspect ratio of the first close-up bullet time video; calculate the pixel distance data of each frame based on the aspect ratio and a first preset threshold to obtain a first ratio value; use the first ratio value to globally and uniformly magnify each frame; perform pixel cropping processing on the pixel distance data of each frame based on the first ratio value to obtain the actual cropped pixel value for each frame; perform pixel cropping processing on each frame of the first close-up bullet time video according to the actual cropped pixel value to obtain the pixel-cropped first close-up bullet time video; and perform image quality enhancement processing on the pixel-cropped first close-up bullet time video to obtain the second close-up bullet time video.

[0105] In some embodiments, the data processing module 4332 is further configured to perform frame sequence uniform playback processing on the video frame data to obtain a distant view video.

[0106] In some implementations, the data processing module 4332 is further configured to stitch together the distant-view video and the second close-up bullet-time video to generate a free-view short video of the user's highlights of the current video content.

[0107] In some implementations, the data acquisition module 4331 is further configured to acquire the user-controlled sliding playback speed in response to the user's sliding video viewing operation on the free-view short video.

[0108] In some implementations, the data processing module 4332 is also used to display the free-view short video on the human-computer interaction interface according to the user-controlled sliding playback speed.

[0109] In some implementations, the data processing module 4332 is further configured to count the number of free-viewpoint short videos shown to the user, the number of free-viewpoint short videos being used to send a prompt to the user to create and share a short video; in response to the user's selection operation on the prompt to create and share a short video, the module calls an editing service to edit multiple free-viewpoint short videos to obtain multiple composite free-viewpoint short videos.

[0110] Those skilled in the art should understand that Figure 2 The functions of each unit in the video generation device shown can be understood by referring to the relevant descriptions of the aforementioned method. Figure 2The functions of each unit in the video generation device shown can be implemented by a program running on a processor or by specific logic circuits.

[0111] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0112] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0113] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0114] This application also provides a computer-readable storage medium for storing computer programs.

[0115] Optionally, the computer-readable storage medium can be applied to the network device in the embodiments of this application, and the computer program causes the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0116] Optionally, the computer-readable storage medium can be applied to the mobile terminal / terminal device in the embodiments of this application, and the computer program causes the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0117] This application also provides a computer program product, including computer program instructions.

[0118] Optionally, the computer program product can be applied to the network device in the embodiments of this application, and the computer program instructions cause the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of this application. For the sake of brevity, they will not be described in detail here.

[0119] Optionally, the computer program product can be applied to the mobile terminal / terminal device in the embodiments of this application, and the computer program instructions cause the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of this application. For the sake of brevity, they will not be described in detail here.

[0120] This application also provides a computer program.

[0121] Optionally, the computer program can be applied to the network device in the embodiments of this application. When the computer program is run on the computer, it causes the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0122] Optionally, the computer program can be applied to the mobile terminal / terminal device in the embodiments of this application. When the computer program is run on a computer, it causes the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0123] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0124] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0128] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A video generation method, characterized in that, The method includes: The video content, including the featured person, is presented in the human-computer interaction interface. In response to a user's free-viewpoint video generation operation for the current video content, video frame data corresponding to the current timestamp is obtained from the video stream of the current video content; wherein, the video frame data includes frame images of all viewpoints corresponding to the current timestamp in the video stream; Based on the video frame data and the focus character, generate a bullet time center array; The dwell time of each frame is calculated based on the bullet time center array, and the first close-up bullet time video is obtained based on the dwell time of each frame. The first close-up bullet time video is subjected to recognition, cropping, and image quality enhancement processing to obtain a second close-up bullet time video; wherein, the resolution of the second close-up bullet time video is higher than that of the first close-up bullet time video; The user's free-view video of the current video content is generated based on the second close-up bullet-time video clip; The free-viewpoint video is displayed on the human-computer interaction interface.

2. The method according to claim 1, characterized in that, The video content also includes key objects; the generation of a bullet-time center array based on the video frame data and the focal figure includes: Obtain the position coordinates of the focus person and the key object in each frame of the video frame data; Based on the position coordinates of the focal person and the position coordinates of the key object, a first viewing direction is determined; the first viewing direction is the viewing direction facing the focal person and the key object. The first viewpoint direction is rotated clockwise based on a first preset threshold to obtain the second viewpoint direction, and the first viewpoint direction is rotated clockwise based on the first preset threshold to obtain the third viewpoint direction; The bullet time center array is generated based on the first view direction, the second view direction, and the third view direction.

3. The method according to claim 2, characterized in that, The step of calculating the dwell time of each frame based on the bullet time center array, and obtaining the first close-up bullet time video based on the dwell time of each frame, includes: Determine the peak position and peak amplitude of each frame; wherein, the peak amplitude represents the duration of the still frame corresponding to the bullet time center array, and the peak position is the bullet time center array; A dwell time function for each frame is constructed based on the Gaussian distribution curve. The dwell time value for each frame is determined based on the dwell time function, the peak position and peak amplitude of each frame. Based on the dwell time value of each frame, the video frame data is re-encoded and re-timing processed to obtain the first close-up bullet time video.

4. The method according to claim 3, characterized in that, The process of recognizing, cropping, and enhancing the image quality of the first close-up bullet time video to obtain the second close-up bullet time video includes: The focus person in each frame of the first close-up bullet time video is identified by the target detection model to obtain the pixel distance data of each frame. Determine the width and height pixel values ​​and aspect ratio of the first close-up bullet time video; The pixel distance data of each frame is calculated based on the aspect ratio and the first preset threshold to obtain a first ratio value; the first ratio value is used to perform global uniform magnification on each frame. Based on the first ratio value, the pixel distance data of each frame is processed to crop pixels to obtain the actual cropped pixel value of each frame. The first close-up bullet time video is pixel-cropped for each frame based on the actual cropped pixel value to obtain the first close-up bullet time video after pixel cropping. The first close-up bullet time video after pixel cropping is subjected to image quality enhancement processing to obtain the second close-up bullet time video.

5. The method according to claim 1, characterized in that, Before generating the user's free-view video of the current video content based on the second close-up bullet-time video clip, the method further includes: The video frame data is processed by playing the frame sequence at a constant speed to obtain a distant view video; The process of generating the user's free-view video of the current video content based on the second close-up bullet-time video clip includes: The distant-view video and the second close-up bullet-time video are stitched together to generate the user's free-view video of the current video content.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to the user's swipe to watch the video from the free-viewpoint perspective, the swipe playback speed controlled by the user is obtained; The free-view video is displayed on the human-computer interaction interface according to the user-controlled sliding playback speed.

7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The number of free-view videos shown to the user is counted, and the number of free-view videos is used to send the user a prompt to create and share a video; In response to the user's selection operation for the prompt to create and share a video, the editing service is invoked to edit multiple free-viewpoint video segments to obtain a composite free-viewpoint video.

8. An electronic device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.