Video processing method and related product thereof

By generating a 3D exploration space within the video scene, users can freely explore and interact with the content, solving the problem of the limited interaction methods in traditional video and enhancing user experience and engagement.

CN121940600APending Publication Date: 2026-04-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In traditional video content production methods, the interaction between users and videos is limited, with users only passively receiving the camera's perspective and lacking interactivity.

Method used

This paper provides a video processing method that generates and displays a three-dimensional exploration space between a terminal device and a server, allowing users to freely explore the video scene, including switching perspectives, manipulating objects, and adding data, thereby enhancing interactivity.

Benefits of technology

It enhances user interaction and engagement with videos, increasing user interest and engagement in watching videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940600A_ABST
    Figure CN121940600A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and related products thereof, the method is applied to terminal equipment, the method comprises the following steps: playing a video, the video comprising a plurality of video pictures; in response to a triggering operation for scene exploration of a target video picture in the video, displaying a three-dimensional exploration space of a picture scene where the target video picture is located; and according to the exploration operation in the three-dimensional exploration space, displaying a three-dimensional scene picture obtained by exploration. By adopting the method and the device, the richness of the mode of interaction with the video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and more particularly to a video processing method and related products. Background Technology

[0002] With the rapid development of multimedia technology, video content has become an important way for people to obtain information and enjoy entertainment. However, traditional video content production methods, whether pre-recorded by a camera or live-streamed, suffer from fixed perspectives and low user participation. Users often passively accept the camera's shooting angle from the video creator, resulting in a very limited way for users to interact with the video. Summary of the Invention

[0003] This application provides a video processing method and related products that can enhance the richness of ways to interact with videos.

[0004] This application provides a video processing method applied to a terminal device, the method comprising:

[0005] Play a video containing multiple video frames;

[0006] In response to a triggered operation of scene exploration of the target video frame in the video, a three-dimensional exploration space of the scene where the target video frame is located is displayed;

[0007] Based on the exploration operations in the three-dimensional exploration space, the obtained three-dimensional scene image is displayed.

[0008] This application provides a video processing method applied to a server, the method comprising:

[0009] Acquire the target video frame sent by the terminal device. The target video frame is the video frame in the video played by the terminal device that triggers scene exploration.

[0010] Obtain scene information associated with the scene where the target video frame is located;

[0011] Based on the target video footage and scene information, generate the first display data of the three-dimensional exploration space of the scene where the target video footage is located;

[0012] The first display data is sent to the terminal device, enabling the terminal device to display the three-dimensional exploration space based on the first display data.

[0013] This application provides a video processing apparatus for use in a terminal device, the apparatus comprising:

[0014] The playback module is used to play videos, which contain multiple video frames;

[0015] The display module is used to respond to the scene exploration trigger operation of the target video frame in the video and display the three-dimensional exploration space of the scene where the target video frame is located.

[0016] The exploration module is used to display the 3D scene obtained from the exploration operations in the 3D exploration space.

[0017] Optionally, exploration actions include perspective switching.

[0018] The exploration module displays the resulting 3D scene visuals based on exploration operations within the 3D exploration space, including:

[0019] Based on the perspective switching operation in the 3D exploration space, the 3D scene image after the perspective switching in the 3D exploration space is displayed.

[0020] Optionally, the trigger method for the view switching operation can be any of the following:

[0021] Triggering methods based on terminal interface, triggering methods based on virtual reality device connected to terminal device, triggering methods based on keyboard connected to terminal device, triggering methods based on mouse connected to terminal device, and triggering methods based on positional changes of virtual image in three-dimensional exploration space.

[0022] Among them, virtual avatars are images set up in the three-dimensional exploration space for scene exploration.

[0023] Optionally, the three-dimensional exploration space includes a virtual avatar set up for scene exploration, and the three-dimensional exploration space includes three-dimensional objects generated for the scene where the target video screen is located. The exploration operation includes the operation of controlling the virtual avatar to enter the three-dimensional objects in the three-dimensional exploration space.

[0024] The exploration module displays the resulting 3D scene visuals based on exploration operations within the 3D exploration space, including:

[0025] In response to the operation of controlling the virtual avatar to enter the target 3D object in the 3D exploration space, the 3D scene screen after the virtual avatar enters the target 3D object is displayed.

[0026] Optionally, the 3D exploration space includes 3D objects generated from the scene where the target video image is located, and the exploration operation includes triggering operations on the 3D objects in the 3D exploration space.

[0027] The exploration module displays the resulting 3D scene visuals based on exploration operations within the 3D exploration space, including:

[0028] Based on the triggering operation on the three-dimensional object in the three-dimensional exploration space, the three-dimensional scene screen is displayed after the three-dimensional object has been processed according to the triggering method;

[0029] Among them, the triggering operations on three-dimensional objects in the three-dimensional exploration space include at least one of the following: movement operation, deletion operation, editing operation, and cutout operation.

[0030] Optionally, exploration operations include adding data to the three-dimensional exploration space;

[0031] The exploration module displays the resulting 3D scene visuals based on exploration operations within the 3D exploration space, including:

[0032] Based on the addition of target data in the 3D exploration space, the 3D scene image after the target data has been added is displayed at the corresponding location in the 3D exploration space.

[0033] Optionally, if the video is playing when scene exploration is triggered, the target video frame is the video frame that the video is playing at the time scene exploration is triggered; and,

[0034] If the video is paused when scene exploration is triggered, the target video frame is the frame where the video was paused when scene exploration was triggered.

[0035] Optionally, in response to a triggered operation of scene exploration of a target video frame in the video, the display module displays the three-dimensional exploration space of the scene where the target video frame is located, including:

[0036] In response to a scene exploration trigger, display the scene loading progress;

[0037] When the scene loading progress indicator shows that loading is complete, the 3D exploration space is displayed.

[0038] Optionally, the above playback module is used for:

[0039] When a scene exploration of the target video frame is triggered, pause video playback;

[0040] After displaying the 3D exploration space, the aforementioned playback module is also used for:

[0041] In response to the action of closing the 3D exploration space, the 3D exploration space is closed, and the video continues to play.

[0042] Optionally, the scene exploration of the target video frame is triggered by the first object, and the video processing device further includes a sharing module, which is used for:

[0043] Based on the sharing operation of the three-dimensional exploration space, obtain the sharing information of the three-dimensional exploration space; use the shared information to invite a second object to conduct collaborative exploration of the three-dimensional exploration space.

[0044] Optionally, the video processing apparatus further includes a storage module, which is used for:

[0045] When a save operation for the 3D exploration space is obtained, the 3D exploration space is exported and saved.

[0046] This application provides a video processing apparatus applied to a server, the apparatus comprising:

[0047] The first acquisition module is used to acquire the target video frame sent by the terminal device. The target video frame is the video frame in the video played by the terminal device that is triggered for scene exploration.

[0048] The second acquisition module is used to acquire scene information associated with the scene where the target video frame is located;

[0049] The generation module is used to generate the first display data of the three-dimensional exploration space of the scene where the target video image is located, based on the target video image and scene information;

[0050] The sending module is used to send the first display data to the terminal device, so that the terminal device can display the three-dimensional exploration space based on the first display data.

[0051] Optionally, scene information associated with the scene where the target video frame is located includes at least one of the following:

[0052] Other video frames in the same scene as the target video frame, video frames adjacent to the target video frame, and text description information related to the scene where the target video frame is located.

[0053] Optionally, the generation module generates the first display data of the 3D exploration space of the scene where the target video frame is located based on the target video frame and scene information, including:

[0054] Obtain a 3D model;

[0055] The 3D model is invoked to perform 3D modeling based on the target video frame and scene information, generating 3D modeling information of the scene where the target video frame is located;

[0056] Based on the 3D modeling information, the first display data of the 3D exploration space is generated.

[0057] Optionally, the generation module can generate the first display data of the 3D exploration space based on the 3D modeling information in the following ways:

[0058] Obtain rendering performance information from the terminal device;

[0059] If the rendering performance information indicates that the rendering performance of the terminal device is sufficient, then the 3D modeling information will be used as the first display data.

[0060] If the rendering performance information indicates that the rendering performance of the terminal device is insufficient, then a 3D scene of the 3D exploration space is rendered based on the 3D modeling information, and the rendered 3D scene is used as the first display data.

[0061] Optionally, if the first display data is 3D modeling information, the terminal device is used to render a 3D scene of the 3D exploration space using the received 3D modeling information, and to display the 3D exploration space based on the rendered 3D scene.

[0062] If the first display data is a 3D scene image rendered by the server, the terminal device is used to display a 3D exploration space based on the 3D scene image sent by the server.

[0063] Optionally, the above-mentioned video processing apparatus is also used for:

[0064] Receive exploration instructions sent by the terminal device, which are generated by the terminal device based on exploration operations in the displayed three-dimensional exploration space;

[0065] Based on the exploration command, generate second display data of the three-dimensional scene explored in the three-dimensional exploration space;

[0066] The second display data is sent to the terminal device, which then displays the three-dimensional scene obtained in the three-dimensional exploration space based on the second display data.

[0067] This application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the method of this application.

[0068] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method described in the above-mentioned aspect.

[0069] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in the various alternative embodiments described above.

[0070] In this application, the terminal device can play a video containing multiple video frames. In response to a trigger operation for scene exploration of a target video frame within the video, the terminal device can display a three-dimensional exploration space of the scene containing the target video frame. Thus, the terminal device can display the corresponding three-dimensional scene obtained through the user's exploration operations within this three-dimensional exploration space. Therefore, the method proposed in this application supports users in exploring the scene containing any video frame (such as a target video frame) in a three-dimensional space during video playback. This enhances the richness of the interaction between the user and the video, increases user engagement during video playback, and increases user viewing stickiness. Attached Figure Description

[0071] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application;

[0073] Figure 2 This is a schematic diagram illustrating the effect of scene exploration provided in an embodiment of this application;

[0074] Figure 3 This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0075] Figure 4 This is a schematic diagram of a video playback interface provided in an embodiment of this application;

[0076] Figure 5 This is a schematic diagram of a scene loading interface provided in an embodiment of this application;

[0077] Figure 6 This is a schematic diagram of the interface of a three-dimensional scene screen for a three-dimensional exploration screen provided in an embodiment of this application;

[0078] Figure 7 This is a flowchart illustrating another video processing method provided in an embodiment of this application;

[0079] Figure 8 This is a scene diagram illustrating how to acquire a 3D scene image according to an embodiment of this application;

[0080] Figure 9This is a schematic diagram of a process for generating a three-dimensional scene image provided in an embodiment of this application;

[0081] Figure 10 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0082] Figure 11 This is a schematic diagram of another video processing device provided in an embodiment of this application;

[0083] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0084] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0085] First, it should be noted that all data collected in this application (such as videos, video footage, exploration operations in the 3D exploration space, scene information, and other related data) were collected with the consent and authorization of the data's owner (such as users, organizations, or enterprises), and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions.

[0086] Here, the relevant technical concepts involved in this application are explained:

[0087] 3D: Three-dimensional, a concept of space.

[0088] 3D engine (such as a 3D game engine): A software framework that provides various tools and functions needed to create and develop 3D games. It simplifies the game development process, allowing developers to focus on game design and content creation without having to write all the underlying code from scratch. A 3D engine typically includes modules for graphics rendering, physics simulation, audio processing, artificial intelligence, and network communication.

[0089] Please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application. Figure 1 As shown, this network architecture may include a server 200 and a terminal device cluster. The terminal device cluster may include one or more terminal devices; the number of terminal devices is not limited here. Figure 1 As shown, multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3, ..., terminal device n; as... Figure 1 As shown, terminal device 1, terminal device 2, terminal device 3, ..., terminal device n can all connect to server 200 via the network, so that each terminal device can interact with server 200 through the network connection.

[0090] like Figure 1 The server 200 shown can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal device can be a smartphone, tablet, laptop, desktop computer, smart TV, in-vehicle terminal, smart home device, or other smart terminal. The following description uses the communication between terminal device 1 and server 200 as an example to illustrate the specific implementation of this application.

[0091] Please see also Figure 2 , Figure 2 This is a schematic diagram illustrating the effect of scene exploration provided in an embodiment of this application. For example... Figure 2 As shown, terminal device 1 can play videos and display the video playback interface (i.e., the interface for playing the video). During video playback, terminal device 1 can respond to a scene exploration trigger operation for a target video frame in the video and display a three-dimensional exploration space of the scene where the target video frame is located. The target video frame can be any video frame in the playing video, such as the video frame that the video played to when scene exploration was triggered, or a video frame in the video.

[0092] When terminal device 1 receives a trigger operation for scene exploration of the target video frame, it can send the corresponding trigger command to server 200, enabling server 200 to generate display data related to the three-dimensional exploration space of the scene where the target video frame is located (as described below). Figure 7 (corresponding to the first display data in the embodiment), so that the server 200 can send the display data to the terminal device 1, and the terminal device 1 can use the received display data to show the three-dimensional exploration space of the scene where the target video screen is located.

[0093] Furthermore, after displaying the three-dimensional exploration space, terminal device 1 can also respond to exploration operations in the three-dimensional exploration space by displaying a three-dimensional scene obtained from the exploration. This obtained three-dimensional scene can also be generated by the terminal device sending exploration commands corresponding to the exploration operations to server 200, which then generates corresponding display data for the three-dimensional exploration space based on these commands (as described below). Figure 7 (corresponding to the second display data in the embodiment), and can send the display data to the terminal device 1, so that the terminal device 1 can display the three-dimensional scene image obtained by exploring the three-dimensional exploration space through the received display data.

[0094] like Figure 2 As shown, for example, the exploration operation of the three-dimensional exploration space may include a perspective switching operation. After receiving the perspective switching operation, the terminal device 1 can display the three-dimensional scene after the perspective of the three-dimensional exploration space has been switched. The exploration operation of the three-dimensional exploration space may also include an object deletion operation. After receiving the object deletion operation, the terminal device 1 can display the three-dimensional scene after deleting the object specified by the user in the three-dimensional exploration space. Here, the object deleted by the user may be a table lamp circled within the dotted line in the three-dimensional exploration space. The exploration operation of the three-dimensional exploration space may also include a data addition operation. After receiving the data addition operation, the terminal device 1 can display the three-dimensional scene after data addition in the three-dimensional exploration space. Here, the data added by the user may be an advertisement for "Delicious Candy" circled within the dotted line. The advertisement may be in any data form, such as image data or product links.

[0095] Using the method provided in this application, users can freely explore the three-dimensional exploration space of the scene where the video is playing at any time while watching the video, which enhances the richness of interaction between the user and the video being watched, thereby increasing the user's interest in watching the video and improving the user experience.

[0096] Please see Figure 3 , Figure 3 This is a flowchart illustrating a video processing method provided in an embodiment of this application. The executing entity in this embodiment can be a terminal device (which can be any user's terminal device).

[0097] like Figure 3 As shown, the method may include:

[0098] Step S101: Play the video, which contains multiple video frames.

[0099] Optionally, the terminal device can play videos. For example, a user can trigger a video playback operation on the terminal device, and the terminal device can respond to the operation by playing the video triggered by the user. The video played can be any video that can be played on the terminal device.

[0100] The video can be any type of video, such as a film or television drama, a live stream, or a short video. It can be a live recording or a video generated using video generation techniques (such as animated videos). The video can contain multiple video frames, each of which can be a single video frame. The specific number of video frames in the video can be determined based on the actual application scenario, and this application does not impose any restrictions on this.

[0101] Optionally, the video can be played in its own video client, and the following operations performed in this application can also be performed in that video client.

[0102] Step S102: In response to the triggered operation of scene exploration of the target video frame in the video, the three-dimensional exploration space of the scene where the target video frame is located is displayed.

[0103] Optionally, during video playback, users can trigger scene exploration operations on video frames (such as target video frames) at any time. The terminal device can respond to the user's trigger operation on scene exploration of the target video frame and display the three-dimensional exploration space of the scene where the target video frame is located.

[0104] The target video frame can be any frame in the video, and it can be the video frame from which the scene exploration is triggered. Optionally, if the video is playing when the scene exploration is triggered (i.e., the video is currently playing), then the target video frame can be the video frame currently being played when the scene exploration is triggered.

[0105] If the video is paused when the scene exploration is triggered (e.g., the video is paused during playback but is still on the playback screen), then the target video screen can be the video screen where the video is paused when the scene exploration is triggered, that is, the video screen where the video is paused when the pause is clicked.

[0106] The scene exploration trigger operation for the target video frame can be executed in any feasible manner. For example, a scene exploration control can exist in the video playback interface (i.e., the interface on the terminal device where the video is played). The scene exploration trigger operation for the target video frame can be a click operation on the scene exploration control, or it can be executed through voice control or gestures, and so on. How the trigger operation is specifically executed can be determined according to the actual application scenario, and this application does not impose any restrictions on it.

[0107] Please see Figure 4 , Figure 4 This is a schematic diagram of a video playback interface provided in an embodiment of this application. Figure 4 As shown, the video playback interface can include a "Explore" control. Users can click this control to begin exploring the 3D scene of the currently viewed or playing video frame. Clicking the "Explore" control triggers the scene exploration of the target video frame.

[0108] More specifically, the scene in which the target video frame is located can refer to the entire scene in which the objects in the target video frame exist, such as an indoor scene, an outdoor scene, a real scene (such as the shooting scene where the target video frame was located when the video was recorded), or a virtual scene (such as the scene created for the target video frame when the video was generated), and so on.

[0109] For example, the target video footage could be footage of an indoor interview with a celebrity. The scene in which the target video footage is located could be the entire indoor scene where the celebrity is being interviewed. This scene could include the celebrity, the people around the celebrity (such as the host interviewing the celebrity), and the objects around the celebrity (such as furniture, decorations, floors, walls, and all the objects in the room where the celebrity is located), and so on.

[0110] Therefore, the three-dimensional exploration space of the scene where the target video is located can be a three-dimensional (i.e., 3D) virtual space generated for the scene where the target video is located. This three-dimensional exploration space can contain three-dimensional objects generated from the objects in the scene where the target video is located. The scene presented by this three-dimensional exploration space can be called a three-dimensional scene. This three-dimensional scene is a three-dimensional virtual scene of the scene where the target video is located. Through this three-dimensional exploration space, the scenes from various perspectives in the scene where the target video is located can be presented. Then, users can freely and comprehensively explore the entire scene where the target video is located through the displayed three-dimensional exploration space.

[0111] Furthermore, in response to a trigger operation that initiates scene exploration of the target video frame, the terminal device can first display the scene loading progress. This progress could be a progress bar indicating the loading of the 3D exploration space of the scene containing the target video frame. The scene loading progress can also be indicated using elements related to the theme of scene exploration in this application, such as travel-related elements (e.g., airplanes, cars) to indicate the current scene loading progress, thereby creating and enhancing the atmosphere for the user to explore the scene containing the current video frame.

[0112] Therefore, when the loading progress of the scene is displayed as completed (e.g., when the loading progress is displayed as 100%), it indicates that the three-dimensional exploration space of the scene where the target video is located has been loaded, and the terminal device can then display the three-dimensional exploration space of the scene where the loaded target video is located.

[0113] Please see Figure 5 , Figure 5 This is a schematic diagram of a scene loading interface provided in an embodiment of this application. The terminal device can display this scene loading interface in response to a trigger operation for scene exploration of the target video frame. The scene loading interface may include a scene loading progress bar, which reflects the loading progress of the three-dimensional exploration space of the scene where the target video frame is located (i.e., the aforementioned scene loading progress). The scene loading progress bar can be represented by a small vehicle. When the vehicle reaches the "End" sign at the end of the progress bar, it indicates that scene loading is complete, and the terminal device can then display the loaded three-dimensional exploration space.

[0114] The scene loading interface can also include a "Don't Go Anymore" control, which users can click at any time to end the current scene exploration and return to the video playback interface to continue playing and watching the video.

[0115] Optionally, the scene loading progress can be displayed in a floating layer (such as a floating window) independent of the video playback interface, or it can be displayed in another terminal interface outside the playback interface, or it can be displayed in an area of ​​the playback interface, etc. The specific method can be determined according to the actual application scenario, and this application does not impose any restrictions on it.

[0116] Furthermore, upon receiving a trigger operation for scene exploration of the target video frame, if the video is currently playing, the terminal device can automatically pause playback and display the scene loading progress. Then, after displaying the 3D exploration space, if the user performs a close operation on the displayed 3D exploration space, the terminal device can close the 3D exploration space and automatically resume playback of the video, that is, resume playback from the point where the video was paused when the trigger operation for scene exploration of the target video frame was received. Alternatively, in another implementation, the terminal device may not automatically resume playback of the video, but may resume playback only after the user performs a playback operation.

[0117] Optionally, corresponding to the above-mentioned scene loading progress display method, the above-mentioned three-dimensional exploration space can also be displayed in a floating layer (such as a floating window) independent of the video playback interface, or the three-dimensional exploration space can also be displayed in another terminal interface outside the playback interface, or the scene loading progress can also be displayed in an area of ​​the playback interface, etc. The specific method can be determined according to the actual application scenario, and this application does not impose any restrictions on it.

[0118] Optionally, the initial display perspective of the 3D exploration space (i.e., the perspective of the initially displayed 3D scene) can be the perspective of the 3D image corresponding to the target video image, i.e., the initial image of the 3D exploration space. It can also be a corresponding 3D image generated from the target video image, which may contain 3D objects generated from the 2D objects in the target video image. For example, if the target video image contains a 2D chair, then the corresponding 3D image can contain a 3D chair of the same style generated from that 2D chair.

[0119] Step S103: Based on the exploration operations in the three-dimensional exploration space, display the three-dimensional scene obtained from the exploration.

[0120] Optionally, the system supports users performing exploration operations within a 3D exploration space. The terminal device can then display the corresponding 3D scene images obtained from these operations. These exploration operations can be any activity performed within the 3D exploration space, such as exploring different perspectives or constructing 3D objects within it. The following examples illustrate some methods for exploring the 3D exploration space; however, it should be noted that any method for exploring the 3D exploration space can be designed based on the specific application scenario.

[0121] For example, exploration operations in a three-dimensional exploration space may include perspective switching operations. Therefore, the terminal device can display a three-dimensional scene image after the perspective of the three-dimensional exploration space has been switched, based on this perspective switching operation in the three-dimensional exploration space. This three-dimensional scene image after the perspective switching is the three-dimensional scene image obtained through the perspective switching operation.

[0122] Optionally, the perspective switching operation can be triggered in any of the following ways: based on the terminal interface, based on a virtual reality device (VR device) connected to the terminal device, based on a keyboard connected to the terminal device, based on a change in the position of the virtual image in the three-dimensional exploration space, or based on a mouse connected to the terminal device. It should be noted that the perspective switching operation can be triggered in any feasible way; this is merely an example description.

[0123] Regarding the aforementioned method of triggering based on the terminal interface: In one feasible implementation, users can slide any distance in any direction on the terminal interface (such as an interface displaying a 3D exploration space) to switch the display perspective of the 3D exploration space. This sliding operation can be a perspective switching operation. Depending on the direction and distance of the user's slide, the perspective switched to the 3D exploration space can also be different.

[0124] For example, regarding the 3D scene currently displayed in the 3D exploration space, if a user wants to view more of the scene on the left, the user can swipe a certain distance to the right on the terminal interface to display the leftmost part of the current 3D scene (i.e., the content of the leftmost part of the scene); while if a user wants to view more of the scene above, the user can swipe a certain distance down on the terminal interface to display the top part of the current 3D scene (i.e., the content of the top part of the scene).

[0125] In another feasible implementation, the terminal interface (such as an interface displaying a 3D exploration space) may also display directional keys (which can be a direction adjustment control), allowing the viewpoint of the 3D exploration space to be adjusted in any direction. The operation of switching the viewpoint using these directional keys (such as pressing and holding the key to tilt in a certain direction) constitutes the aforementioned viewpoint switching operation.

[0126] Similarly, for the 3D scene currently displayed in the 3D exploration space, if the user wants to view more of the scene on the right, the user can press and hold the directional key to move to the left to display the rightmost part of the current 3D scene (i.e., the content of the rightmost part of the scene); and if the user wants to view more of the scene at the bottom, the user can press and hold the directional key to move up to display the bottom part of the current 3D scene (i.e., the content of the bottom part of the scene).

[0127] Please see Figure 6 , Figure 6 This is a schematic diagram of the interface of a three-dimensional scene in a three-dimensional exploration screen provided in an embodiment of this application. For example... Figure 6 The display interface for the three-dimensional exploration space can include a directional key for adjusting the viewing angle, allowing users to freely adjust the viewing angle of the three-dimensional exploration space using this directional key.

[0128] The display interface may also include an "Invite Friends" control and a "Stop Exploring" control. Users can click the "Invite Friends" control to invite more friends (such as the second object below) to explore the current 3D exploration space together. Users can click the "Stop Exploring" control to exit the current scene exploration, close the 3D exploration space, and return to the video playback interface to continue playing and watching the video.

[0129] Regarding the above-mentioned triggering method based on a mouse connected to a terminal device: the mouse can be connected to the terminal device via wired or wireless means, and the perspective of the three-dimensional exploration space can be adjusted in any direction using the mouse. The operation of switching perspectives using the mouse (such as holding down the mouse button and moving it in a certain direction) can be considered as the above-mentioned perspective switching operation.

[0130] Similarly, for the 3D scene currently displayed in the 3D exploration space, if the user wants to view more of the scene on the right, the user can hold down the mouse button and move it to the left to display the rightmost part of the current 3D scene (i.e., the content of the rightmost part of the scene); and if the user wants to view more of the scene at the bottom, the user can move the mouse button up to display the bottom part of the current 3D scene (i.e., the content of the bottom part of the scene).

[0131] Regarding the above-mentioned triggering method based on a virtual reality device connected to a terminal device: the virtual reality device can also be connected to the terminal device via wired or wireless means. The virtual reality device can be a head-mounted device, and users can wear the virtual reality device to see a three-dimensional exploration space in the virtual reality device. When wearing the virtual reality device, users can switch the display perspective of the three-dimensional exploration space at will by turning their heads.

[0132] For example, when a user wears the virtual reality device and tilts their head upwards, they can see the upper part of the currently displayed 3D scene in the 3D exploration space (i.e., the content of the upper part of the scene). When the user tilts their head to the left, they can see the leftmost part of the currently displayed 3D scene in the 3D exploration space (i.e., the content of the leftmost part of the scene). In this case, the viewpoint of the 3D exploration space can also change (i.e., switch) in the direction the user's head turns. The user's head-turning operation while wearing the virtual reality device is the aforementioned viewpoint switching operation, which can be transmitted from the virtual reality device to the terminal device.

[0133] Regarding the aforementioned triggering method based on a keyboard connected to a terminal device: the keyboard can be connected to the terminal device via wired or wireless means. The keyboard can contain many keys. In this application, some keys on the keyboard can be defined for switching the display perspective of the 3D exploration space. For example, the keyboard can define four directional keys: up, down, left, and right (or more directional keys, such as upper left, lower left, upper right, and lower right), to adjust the display perspective of the 3D exploration space upwards, downwards, leftwards, and rightwards.

[0134] For example, pressing the "Up" key will adjust the viewpoint of the 3D exploration space upwards (the degree of viewpoint adjustment can be set by the developers), allowing you to view the upper part of the currently displayed 3D scene; and pressing the "Down" key will adjust the viewpoint of the 3D exploration space downwards (the degree of viewpoint adjustment can be set by the developers), allowing you to view the lower part of the currently displayed 3D scene. The screen is divided into two parts: pressing the "left" key will adjust the viewpoint of the 3D exploration space further to the left (the degree of viewpoint adjustment by pressing the "left" key can be set by the developers), thus allowing you to view the leftmost part of the currently displayed 3D scene; pressing the "right" key will adjust the viewpoint of the 3D exploration space further to the right (the degree of viewpoint adjustment by pressing the "right" key can be set by the developers), thus allowing you to view the rightmost part of the currently displayed 3D scene.

[0135] Regarding the aforementioned method of triggering changes in the position of a virtual avatar within a 3D exploration space: This application allows setting up a virtual avatar for scene exploration within the created 3D exploration space. This virtual avatar can be a virtual cartoon character. Initially, the virtual avatar can be placed at any position on the initial screen of the 3D exploration space, and the user can adjust its position (e.g., by sliding the terminal interface, pressing the directional keys, or other methods). The effect is that the user can control the virtual avatar, moving it in any direction within the 3D exploration space (e.g., walking, running, jumping, sliding, crawling, etc.), thus changing the virtual avatar's position within the 3D exploration space.

[0136] Therefore, this application can adjust the display perspective of the three-dimensional exploration space by changing the position of the virtual image in the three-dimensional exploration space. The user's operation of moving the virtual image in the three-dimensional exploration space can be used to switch the perspective.

[0137] For example, if the user controls the virtual avatar to move to the left in the 3D exploration space, the perspective of the 3D exploration space can also change to the left, thereby displaying the leftmost part of the 3D scene currently displayed in the 3D exploration space; if the user controls the virtual avatar to move downwards in the 3D exploration space, the perspective of the 3D exploration space can also change downwards, thereby displaying the lower part of the 3D scene currently displayed in the 3D exploration space.

[0138] Among them, the aforementioned directions such as up, down, left, and right can be the directions in which users view the three-dimensional exploration space on the terminal interface. Specifically, in the three-dimensional exploration space, it can be any direction that conforms to the perspective of the virtual image in the three-dimensional exploration space can change.

[0139] The aforementioned three-dimensional exploration space may contain three-dimensional objects generated from the scene where the target video is located. These three-dimensional objects may include all objects appearing in the scene where the target video is located, such as furniture, walls, ground, plants, pedestrians, ornaments, and all other visible objects, whether three-dimensional or non-three-dimensional.

[0140] Optionally, the three-dimensional exploration space is not a complete control generated at the beginning. The three-dimensional exploration space can be a space that continuously expands and changes in range as the user explores.

[0141] For example, exploration operations in a 3D exploration space can also include controlling a virtual avatar to enter a 3D object within the space, allowing the user to explore the internal structure of the 3D object. In this case, the resulting 3D scene can be the view of the virtual avatar within that 3D object. For instance, a 3D object in the exploration space could include a 3D house or a box, allowing the user to control a virtual avatar to enter that 3D house or box.

[0142] Specifically, the terminal device can respond to the operation of controlling a virtual avatar to enter a target 3D object in the 3D exploration space, and display the 3D scene after the virtual avatar enters the target 3D object. The target 3D object can be any 3D object that can be entered in the 3D exploration space.

[0143] Optionally, if the target 3D object is larger than the virtual image, the size of the virtual image can be maintained, and the virtual image can be manipulated to enter the target 3D object; if the target 3D object is smaller than the virtual image, the virtual image can be shrunk (the extent to which it is shrunk can be set according to the actual application), so that the shrunk virtual image can enter the target 3D object.

[0144] For example, exploration operations in a 3D exploration space may also include triggering operations on 3D objects within the 3D exploration space. These triggering operations can include any operation that can be performed on the 3D objects. Optionally, triggering operations on 3D objects in the 3D exploration space may include at least one of the following: a move operation (i.e., moving the triggered 3D object to a new position within the 3D exploration space), a delete operation (e.g., deleting the triggered 3D object from the 3D exploration space), an edit operation (i.e., editing the triggered 3D object, including editing its shape, size, color, style, etc.), a cutout operation (e.g., cutting out the triggered 3D object from the 3D exploration space; the cutout 3D object can be saved as a different object), and so on.

[0145] Therefore, the terminal device can display a 3D scene image after the triggered 3D object has been processed (such as moved, deleted, edited, or cut out) according to the triggering method (such as moving, deleting, editing, or cutting out) based on the triggering operation on the 3D object in the 3D exploration space. The processed 3D scene image can be the 3D scene image obtained from the exploration.

[0146] For example, exploration operations in a three-dimensional exploration space may also include operations of adding data to the three-dimensional exploration space. The added data can be any data that the user wants to add. The data can be data from an optional list provided by the terminal device, or the data can be data uploaded or imported by the user. This application does not limit this.

[0147] Therefore, based on the addition of target data in the 3D exploration space, the terminal device can display a 3D scene image at the corresponding location in the 3D exploration space (which can be a user-specified location; if not specified, it can be any location in the currently displayed 3D scene image). This 3D scene image after adding the target data can be the 3D scene image obtained through the aforementioned exploration. The target data can be advertising data, object data (such as added 3D or 2D objects), link data, or multimedia data (such as video data or image data), etc.

[0148] Furthermore, the aforementioned terminal device can be the terminal device of the first object (which can be any user). Therefore, the scene exploration of the target video screen can be triggered by the first object, that is, the first object triggers the scene exploration of the target video screen, thereby displaying the three-dimensional exploration space of the scene where the target video screen is located.

[0149] This application also supports a first object triggering a sharing operation on the 3D exploration space. The terminal device can respond to the sharing operation on the 3D exploration space, obtain the sharing information of the 3D exploration space, and display the sharing information to the first object. This sharing information could be a link to the 3D exploration space, a QR code, etc. Thus, the first object can use this sharing information to invite a second object to collaboratively explore the 3D exploration space, achieving multi-person online interaction. The second object can be one or more other objects in the video client besides the first object.

[0150] In other words, this application supports the first object that triggers the creation of the three-dimensional exploration space, and shares the three-dimensional exploration space with other more objects (such as more users) to explore together. The three-dimensional exploration space can be continuously expanded in range as each object explores, and each object can also see each other's newly added and explored spaces.

[0151] Furthermore, the explored 3D exploration space can be saved. When a save operation is initiated, the terminal device can export the 3D exploration space (e.g., export a file in the appropriate format) and save the exported file. Later, when the user wants to continue exploring and viewing the 3D exploration space, they can open the exported 3D exploration space using a compatible tool or client to display the 3D scene and continue exploring. Alternatively, the exported 3D exploration space can also be uploaded and published on other platforms as a user-created work for other users to view. When publishing, the source of the creation of the 3D exploration space can be indicated, such as originating from the aforementioned video.

[0152] Furthermore, some attributes in the 3D exploration space can also be set as time-related attributes, such as the lighting direction, lighting color, and / or lighting intensity in the 3D exploration space. These attributes can change accordingly depending on the length of time the 3D exploration space has been open (i.e., the time it has been explored). The specific settings can be configured according to the actual application scenario.

[0153] This application provides a video processing method based on AIGC (Artificial Intelligence Generated Content), enabling users to more actively participate in the video content experience, freely explore and observe the video scenes from multiple angles. This not only enhances the user's sense of participation and immersion in the video, allowing interaction between the user and the video, but also provides video content creators with more diverse creative means and display platforms. For example, video content creators can create rich 3D exploration spaces for each video scene. Furthermore, the introduction of social sharing functions, such as sharing the 3D exploration space with a second person for joint exploration, further enriches the user's social experience and promotes the dissemination and exchange of video content.

[0154] In this application, the terminal device can play a video containing multiple video frames. In response to a trigger operation for scene exploration of a target video frame within the video, the terminal device can display a three-dimensional exploration space of the scene where the target video frame is located. Thus, the terminal device can display the corresponding three-dimensional scene obtained through the user's exploration operations within this three-dimensional exploration space. Therefore, the method proposed in this application supports users in exploring the scene of any video frame (such as a target video frame) within the played video in three-dimensional space (e.g., exploring within a three-dimensional exploration space) during video playback. This enhances the richness of the interaction between the user and the video, increases the user's participation in the video playback, and increases user engagement while watching the video.

[0155] Please see Figure 7 , Figure 7 This is a flowchart illustrating another video processing method provided in an embodiment of this application. The executing entity in this embodiment can be a server, such as the backend server of the video client playing the video described above. Figure 7 As shown, the method may include:

[0156] Step S201: Obtain the target video frame sent by the terminal device. The target video frame is the video frame of the scene exploration triggered in the video played by the terminal device.

[0157] Optionally, the server can obtain the target video frame sent by the terminal device. The target video frame is the video frame in the video played by the terminal device that has been triggered for scene exploration. That is, the target video frame is the target video frame in step S102 where the scene exploration trigger operation was performed. The terminal device can send the target video frame that has been triggered for scene exploration to the server.

[0158] Step S202: Obtain scene information associated with the scene where the target video frame is located.

[0159] Optionally, the server can obtain scene information associated with the scene where the received target video frame is located. This scene information can be used to assist in the subsequent generation of a 3D exploration space for the scene where the target video frame is located.

[0160] The scene information can include any information that can be used to assist in generating the 3D exploration space, such as at least one of the following: video (i.e., the above). Figure 3 Other video frames in the same scene as the target video frame (in the video played in the corresponding embodiment), video frames adjacent to the target video frame, and text description information related to the scene where the target video frame is located.

[0161] The method of acquiring other video frames that are in the same scene as the target video frame can include: the server can call an object detection model to perform object detection on the target video frame to detect objects appearing (i.e., contained) in the target video frame; and the server can also perform object detection on each of the other video frames in the video besides the target video frame to detect objects appearing in each of the other video frames. Thus, the server can consider video frames among the other video frames that contain the same objects (at least one of the same objects) as the target video frame as video frames in the same scene as the target video frame. The object detection model can be a trained model capable of detecting objects contained in an image.

[0162] Alternatively, the process of acquiring other video frames in the same scene as the target video frame may also include: the server can acquire the screen similarity (which may be image similarity) between the target video frame and each other video frame in the video, so that the server can regard video frames whose screen similarity with the target video frame is greater than or equal to a set feature similarity threshold as video frames in the same scene as the target video frame.

[0163] For example, a server can invoke an image feature encoding model to encode the target video frame and other video frames within the video, thereby obtaining the encoded features of the target video frame and the encoded features of the other video frames. These encoded features can be feature vectors. The server can calculate the feature similarity (e.g., cosine similarity) between the encoded features of the target video frame and the encoded features of the other video frames. Video frames whose encoded features have a feature similarity greater than or equal to a set feature similarity threshold are considered to be in the same scene as the target video frame. This image feature encoding model can be a pre-trained model used to encode the image's features.

[0164] Methods for obtaining video frames adjacent to the target video frame in a video can include: when the terminal device sends the target video frame to the server, it can also send the playback progress corresponding to the target video frame in the video to the server. For example, the playback progress can be the playback time of the target video frame in the video's playback progress bar (e.g., 50 minutes and 20 seconds). Therefore, the server can consider multiple video frames adjacent to the playback progress corresponding to the target video frame in the video, or video frames within a few seconds before or after, as video frames adjacent to the target video frame. For example, it can consider 5 video frames before and after the playback progress corresponding to the target video frame in the video (i.e., 5 frames before and after) as video frames adjacent to the target video frame. Alternatively, it can consider video frames within 5 seconds before and after the playback progress corresponding to the target video frame in the video as video frames adjacent to the target video frame.

[0165] The method of obtaining text description information related to the scene where the target video is located may include: if the video played in this application is a video with a script, the server can obtain the script of the video. For example, for film and television dramas, there is usually a script, or for short videos, there may be a pre-edited script.

[0166] Specifically, when a terminal device sends a target video frame to a server, it can also send the video identifier (such as ID) of the video to which the target video frame belongs. If the video has a script, the script can be stored in association with the video identifier of the video. Therefore, the server can find and obtain the script of the video to which the video identifier belongs through the video identifier sent by the terminal device.

[0167] The server can invoke a large language model (such as GPT) to segment the acquired script. This model can understand and recognize logical identifiers such as titles, paragraphs, or chapters within the script, thus segmenting the script into multiple segments. This ensures that each segment is complete and logically clear; for example, text under different chapters or major headings can be divided into different segments. This large language model can be a pre-trained model capable of performing various text understanding operations.

[0168] Optionally, before segmenting the script, it can be preprocessed. This preprocessing may include text extraction, file format conversion, and data cleaning. The script can be in any file format such as TXT (plain text), PDF (portable document format), JSON (lightweight data interchange format), or XML (markup language). The server can first use a file extraction tool that matches the script's file format to extract the text content. Then, the extracted text content can be converted to a standard plain text format (such as TXT). Therefore, if the script itself is in TXT file format, the text extraction step can be omitted.

[0169] The server can perform data cleaning on the converted plain text content. This includes removing meaningless information (such as fixed formatting descriptions related to the original script's file format) and useless characters (such as whitespace or special characters). After cleaning, the cleaned text content is obtained. The server can then segment the cleaned text content, resulting in multiple segments of the script. These segments must be in plain text format and conform to the input data format of the Large Language Model (GPT).

[0170] Therefore, the server can also call a large language model to locate the segment corresponding to the target video frame in multiple segments of the script (i.e., the segment corresponding to the target video frame is the segment depicted by the content of the target video frame). For example, this process may include: the server can obtain the audio between the playback progress corresponding to the target video frame in the video, and can perform text recognition on the audio to obtain the text content between the playback progress corresponding to the target video frame (such as spoken content (e.g., lines) or narration). For example, if the playback progress corresponding to the target video frame is 50 minutes and 20 seconds of the video, the server can obtain the audio within 10 seconds before and after 50 minutes and 20 seconds (i.e., 50 minutes and 10 seconds to 50 minutes and 30 seconds, or other durations in specific implementations), and perform text recognition on the audio to obtain the text content between 50 minutes and 10 seconds and 50 minutes and 30 seconds of the video.

[0171] Therefore, the server can call the large language model to perform text understanding on the content of each section of the script and the identified text content, thereby locating the section in the script that describes the text content, and the located section can be the section corresponding to the target video frame.

[0172] Furthermore, segments other than the segment corresponding to the target video frame can be referred to as candidate segments, and the segment corresponding to the target video frame can be referred to as the target segment. The server can also obtain the similarity between each candidate segment and the target segment (e.g., the text similarity between each candidate segment and the target segment), and can use candidate segments whose similarity with the target segment is greater than or equal to a set text similarity threshold as matching segments of the target segment. These matching segments can be considered as segments describing the same scene as the target segment. Any feasible and suitable method can be used to obtain the similarity between candidate segments and target segments. For example, candidate segments can be encoded into corresponding text features (which can be feature vectors), and target segments can also be encoded into corresponding text features (which can be feature vectors). Thus, the feature similarity (e.g., cosine similarity) between the text features of the candidate segment and the text features of the target segment can be obtained as the similarity between the candidate segment and the target segment.

[0173] Optionally, the target segment and the matching segment mentioned above can be associated with the video identifier of the video. This allows the matching segment of the target segment to be quickly found by using the associated video identifier, target segment, and matching segment when other objects (such as other objects besides the first object) trigger scene exploration of the video that also corresponds to the target segment. This enables the rapid construction of the three-dimensional exploration space explored by other objects.

[0174] The server can input both the matching segment and the target segment into a large language model to understand the text content of the matching segment and the target segment, thereby generating a comprehensive summary and key information of the understanding of the matching segment and the target segment. This summary is a general information generated after understanding the text content of the matching segment and the target segment. The summary summarizes the overall text content described by the matching segment and the target segment. The key information may include some key information described in the matching segment and the target segment, such as some scene descriptions of the scene where the target video screen is located (such as environmental descriptions related to the scene, interior furnishings, etc.), the time of occurrence, the location of occurrence, the main characters (such as existing characters), and key events (such as events that occur in the current scene, such as interviews, meals, etc.).

[0175] The server can integrate (such as combine or splice together) the summary and the key information as the final text description information related to the scene where the target video is located.

[0176] Alternatively, for videos without a script, some text descriptions related to the video (such as text describing the content of the video, the title information of the video, etc.) can be used as text description information related to the scene where the target video is located.

[0177] The specific method for obtaining text description information related to the scene where the target video is located can be determined based on the actual application scenario; any suitable method is acceptable.

[0178] Step S203: Based on the target video image and scene information, generate the first display data of the three-dimensional exploration space of the scene where the target video image is located.

[0179] Optionally, the server can combine the aforementioned target video frame and scene information associated with the scene where the target video frame is located to generate first display data for the three-dimensional exploration space of the scene where the target video frame is located. This first display data can be used to initially display the three-dimensional exploration space, such as displaying the initial frame of the three-dimensional exploration space. The process of generating the first display data from the target video frame and scene information can be described as follows.

[0180] The server can acquire several multi-view images, which can be images of different perspectives of the scene where the target video is located. The multi-view images can include one or more of the following: the target video (which is required), other video frames in the same scene as the target video in the video described above, and video frames adjacent to the target video in the video described above.

[0181] The server can call an object tracking algorithm to track and detect all objects appearing in the multi-view image, thereby detecting the category of each object, its position and viewpoint in the multi-view image, and the object's own attributes. The object category indicates what the object is; for example, if an object is a bed, then the object category is bed. Regarding the object's position in multi-view images, this application can establish a coordinate system in 3D space. The origin of this coordinate system can be arbitrarily set. An object tracking algorithm can track and identify the object's position in different multi-view images within this coordinate system. This position can be represented as a coordinate location in the 3D space. As for the object's perspective in multi-view images, the object tracking algorithm can track and identify the object's perspective in the multi-view image within this coordinate system. This perspective can be represented as a directed line segment (i.e., a line segment with a direction) in the 3D space. This directed line segment indicates the object's shooting perspective (i.e., the direction of viewing the object, which can be understood as the camera's direction) in the multi-view image within this coordinate system. The object's attributes can include some of the object's own characteristics, such as its size, shape, and color.

[0182] For example, the object tracking algorithm mentioned above could be the SORT algorithm (Simple Online and Realtime Tracking, a target tracking algorithm) or the DeepSORT algorithm (an improved target tracking algorithm based on SORT). SORT is a simple and efficient online real-time tracking algorithm based on Kalman filtering and the Hungarian algorithm. Its advantages are simple implementation, fast computation speed, and suitability for real-time applications. However, its robustness to object occlusion and appearance changes is slightly weaker. DeepSORT, on the other hand, adds deep learning feature extraction and Mahalanobis distance calculation to SORT. By combining appearance features and motion information for data association, it significantly improves the robustness and accuracy of tracking, especially performing better when dealing with occlusion and appearance changes. However, it has higher computational complexity and resource requirements. Therefore, when server performance is sufficient, we can use the DeepSORT algorithm to pursue higher object tracking and detection performance. When server performance is less sufficient, we can use the SORT algorithm to pursue higher object tracking and detection efficiency and reduce the overhead of object tracking and detection.

[0183] Furthermore, the server can also obtain the object tag (such as a user tag) of the first object. This object tag can be used to indicate the preferences and attributes of the first object, such as the object tag of the first object being a "2D boy" or similar. This application can also combine the object tags of the objects to construct a 3D exploration space that matches the object's preferences. If the first object also invites a second object to explore the 3D exploration space together, the server can also obtain the object tag of the second object. Since the 3D exploration space expands incrementally with the user's exploration operations, after inviting the second object to explore the 3D exploration space together, the server can simultaneously use the object tags of the first object and the second object to comprehensively generate the newly added exploration screen in the 3D exploration space. That is, after inviting the second object, the server can use the object tags of the first object and the second object to comprehensively generate the space newly explored by the first and second objects, matching their combined preferences. This is described below.

[0184] The server can also obtain a 3D construction model, which can be a trained model that can be used to construct and generate 3D objects. For example, the 3D construction model can be a NeRF model (Neural Radiance Fields). The NeRF model uses a technique of 3D reconstruction using multi-view images. It can learn and model static 3D objects or scenes through implicit representation.

[0185] The server can organize and convert at least one of the following: the obtained multi-view images, the results of object tracking and detection (including object category, object position and viewpoint in the multi-view images, and object attributes), the text description information related to the scene where the target video frame is located, and the object label (i.e., user label) of the first object, into a data format (such as encoded as a corresponding vector or matrix) that conforms to the input format of the 3D construction model (such as a NeRF model), to obtain the input data of the 3D construction model. This input data may contain at least one of the following: information related to the multi-view images, information related to the results of object tracking and detection, information related to the text description information related to the scene where the target video frame is located, and information related to the object label of the first object. The multi-view images are necessary to obtain this input data. Here, whatever the input data of the 3D construction model is, the 3D construction model can also be pre-trained using sample data with the same concept (or type) as the input data.

[0186] The server can input the input data (including data obtained by format conversion of the target video footage and scene information) into the 3D model. The server can then call the 3D model to perform 3D modeling of objects (including objects appearing in multi-view images and / or objects derived from objects appearing in multi-view images) in the aforementioned 3D space (which may be combined with the coordinate system in the 3D space) using the input data. This generates 3D modeling information of the scene where the target video footage is located. This 3D modeling information may include the individual 3D meshes (a 3D spatial structure composed of vertices, edges, and faces used to describe the shape and appearance of the objects), texture information (such as the texture of the object's surface), and lighting information (such as the direction of light) of each object modeled in the scene where the target video footage is located. If the 3D model can perform feature learning on the input data (the specific learning method and effect can be obtained during training) to learn features for modeling 3D objects, the 3D model can then generate the 3D modeling information using the learned features.

[0187] The input data will definitely include data obtained by converting the target video image into a different format (this data is used to represent or refer to the target video image). The server can mark the data obtained by converting the target video image into a different format. This mark indicates that the 3D model can use this data to generate the initial image of the 3D exploration space. The 3D scene images obtained through subsequent exploration operations can all be incrementally derived and generated by the 3D model based on the input data and the generated initial image.

[0188] In addition, the server can also obtain the material information of each object modeled by the 3D construction model. This material information can be determined by the category identified by the object (or, for derived objects, by the category derived from the object), such as the material of a wooden board being wood; or it can be obtained by detecting the object during the object tracking and detection process mentioned above.

[0189] Optionally, the 3D modeling generated from the 3D construction model can be initial 3D modeling information. This initial 3D modeling information can include the 3D mesh information (which can be point cloud data or voxel data), texture information, and lighting information of each object modeled. The server can convert this initial 3D modeling information into a format compatible with the 3D engine's input format, thus obtaining 3D modeling information. This 3D modeling information can include: 3D mesh information converted from the initial 3D modeling information into a 3D mesh file compatible with the 3D engine's input format (such as OBJ or FBX), texture images (such as texture maps) extracted from the texture information in the initial 3D modeling information, and lighting information exported from the initial 3D modeling information. This 3D modeling information can then be used as input to the 3D engine. Furthermore, this 3D modeling information can also include the material information of each object modeled by the 3D construction model.

[0190] The server can generate the first display data of the 3D exploration space using the 3D modeling information obtained above. This process may include: the server obtaining rendering performance information of the terminal device, which can be used to indicate the performance of the terminal device in rendering 3D images. For example, the server may obtain this rendering performance information by sending an example 3D modeling information to the terminal device. This example 3D modeling information can be used to render multiple example 3D images. After receiving the example 3D modeling information, the terminal device can render 3D images using it. If the terminal device can smoothly render the multiple 3D images corresponding to the example 3D modeling information, such as rendering the multiple 3D images corresponding to the example 3D modeling information within a set time (e.g., 2 seconds), the terminal device can generate rendering performance information indicating sufficient performance and provide it to the server. Conversely, if the terminal device cannot smoothly render the multiple 3D images corresponding to the example 3D modeling information, such as not rendering the multiple 3D images corresponding to the example 3D modeling information within a set time (e.g., 1 second), the terminal device can generate rendering performance information indicating insufficient performance and provide it to the server. It should be noted that in practical application scenarios, other feasible methods can also be used to determine the rendering performance of the terminal device.

[0191] After obtaining the rendering performance information of the terminal device, if the rendering performance information indicates that the terminal device's rendering performance is sufficient (i.e., the terminal device's rendering performance is adequate or good), then the server can directly use the 3D modeling information as the primary display data. In other words, if the terminal device's rendering performance is sufficient, the terminal device can automatically render the 3D scene locally using its own 3D engine (such as Unity) and graphics card.

[0192] If the rendering performance information indicates insufficient rendering performance of the terminal device (i.e., poor rendering performance), the server can directly render the 3D scene of the 3D exploration space (such as the initial image of the 3D exploration space) in the cloud using the 3D modeling information, and use the rendered 3D scene as the first display data. In other words, when the terminal device's rendering performance is insufficient, the 3D scene can be remotely rendered in the cloud using the background 3D engine and graphics card (i.e., cloud rendering), without requiring the terminal device to render it itself. This improves the real-time performance of the 3D scene in the 3D exploration space displayed on the terminal device and enhances the user experience.

[0193] Optionally, to improve the efficiency of server data transmission to terminal devices, when the server renders the corresponding 3D scene image using 3D modeling information, it can also encode the rendered 3D scene image to obtain a corresponding data stream. For example, the server can use low-latency encoding technology to encode the rendered 3D scene image to obtain a corresponding data stream, which can then be used as the first display data. This encoding technology could be H.264 (a video encoding technology) or H.265 (a video encoding technology), etc.

[0194] Step S204: Send the first display data to the terminal device, so that the terminal device can display the three-dimensional exploration space based on the first display data.

[0195] Optionally, the server can send the first display data obtained above to the terminal device, and the terminal device can display the three-dimensional exploration space through the first display data, such as displaying the initial screen of the three-dimensional exploration space (i.e., the initial three-dimensional scene screen).

[0196] Specifically, if the first displayed data is 3D modeling information, the terminal device can use the received 3D modeling information to render a 3D scene of the 3D exploration space. For example, if the terminal device includes a 3D engine and a graphics card, it can input the 3D modeling information into its 3D engine, which will then render and generate image rendering data. Finally, the image rendering data generated by the 3D engine is sent to the graphics card, which will perform the final image rendering to generate the final 3D scene.

[0197] Terminal devices can display the three-dimensional exploration space by rendering the three-dimensional scene themselves. For example, the terminal device can display the three-dimensional scene it renders on the terminal interface, thus realizing the display of the three-dimensional exploration space.

[0198] If the first displayed data is a 3D scene rendered by the server, the terminal device can display the 3D exploration space using this 3D scene sent by the server. For example, the terminal device can display the 3D scene sent by the server on its interface, thus realizing the display of the 3D exploration space. The principle by which the server renders the 3D scene using 3D modeling information in the cloud (e.g., through a cloud server) is the same as the principle by which the terminal device renders the 3D scene using 3D modeling information.

[0199] Furthermore, if the first display data is a data stream obtained by the server encoding the rendered 3D scene image, the terminal device can decode the data stream to obtain the 3D scene image rendered by the server, and can display the 3D exploration space through the 3D scene image. For example, the terminal device can display the decoded 3D scene image on the terminal interface, thus realizing the display of the 3D exploration space.

[0200] Furthermore, when rendering 3D scenes, the 3D engine can set scene elements such as light sources and cameras in the 3D space used for scene rendering. The initial viewpoint of the camera (such as the viewpoint of the initial image of the 3D exploration space) and the initial attributes of the light source (such as initial direction, light intensity, light color, etc.) can be set by the user. This camera is a virtual camera set in the 3D exploration space (which can be the aforementioned 3D space) used to construct the scene where the target video image is located. It is used to control the viewpoint of the 3D exploration space (i.e., the viewing angle of the 3D exploration space, or the display angle of the 3D exploration space).

[0201] Subsequently, based on the perspective switching operation in the 3D exploration space displayed on the user terminal device, the perspective of the 3D exploration space can be adjusted accordingly using the camera in that 3D space. For example, after the user performs a perspective switching operation, the terminal device can generate a perspective switching command. This perspective switching command can include the direction of the user's perspective switching (such as the direction of sliding on the terminal interface) and the range of the switching (such as the distance of sliding on the terminal interface). The direction of the switching and the range of the switching indicate the range of change of the camera's shooting direction (i.e., the camera's viewing angle) (such as the range of viewpoint movement or rotation). The terminal device can send this perspective switching command to the server, and the server can then use this perspective switching command to generate display data for the newly added 3D scene within the switched perspective (the concept is similar to the first display data mentioned above).

[0202] Furthermore, the server can receive exploration instructions sent by the terminal device. These instructions are generated by the terminal device based on the exploration operations performed by the user in the displayed 3D exploration space. The exploration instructions correspond to the exploration operations performed by the terminal device in the 3D exploration space, and the corresponding exploration instructions can carry information related to those exploration operations.

[0203] If the exploration operation is a perspective switching operation, then the exploration command can be the aforementioned perspective switching command, which can carry the direction and range of the perspective switching; if the exploration operation is the aforementioned triggering operation on a 3D object in the 3D exploration space, then the exploration command can be the corresponding triggering command on the 3D object in the 3D exploration space (which can carry the triggering method, such as moving, deleting, editing, or peek-through); if the exploration operation is the aforementioned operation of adding data in the 3D exploration space, then the exploration command can be the corresponding command of adding data in the 3D exploration space, which can carry the data to be added (such as the aforementioned target data); and so on.

[0204] The server can generate second display data of the three-dimensional scene explored in the three-dimensional exploration space by receiving exploration instructions. The generation principle of the second display data is the same as that of the first display data. However, the first display data can be used to display the initial screen of the three-dimensional exploration space, while the second display data can be used to display the newly explored screen of the three-dimensional exploration space.

[0205] It should be noted that if the first object has already invited the second object to collaboratively explore the 3D exploration space before the server receives the exploration command, then when generating the aforementioned second display data, compared to generating the first display data, the server can incorporate more object tags from the second object to generate the second display data. This allows for the construction of more 3D scene images that conform to the combined preferences of the first and second objects within the 3D exploration space. The method of incorporating the object tags from the second object is the same as the method described for processing the object tags from the first object.

[0206] The server can send the generated second display data to the terminal device, so that the terminal can use the second display data to display the corresponding three-dimensional scene image obtained in the three-dimensional exploration space.

[0207] Developers can write scripts to change the camera's viewpoint in the 3D exploration space through viewpoint switching operations on the terminal device, and enable interaction between the user and the triggered 3D object through user triggering operations on the terminal device.

[0208] Optionally, during the rendering process of the 3D engine (such as a 3D engine in a terminal device or a 3D engine in the cloud) based on the input data, some rendering techniques can be added to optimize the rendering and improve the rendering effect or efficiency. Such rendering techniques may include at least one of the following: LOD (Level of Detail, a rendering optimization technique), light mapping (a texture mapping technique), occlusion culling (a technique used to optimize 3D graphics rendering performance), etc.

[0209] The Level of Detail (LOD) technique allocates rendering resources based on the location and importance of an object's nodes within the display environment, reducing the face count and detail of less important objects to achieve high-efficiency rendering computation. By employing lightmaps, the lighting effects on object surfaces can be dynamically calculated, resulting in more realistic lighting simulation. Occlusion culling improves rendering efficiency by reducing the rendering of invisible objects, thus enhancing rendering performance.

[0210] Please see Figure 8 , Figure 8 This is a scene illustration provided in an embodiment of this application for acquiring a three-dimensional scene image. For example... Figure 8As shown, the server can obtain the target video frame and the scene information associated with the scene where the target video frame is located. The server can obtain input data that matches the 3D construction model through the target video frame and scene information. The server can input the input data into the 3D construction model to call the 3D construction model to perform 3D modeling based on the input data and generate 3D modeling information.

[0211] The 3D modeling information can then be input into a 3D engine (such as a 3D engine on a terminal device or a 3D engine in the cloud) for rendering, obtaining rendering data. This rendering data can then be sent to a graphics card (such as a graphics card on a terminal device or a graphics card in the cloud) for final rendering processing, generating the final 3D scene to be displayed. For the sake of visual continuity, the generated 3D scene can be a series of multiple frames. For example, when the user switches their perspective on the 3D exploration space, many new frames of the 3D scene can be added within the switched perspective.

[0212] The method provided in this application can efficiently transform video content into an interactive viewing experience that users can freely explore (such as free exploration in a 3D exploration space within the scene depicted in the video), significantly enhancing the viewing experience. Furthermore, this solution can seamlessly integrate video content and games; for example, the 3D exploration space can be used as a game space where users can perform any feasible exploration operation (which could be a game operation), breaking down the boundaries between video and games. This efficiently creates an open game world through video content, thereby significantly increasing the value of video products.

[0213] Please see Figure 9 , Figure 9 This is a schematic diagram illustrating a process for generating a 3D scene image according to an embodiment of this application. Figure 9 As shown, the process may include:

[0214] In step S301, the server can acquire images from multiple angles.

[0215] In step S302, the server can use an object tracking algorithm to track and detect objects in multi-angle images (i.e., track and recognize them) to obtain the tracking and detection results of objects in multi-angle images.

[0216] In step S303, the server can combine the tracking and detection results of objects in multi-angle images with the three-dimensional construction model to perform 3D construction of the objects and obtain the above-mentioned three-dimensional modeling information.

[0217] In step S304, the server can start the process of real-time 3D rendering of the obtained 3D modeling information. If the rendering performance of the terminal device is sufficient, the following step S305 can be executed to render the image. If the rendering performance of the terminal device is insufficient, the following step S306 can be executed to render the image.

[0218] In step S305, the server can send the above-mentioned 3D modeling information to the terminal device so that the terminal device can perform 3D rendering on the 3D modeling information to obtain the corresponding 3D scene image.

[0219] In step S306, the server can perform 3D rendering of the 3D modeling information in the cloud to obtain the corresponding 3D scene image.

[0220] By using the method provided in this application, the rendering method can be adaptively selected according to the rendering performance of the terminal device. For example, if the rendering performance of the terminal device is excellent, the terminal device can perform local rendering, and if the rendering performance of the terminal device is poor, cloud rendering can be performed. This can improve the efficiency of the terminal device in displaying the three-dimensional scene of the three-dimensional exploration space, achieve low latency, and improve the user experience.

[0221] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this application. This video processing apparatus 1 can be applied to terminal devices. Figure 10 As shown, the video processing device 1 may include a playback module 11, a display module 12, and an exploration module 13.

[0222] Playback module 11 is used to play videos, which contain multiple video frames;

[0223] Display module 12 is used to respond to the trigger operation of scene exploration of the target video frame in the video and display the three-dimensional exploration space of the scene where the target video frame is located.

[0224] The exploration module 13 is used to display the three-dimensional scene obtained from the exploration based on the exploration operations in the three-dimensional exploration space.

[0225] Optionally, exploration actions include perspective switching.

[0226] The exploration module 13 displays the resulting 3D scene based on exploration operations in the 3D exploration space in the following ways:

[0227] Based on the perspective switching operation in the 3D exploration space, the 3D scene image after the perspective switching in the 3D exploration space is displayed.

[0228] Optionally, the trigger method for the view switching operation can be any of the following:

[0229] Triggering methods based on terminal interface, triggering methods based on virtual reality device connected to terminal device, triggering methods based on keyboard connected to terminal device, triggering methods based on mouse connected to terminal device, and triggering methods based on positional changes of virtual image in three-dimensional exploration space.

[0230] Among them, virtual avatars are images set up in the three-dimensional exploration space for scene exploration.

[0231] Optionally, the three-dimensional exploration space includes a virtual avatar set up for scene exploration, and the three-dimensional exploration space includes three-dimensional objects generated for the scene where the target video screen is located. The exploration operation includes the operation of controlling the virtual avatar to enter the three-dimensional objects in the three-dimensional exploration space.

[0232] The exploration module 13 displays the resulting 3D scene based on exploration operations in the 3D exploration space in the following ways:

[0233] In response to the operation of controlling the virtual avatar to enter the target 3D object in the 3D exploration space, the 3D scene screen after the virtual avatar enters the target 3D object is displayed.

[0234] Optionally, the 3D exploration space includes 3D objects generated from the scene where the target video image is located, and the exploration operation includes triggering operations on the 3D objects in the 3D exploration space.

[0235] The exploration module 13 displays the resulting 3D scene based on exploration operations in the 3D exploration space in the following ways:

[0236] Based on the triggering operation on the three-dimensional object in the three-dimensional exploration space, the three-dimensional scene screen is displayed after the three-dimensional object has been processed according to the triggering method;

[0237] Among them, the triggering operations on three-dimensional objects in the three-dimensional exploration space include at least one of the following: movement operation, deletion operation, editing operation, and cutout operation.

[0238] Optionally, exploration operations include adding data to the three-dimensional exploration space;

[0239] The exploration module 13 displays the resulting 3D scene based on exploration operations in the 3D exploration space in the following ways:

[0240] Based on the addition of target data in the 3D exploration space, the 3D scene image after the target data has been added is displayed at the corresponding location in the 3D exploration space.

[0241] Optionally, if the video is playing when scene exploration is triggered, the target video frame is the video frame that the video is playing at the time scene exploration is triggered; and,

[0242] If the video is paused when scene exploration is triggered, the target video frame is the frame where the video was paused when scene exploration was triggered.

[0243] Optionally, in response to a trigger operation for scene exploration of a target video frame in the video, the display module 12 displays the three-dimensional exploration space of the scene where the target video frame is located, including:

[0244] In response to a scene exploration trigger, display the scene loading progress;

[0245] When the scene loading progress indicator shows that loading is complete, the 3D exploration space is displayed.

[0246] Optionally, the playback module 11 described above is used for:

[0247] When a scene exploration of the target video frame is triggered, pause video playback;

[0248] After displaying the three-dimensional exploration space, the playback module 11 is also used for:

[0249] In response to the action of closing the 3D exploration space, the 3D exploration space is closed, and the video continues to play.

[0250] Optionally, the scene exploration of the target video frame is triggered by the first object, and the video processing device 1 further includes a sharing module 14, which is used for:

[0251] Based on the sharing operation of the three-dimensional exploration space, obtain the sharing information of the three-dimensional exploration space; use the shared information to invite a second object to conduct collaborative exploration of the three-dimensional exploration space.

[0252] Optionally, the video processing device 1 further includes a storage module 15, which is used for:

[0253] When a save operation for the 3D exploration space is obtained, the 3D exploration space is exported and saved.

[0254] According to one embodiment of this application, Figure 3 The steps involved in the video processing method shown can be derived from... Figure 10 The video processing apparatus 1 shown is executed by each module. For example, Figure 3 Step S101 shown can be performed by Figure 10 The playback module 11 in the middle is used to execute it. Figure 3Step S102 shown can be performed by Figure 10 The display module 12 in the middle is used to execute; Figure 3 Step S103 shown can be performed by Figure 10 The exploration module 13 in the middle is used to execute.

[0255] Please see Figure 11 , Figure 11 This is a schematic diagram of another video processing apparatus provided in an embodiment of this application. This video processing apparatus 2 can be applied to a server. Figure 11 As shown, the video processing device 2 may include: a first acquisition module 21, a second acquisition module 22, a generation module 23, and a sending module 24.

[0256] The first acquisition module 21 is used to acquire the target video frame sent by the terminal device. The target video frame is the video frame in the video played by the terminal device that is triggered for scene exploration.

[0257] The second acquisition module 22 is used to acquire scene information associated with the scene where the target video frame is located;

[0258] The generation module 23 is used to generate the first display data of the three-dimensional exploration space of the scene where the target video image is located, based on the target video image and scene information;

[0259] The sending module 24 is used to send the first display data to the terminal device, so that the terminal device can display the three-dimensional exploration space based on the first display data.

[0260] Optionally, scene information associated with the scene where the target video frame is located includes at least one of the following:

[0261] Other video frames in the same scene as the target video frame, video frames adjacent to the target video frame, and text description information related to the scene where the target video frame is located.

[0262] Optionally, the generation module 23 generates the first display data of the three-dimensional exploration space of the scene where the target video frame is located based on the target video frame and scene information, including:

[0263] Obtain a 3D model;

[0264] The 3D model is invoked to perform 3D modeling based on the target video frame and scene information, generating 3D modeling information of the scene where the target video frame is located;

[0265] Based on the 3D modeling information, the first display data of the 3D exploration space is generated.

[0266] Optionally, the generation module 23 generates the first display data of the 3D exploration space based on the 3D modeling information in the following ways:

[0267] Obtain rendering performance information from the terminal device;

[0268] If the rendering performance information indicates that the rendering performance of the terminal device is sufficient, then the 3D modeling information will be used as the first display data.

[0269] If the rendering performance information indicates that the rendering performance of the terminal device is insufficient, then a 3D scene of the 3D exploration space is rendered based on the 3D modeling information, and the rendered 3D scene is used as the first display data.

[0270] Optionally, if the first display data is 3D modeling information, the terminal device is used to render a 3D scene of the 3D exploration space using the received 3D modeling information, and to display the 3D exploration space based on the rendered 3D scene.

[0271] If the first display data is a 3D scene image rendered by the server, the terminal device is used to display a 3D exploration space based on the 3D scene image sent by the server.

[0272] Optionally, the video processing device 2 described above is also used for:

[0273] Receive exploration instructions sent by the terminal device, which are generated by the terminal device based on exploration operations in the displayed three-dimensional exploration space;

[0274] Based on the exploration command, generate second display data of the three-dimensional scene explored in the three-dimensional exploration space;

[0275] The second display data is sent to the terminal device, which then displays the three-dimensional scene obtained in the three-dimensional exploration space based on the second display data.

[0276] According to one embodiment of this application, Figure 7 The steps involved in the video processing method shown can be derived from... Figure 11 The video processing apparatus 2 shown is used to perform the operation. For example, Figure 7 Step S201 shown can be performed by Figure 11 The first acquisition module 21 in the process is executed. Figure 7 Step S202 shown can be performed by Figure 11 The second acquisition module 22 in the middle is used to execute; Figure 7 Step S203 shown can be performed by Figure 11 The generation module 23 in the middle is used to execute, Figure 7 Step S204 shown can be performed by Figure 11 The sending module 24 in the middle is used to execute this.

[0277] In this application, the terminal device can play a video containing multiple video frames. In response to a trigger operation for scene exploration of a target video frame within the video, the terminal device can display a three-dimensional exploration space of the scene where the target video frame is located. Thus, the terminal device can display the corresponding three-dimensional scene obtained through the user's exploration operations within this three-dimensional exploration space. Therefore, the method proposed in this application supports users in exploring the scene of any video frame (such as a target video frame) within the played video in three-dimensional space (e.g., exploring within a three-dimensional exploration space) during video playback. This enhances the richness of the interaction between the user and the video, increases the user's participation in the video playback, and increases user engagement while watching the video.

[0278] According to one embodiment of this application, Figure 10 The video processing device 1 shown and Figure 11 The modules in the video processing device 2 shown can be individually or entirely combined into one or more units, or some of these units can be further divided into multiple functionally smaller sub-units to achieve the same operation without affecting the technical effects of the embodiments of this application. The above modules are based on logical functional division. In practical applications, the function of one module can be implemented by multiple units, or the function of multiple modules can be implemented by one unit. In other embodiments of this application, video processing device 1 and video processing device 2 may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0279] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0280] According to one embodiment of this application, a computer program capable of executing the steps involved in the corresponding methods shown in the various embodiments of this application can be run on a general-purpose computer device (which may include processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM)) to construct, as described in the embodiments of this application. Figure 10 The video processing device 1 shown herein and Figure 11The video processing apparatus 2 shown above. The computer program described above can be recorded on a computer-readable recording medium, and can be loaded into the computer device described above through the computer-readable recording medium and run therein.

[0281] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 12 As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. In some embodiments, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 12 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0282] exist Figure 12 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0283] Play a video containing multiple video frames;

[0284] In response to a triggered operation of scene exploration of the target video frame in the video, a three-dimensional exploration space of the scene where the target video frame is located is displayed;

[0285] Based on the exploration operations in the three-dimensional exploration space, the obtained three-dimensional scene image is displayed.

[0286] Alternatively, processor 1001 can also be used to call device control applications stored in memory 1005 to achieve:

[0287] Acquire the target video frame sent by the terminal device. The target video frame is the video frame in the video played by the terminal device that triggers scene exploration.

[0288] Obtain scene information associated with the scene where the target video frame is located;

[0289] Based on the target video footage and scene information, generate the first display data of the three-dimensional exploration space of the scene where the target video footage is located;

[0290] The first display data is sent to the terminal device, enabling the terminal device to display the three-dimensional exploration space based on the first display data.

[0291] It should be understood that the computer device 1000 described in the embodiments of this application can execute the video processing methods described in the embodiments of this application, and can also execute the methods described above. Figure 10 The description of the video processing device 1 in the corresponding embodiments, and the foregoing text Figure 11 The description of the video processing apparatus 2 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0292] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program. When a processor executes this computer program, it can perform the video processing methods described in the various embodiments of this application; therefore, these descriptions will not be repeated here. Additionally, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer storage medium involved in this application, please refer to the description of the method embodiments of this application.

[0293] As an example, the aforementioned computer program can be deployed and executed on a single computer device, or deployed and executed on multiple computer devices located in one location, or executed on multiple computer devices distributed across multiple locations and interconnected via a communication network. These multiple computer devices distributed across multiple locations and interconnected via a communication network can form a blockchain network.

[0294] The aforementioned computer-readable storage medium can be an internal storage unit of the computer device, such as a hard drive or memory. It can also be an external storage device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD) card, or flash card. Furthermore, the computer-readable storage medium can include both internal and external storage units of the computer device. This computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. It can also be used to temporarily store data that has been output or will be output.

[0295] This application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the video processing methods described in the embodiments of this application; therefore, these descriptions will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application.

[0296] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0297] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0298] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A video processing method, characterized in that, The method is applied to a terminal device, and the method includes: Play a video, which contains multiple video frames; In response to a triggered operation of scene exploration of a target video frame in the video, a three-dimensional exploration space of the scene where the target video frame is located is displayed; Based on the exploration operations in the three-dimensional exploration space, the obtained three-dimensional scene image is displayed.

2. The method as described in claim 1, characterized in that, The exploration operation includes a perspective switching operation; The step of displaying the obtained three-dimensional scene image based on the exploration operation in the three-dimensional exploration space includes: Based on the perspective switching operation in the three-dimensional exploration space, the three-dimensional scene after the perspective switching in the three-dimensional exploration space is displayed.

3. The method as described in claim 2, characterized in that, The triggering method for the viewpoint switching operation is any of the following: Triggered by the terminal interface, triggered by the virtual reality device connected to the terminal device, triggered by the keyboard connected to the terminal device, triggered by the mouse connected to the terminal device, or triggered by the positional change of the virtual image in the three-dimensional exploration space. The virtual avatar is an image set up in the three-dimensional exploration space for scene exploration.

4. The method as described in claim 1, characterized in that, The three-dimensional exploration space contains a virtual avatar set up for scene exploration, and the three-dimensional exploration space contains three-dimensional objects generated for the scene where the target video image is located. The exploration operation includes the operation of controlling the virtual avatar to enter the three-dimensional objects in the three-dimensional exploration space. The step of displaying the obtained three-dimensional scene image based on the exploration operation in the three-dimensional exploration space includes: In response to the operation of manipulating the virtual avatar to enter the target three-dimensional object in the three-dimensional exploration space, a three-dimensional scene screen is displayed after the virtual avatar enters the target three-dimensional object.

5. The method as described in claim 1, characterized in that, The three-dimensional exploration space contains three-dimensional objects generated from the scene where the target video image is located, and the exploration operation includes triggering operations on the three-dimensional objects in the three-dimensional exploration space; The step of displaying the obtained three-dimensional scene image based on the exploration operation in the three-dimensional exploration space includes: Based on the triggering operation of the three-dimensional object in the three-dimensional exploration space, the three-dimensional scene screen after the triggered three-dimensional object has been processed according to the triggering method is displayed. The triggering operations on three-dimensional objects in the three-dimensional exploration space include at least one of the following: movement operation, deletion operation, editing operation, and cutout operation.

6. The method as described in claim 1, characterized in that, The exploration operation includes adding data to the three-dimensional exploration space; The step of displaying the obtained three-dimensional scene image based on the exploration operation in the three-dimensional exploration space includes: Based on the addition operation of target data in the three-dimensional exploration space, a three-dimensional scene image is displayed at the corresponding position in the three-dimensional exploration space after the target data has been added.

7. The method as described in claim 1, characterized in that, If the video is playing when the scene exploration is triggered, then the target video screen is the video screen that the video is playing when the scene exploration is triggered; as well as, If the video is paused when the scene exploration is triggered, then the target video screen is the video screen where the video is paused when the scene exploration is triggered.

8. The method as described in claim 1, characterized in that, The trigger operation in response to scene exploration of the target video frame in the video, displaying a three-dimensional exploration space of the scene where the target video frame is located, includes: In response to a scene exploration trigger operation on the target video frame, the scene loading progress is displayed; When the scene loading progress display shows that loading is complete, the three-dimensional exploration space is displayed.

9. The method as described in claim 1, characterized in that, The method further includes: When a scene exploration of the target video frame is triggered, the video playback is paused; After displaying the three-dimensional exploration space, the method further includes: In response to the closing operation of the three-dimensional exploration space, the three-dimensional exploration space is closed, and the video continues to play.

10. The method as described in claim 1, characterized in that, The scene exploration of the target video frame is triggered by a first object, and the method further includes: Based on the sharing operation of the three-dimensional exploration space, obtain the sharing information of the three-dimensional exploration space; Using the shared information, a second object is invited to collaboratively explore the three-dimensional exploration space.

11. The method as described in claim 1, characterized in that, The method further includes: When a save operation for the three-dimensional exploration space is obtained, the three-dimensional exploration space is exported and saved.

12. A video processing method, characterized in that, The method is applied to a server, and the method includes: Acquire the target video frame sent by the terminal device, wherein the target video frame is the video frame in the video played by the terminal device that triggers scene exploration; Obtain scene information associated with the scene where the target video frame is located; Based on the target video frame and the scene information, first display data of the three-dimensional exploration space of the scene where the target video frame is located is generated; The first display data is sent to the terminal device, so that the terminal device displays the three-dimensional exploration space based on the first display data.

13. The method as described in claim 12, characterized in that, The scene information associated with the scene where the target video frame is located includes at least one of the following: Other video frames in the same scene as the target video frame, video frames adjacent to the target video frame, and text description information related to the scene where the target video frame is located.

14. The method as described in claim 12, characterized in that, The first display data for generating a three-dimensional exploration space of the scene where the target video image is located, based on the target video image and the scene information, includes: Obtain a 3D model; The three-dimensional construction model is invoked to perform three-dimensional modeling based on the target video frame and the scene information, thereby generating three-dimensional modeling information of the scene where the target video frame is located; Based on the 3D modeling information, the first display data of the 3D exploration space is generated.

15. The method as described in claim 14, characterized in that, The step of generating the first display data of the three-dimensional exploration space based on the three-dimensional modeling information includes: Obtain the rendering performance information of the terminal device; If the rendering performance information indicates that the rendering performance of the terminal device is sufficient, then the 3D modeling information is used as the first display data; If the rendering performance information indicates that the rendering performance of the terminal device is insufficient, then a three-dimensional scene of the three-dimensional exploration space is rendered based on the three-dimensional modeling information, and the rendered three-dimensional scene is used as the first display data.

16. The method as described in claim 15, characterized in that, If the first displayed data is the three-dimensional modeling information, the terminal device is used to render a three-dimensional scene of the three-dimensional exploration space using the received three-dimensional modeling information, and to display the three-dimensional exploration space based on the rendered three-dimensional scene. If the first display data is a 3D scene image rendered by the server, then the terminal device is used to display the 3D exploration space based on the 3D scene image sent by the server.

17. The method as described in claim 12, characterized in that, The method further includes: The terminal device receives exploration instructions sent by the terminal device, the exploration instructions being generated by the terminal device based on exploration operations in the displayed three-dimensional exploration space; Based on the exploration command, second display data of the three-dimensional scene explored in the three-dimensional exploration space is generated; The second display data is sent to the terminal device, so that the terminal device displays the three-dimensional scene image obtained in the three-dimensional exploration space based on the second display data.

18. A video processing apparatus, characterized in that, The device is used in a terminal device, and the device includes: A playback module is used to play videos, which contain multiple video frames; The display module is used to display the three-dimensional exploration space of the scene where the target video frame is located in response to the trigger operation of scene exploration of the target video frame in the video; The exploration module is used to display the three-dimensional scene obtained from the exploration based on the exploration operations in the three-dimensional exploration space.

19. A video processing apparatus, characterized in that, The device is used in a server, and the device includes: The first acquisition module is used to acquire the target video frame sent by the terminal device, wherein the target video frame is the video frame of the triggered scene exploration in the video played by the terminal device. The second acquisition module is used to acquire scene information associated with the scene scene where the target video frame is located; The generation module is used to generate first display data of the three-dimensional exploration space of the scene where the target video image is located, based on the target video image and the scene information; The sending module is used to send the first display data to the terminal device, so that the terminal device can display the three-dimensional exploration space based on the first display data.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and to execute the steps of the method according to any one of claims 1-17.