Data processing method, apparatus, electronic device, and computer program

The integration of virtual reality technology in video processing allows users to edit and produce personalized videos within a virtual space, enhancing interaction and reducing resource requirements for secondary editing.

JP7714779B2Active Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024510713
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-25
Filing Date
2022-07-26
Publication Date
2025-07-29
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

Traditional video playback scenarios lack the ability to efficiently perform secondary editing on monotonous videos to create personalized content, limiting viewer interaction and engagement.

Method used

A data processing method that utilizes virtual reality technology to enable users to interact with multi-view videos, allowing for scene editing and production within a virtual video space, incorporating object data such as sound, performance, and collaboration, thereby creating a personalized production video without the need for additional shooting.

Benefits of technology

Enriches video display and interaction methods by enabling users to integrate into the video content with minimal data processing resources, improving man-machine interaction efficiency and allowing for personalized video creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714779000001
    Figure 0007714779000001
  • Figure 0007714779000002
    Figure 0007714779000002
  • Figure 0007714779000003
    Figure 0007714779000003
Patent Text Reader

Abstract

The present application discloses a data processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, the method including the steps of: displaying a virtual video space scene corresponding to a multi-view video in response to a trigger operation on the multi-view video; and, in response to a scene editing operation on the virtual video space scene, a first object acquiring object data of the virtual video space scene, the first object referring to an object that initiates a trigger operation on the multi-view video; and, a production video associated with the multi-view video in a virtual display interface, the production video being obtained by performing an editing process on the virtual video space scene based on the object data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of computers, and in particular to data processing methods, apparatus, electronic devices, computer-readable storage media, and computer program products.

[0002] The examples of the present application are based on and claim priority to Chinese patent application No. 202210009571.6, filed on January 5, 2022, and Chinese patent application No. 202210086698.8, filed on January 25, 2022, the entire contents of which are hereby incorporated by reference into the examples of the present application. [Background technology]

[0003] Films and television programs are art forms that use copies, magnetic tapes, film, and memory media as carriers, with the aim of broadcasting them on screens and displays, allowing viewers to enjoy them through a combination of sight and sound. They are a comprehensive form of modern art, and include films, television dramas, programs, and animations.

[0004] For example, videos created based on virtual reality technology use real-life data, which is converted into a phenomenon that people can feel through electronic signals generated by computer technology and combined with various output devices, helping users have an immersive experience during the viewing process.

[0005] However, in traditional video playback scenarios, the displayed content of the video is fixed and too monotonous, and the related art lacks a solution for efficiently performing secondary editing on the monotonous video to create a personalized video according to the user's needs. Summary of the Invention [Problem to be solved by the invention]

[0006] Embodiments of the present application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can fuse viewers with the content of a video in a lightweight manner using relatively few data processing resources, improve the man-machine interaction efficiency, and enrich the video display method and interaction method.

Means for Solving the Problem

[0007] Embodiments of the present application provide a data processing method, which is executed by a computer device, responding to a trigger operation on a multi-view video, and displaying a virtual video space scene corresponding to the multi-view video; responding to a scene editing operation on the virtual video space scene, and obtaining object data of an object in the virtual video space scene by a first object, where the first object refers to an object that starts a trigger operation on the multi-view video; playing a production video associated with the multi-view video in a virtual display interface, where the production video is obtained by performing an editing process on the virtual video space scene based on the object data.

[0008] Embodiments of the present application provide a data processing apparatus, a first response module configured to display a virtual video space scene corresponding to a multi-view video in response to a trigger operation on the multi-view video; a second response module configured to obtain object data of an object in the virtual video space scene by a first object in response to a scene editing operation on the virtual video space scene, where the first object refers to an object that starts a trigger operation on the multi-view video; and a video playback module configured to play a produced video associated with the multi-view video in a virtual display interface, the produced video being obtained by performing an editing process on the virtual video space scene based on the object data.

[0009] An embodiment of the present application provides a computer device, including a processor, a memory, and a network interface; The processor is connected to the memory and the network interface, the network interface is used to provide a data communication network element, the memory is used to store a computer program, and the processor is used to call the computer program and thereby perform the method in the embodiment of the present application.

[0010] An embodiment of the present application provides a computer-readable storage medium having a computer program stored therein, the computer program being suitable for being loaded by a processor to perform a method according to an embodiment of the present application.

[0011] An embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium, a processor of a computing device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to cause the computing device to perform a method in an embodiment of the present application. [Effects of the Invention]

[0012] According to the embodiments of the present application, the first object can view the virtual video space scene from any view in the virtual video space scene. The first object can perform secondary production on the multi-view video according to its own production idea in the virtual video space scene to obtain a production video, which can enrich the video display method and interaction method, and the first object can obtain the production video involved without the need for secondary shooting. Therefore, with relatively few data processing resources, the viewer can be integrated into the video content in a lightweight manner, improving the man-machine interaction efficiency and enriching the video display method and interaction method on the premise of saving data processing resources.

[0013] To more clearly explain the technical solutions in the embodiments of the present application or the prior art, the drawings necessary for use in the following description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are merely some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without the need for creative labor.

Brief Description of the Drawings

[0014] [Figure 1] It is a schematic diagram of the system architecture provided by the embodiments of the present application. [Figure 2a] It is a schematic diagram of a scene of virtual reality-based scene production provided by the embodiments of the present application. [Figure 2b] It is a comparison schematic diagram of the videos provided by the embodiments of the present application. [Figure 3] It is a flowchart of the data processing method provided by the embodiments of the present application. [Figure 4] It is a flowchart of the data processing method for dubbing based on virtual reality provided by the embodiments of the present application. [Figure 5a] It is a schematic diagram of a scene for displaying the dubbing mode list provided by the embodiments of the present application. [Figure 5b]It is a schematic diagram of a scene for displaying a video object capable of voice insertion provided in an embodiment of the present application. [Figure 5c] It is a schematic diagram of a scene for object voice insertion provided in an embodiment of the present application. [Figure 6] It is a flowchart of a data processing method for performing a performance based on virtual reality provided in an embodiment of the present application. [Figure 7a] It is a schematic diagram of a scene for displaying a performance mode list provided in an embodiment of the present application. [Figure 7b] It is a schematic diagram of a scene for selecting a replaceable video object provided in an embodiment of the present application. [Figure 7c] It is a schematic diagram of another scene for selecting a replaceable video object provided in an embodiment of the present application. [Figure 7d] It is a schematic diagram of a scene for object performance based on virtual reality provided in an embodiment of the present application. [Figure 7e] It is a schematic diagram of a scene for transparent display of an object based on virtual reality provided in an embodiment of the present application. [Figure 7f] It is a schematic diagram of a scene for mirror preview based on virtual reality provided in an embodiment of the present application. [Figure 7g] It is a schematic diagram of a first image customization list provided in an embodiment of the present application. [Figure 7h] It is a schematic diagram of a second image customization list provided in an embodiment of the present application. [Figure 7i] It is a schematic diagram of a scene for displaying purchasable virtual items provided in an embodiment of the present application. [Figure 8] It is a flowchart of a data processing method for performing multi-object video production based on virtual reality provided in an embodiment of the present application. [Figure 9a] It is a schematic diagram of a scene for object invitation based on virtual reality provided in an embodiment of the present application. [Figure 9b] It is a schematic diagram of a scene for second object display based on virtual reality provided in an embodiment of the present application. [Figure 10] It is a flowchart of a data processing method for video recording based on virtual reality provided in an embodiment of the present application. [Figure 11a] It is a schematic diagram of a scene of mobile view switching based on virtual reality provided in an embodiment of the present application. [Figure 11b] It is a schematic diagram of a scene of pointing view switching based on virtual reality provided in an embodiment of the present application. [Figure 11c] It is a schematic diagram of a scene of video shooting based on virtual reality provided in an embodiment of the present application. [Figure 12] It is a schematic structural diagram of a data processing device provided in an embodiment of the present application. [Figure 13] It is a schematic structural diagram of a computer device provided in an embodiment of the present application.

DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, in conjunction with the drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only some of the embodiments of the present application, not all of them. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0016] Artificial Intelligence (AI) is a theory, method, and technology that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain optimal results, and it is an application system. In other words, artificial intelligence is a comprehensive technology in computer science, which aims to understand the essence of intelligence and manufacture a new intelligent machine that can react in a way similar to human intelligence. That is, artificial intelligence studies the design principles and implementation methods of various intelligent machines, and enables the machine to have the functions of perception, reasoning, and decision-making.

[0017] Artificial intelligence technology is a comprehensive discipline with a wide range of related fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0018] Computer Vision (CV) is a science that studies how to "make a machine see". More specifically, it refers to machine vision that uses a video camera and a computer to perform recognition and measurement on a target instead of the human eye. Then, further graphics processing is performed to make the computer process it into an image that is more suitable for human eye observation or transmission to an instrument for detection. As a scientific field, computer vision attempts to establish an artificial intelligence system that can study related theories and technologies and obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image search, OCR, video processing, video semantic understanding, video content recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning, and map construction.

[0019] Virtual reality (VR) technology integrates computer, electronic information, and simulation technologies. Its basic implementation method is that a computer simulates a virtual environment, thereby giving people a sense of immersion in the environment. So-called virtual reality, as the name implies, combines the virtual and the real. Theoretically, virtual reality technology is a computer simulation system that can set up and experience a virtual world. It uses a computer to generate a simulated environment and enables users to immerse themselves in that environment. Virtual reality technology uses data in real life and combines it with electronic signals generated by computer technology to convert it into phenomena that people can feel through various output devices. Since these phenomena are not directly visible to us but are the world in reality simulated by computer technology, they are called virtual reality. Virtual reality technology is a combination of simulation technology and multiple technologies such as computer graphics, man-machine interface technology, multimedia technology, sensing technology, and network technology. It is a challenging interdisciplinary technology frontier field and a research field. Virtual reality technology mainly includes aspects such as simulated environment, operation, perception, and sensing devices. The simulated environment includes real-time dynamic three-dimensional panoramic images and sounds generated by a computer.

[0020] The solution provided in the embodiments of the present application relates to the technical field of virtual reality in the computer vision technology of artificial intelligence, and will be specifically described by the following embodiments.

[0021] FIG. 1 is a system architecture diagram provided in an embodiment of the present application. As shown in FIG. 1, the system architecture may include a data processing device 100, a virtual display device 200, and a controller (FIG. 1 takes the controller 300a and the controller 300b as examples. Of course, the number of controllers may be one or more. That is, the controller may include the controller 300a and / or the controller 300b). Both the controller and the virtual display device 200 can be communicatively connected to the data processing device 100. The communication connection does not limit the connection method and may be directly or indirectly connected by a wired communication method, or may be directly or indirectly connected by a wireless communication method. Further, it may be by other methods, and the embodiments of the present application do not limit here. Also, when the data processing device 100 is integrated into the virtual display device 200, the controller 300a and the controller 300b may be further directly wired or wirelessly connected to the virtual display device 200 having data processing capabilities. Here, the virtual display device 200 may be a VR device, or a computer device having an augmented reality (AR) function, or a computer device having a mixed reality (MR) function.

[0022] As shown in FIG. 1, a controller (i.e., controller 300a and / or controller 300b) can send control instructions to data processing device 100. The data processing device 100 can generate related animation data according to the control instructions, and then send the animation data to the virtual display device 200 for display. The virtual display device 200 can be worn on the user's head, such as a virtual reality helmet, and is used to display a virtual world for the user (the virtual world refers to a world similar to the earth or the universe that is independent of the real world, related to the real world, and people can enter in the form of consciousness through virtual reality devices by utilizing computer technology, Internet technology, satellite technology, and the potential capabilities of human consciousness). The controller may be a handle in the virtual reality system, or may be a somatosensory device worn on the user's body, or may be a smart wearable device (e.g., a smart bracelet). The data processing device 100 may be a server or a terminal with data processing capabilities. Here, the server may be an independent physical server, or a server cluster composed of multiple physical servers, or a distributed system. Furthermore, it may be a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Here, the terminal may be an intelligent terminal such as a smartphone, tablet computer, notebook computer, desktop computer, palmtop computer, and mobile internet device (MID).

[0023] The virtual display device 200 and the data processing device 100 may be independent devices or may be integrated (i.e., the data processing device 100 is integrated into the virtual display device 200). To better understand the data processing method provided in the embodiments of the present application, the following embodiments will be described in detail using an example in which a control device implements the data processing method provided in the present application. The control device may also be called a virtual reality device, which refers to the device obtained after integrating the virtual display device 200 and the data processing device 100. The virtual reality device can be connected to a controller to process data and can provide a virtual display interface to display a video screen corresponding to the data.

[0024] Embodiments of the present application provide a data processing method that can realize a selection for a multi-view video in response to a trigger operation on the multi-view video by a first object. A virtual reality device can display a virtual video space scene corresponding to the multi-view video in response to a trigger operation on the multi-view video by the first object. The virtual video space scene is a kind of simulated environment generated by the virtual reality device, providing the first object with simulated sensations about senses such as vision, hearing, and touch, and enabling the first object to have an immersive feeling as if it is on the spot. The first object wearing the virtual reality device can realize that it has entered the virtual video space scene. In the perception of the first object, the virtual video space scene is a three-dimensional spatial scene, and the virtual reality device can display the virtual video space scene in the view where the first object is located to the first object. When the first object walks, the view where the first object is located can change accordingly, and the virtual reality device can obtain the view where the first object is located in real time and display the virtual video space scene in the view where the first object is located in real time to the first object, so that the first object can perceive that it is walking in the virtual video space scene. In addition, the first object can interact with the virtual reality device and the virtual video space scene.

[0025] As shown in conjunction with FIGS. 2a to 2b, FIG. 2a is a schematic diagram of a scene for virtual reality-based scene production provided in an embodiment of the present application. The realization process of the scene of the application process can be realized based on virtual reality equipment. As shown in FIG. 2a, it is assumed that a first object (for example, object A) is wearing a virtual reality device 21 (that is, the device after integrating the virtual display device 200 and the data processing device 100 described in FIG. 1 above) and a virtual reality handle 22 (that is, the controller described in FIG. 1 above). After the virtual reality device 21 responds to a trigger operation on the multi-view video by the object A, it can enter a virtual video space scene 2000 corresponding to the multi-view video. The multi-view video may refer to a video with a six degrees of freedom (6DOF, Six degrees of freedom tracking) browsing method. The multi-view video may be composed of a plurality of virtual video screens in different views. For example, the multi-view video may include a target virtual video screen in a target view. The target virtual video screen may be obtained by an image acquisition device (for example, a camera) acquiring the real space scene in the target view. The virtual reality device outputs the target virtual video screen by three-dimensional (3D, three-dimensional) display technology at a virtual display interface 23, whereby the real space scene in the target view can be simulated. The simulated real space scene in the target view is the virtual video space scene in the target view. As shown in FIG. 2a, the virtual reality device 21 can display a virtual video screen 201 on the virtual display interface 23. In the virtual video screen 201, a video object B corresponding to character A and a video object C corresponding to character B are talking. The virtual video screen 201 may be a corresponding video screen in the default view of the virtual video space scene 2000. The default view may be the main lens view, that is, the ideal lens view for the director.The virtual reality device 21 can display a virtual video screen 201 for object A by means of 3D display technology. Object A can feel as if video object B and video object C are standing right in front of it talking. Object A is viewing the virtual video space scene 2000 in the default view. In the perception of object A, the virtual video space scene 2000 is a three-dimensional real space scene, and object A can feel that it can walk freely in the virtual video space scene 2000. Therefore, the viewing view of object A with respect to the virtual video space scene 2000 can change as object A walks. To immerse object A in the virtual video space scene 2000, the virtual reality device 21 can obtain the position where object A is in the virtual video space scene 2000 and determine the view where object A is located. Subsequently, by obtaining the virtual video screen in the view where object A is located in the virtual video space scene 2000 and continuously displaying it through the virtual display interface, object A can see the virtual video space scene 2000 in the located view, thereby giving object A a sense of immersion. That is, object A is made to perceive that it is located in the virtual video space scene 2000 at this moment and can freely walk and view the virtual video space scene 2000 in different views. In addition to realizing the simulation of object A's vision by means of 3D display technology, the virtual reality device 21 can further generate electronic signals by computer technology and combine them with various output devices to realize the simulation of object A's other perceptions. For example, the virtual reality device 21 can utilize surround sound technology, that is, adjust parameters such as the volume of different channels to simulate the sense of direction of sound, thereby making it possible to bring a real auditory experience to object A. However, the embodiments of the present application are not limited thereto here.

[0026] With the virtual reality device 21, the object A is positioned to immerse in the virtual video space scene corresponding to the multi-view video in the first view, and can feel the scenario of the multi-view video. When the object A has another production idea for the scenario of the multi-view video, the object A can perform scene production, such as adding sound and performance, etc. in the virtual video space scene 2000 corresponding to the multi-view video by the virtual reality device 21. The object A can further invite its friends to perform scene production for the multi-view video together. The virtual reality device 21 can further independently display a scene production bar 24 in the virtual display interface 23. As shown in FIG. 2a, the object A can see that the scene production bar 24 is floating forward and displayed. The scene production bar 24 may include a sound addition control 241, a performance control 242, and an object invitation control 243. With the virtual reality handle 22, the object A can trigger a scene editing operation for the virtual video space scene 2000, that is, a trigger operation for a certain control in the scene production bar 24. Subsequently, the object A can change the scenario of the multi-view video according to its own production idea, and obtain the production video generated after performing the scenario change for the multi-view video by the virtual reality device 21. For example, by triggering the sound addition control 241, the object A can add a narration to the multi-view video or change the line expression of a certain character. Or, by triggering the performance control 242, the object A can add a new character in the multi-view video or replace a certain character in the video object, thereby performing production performance. Or, by triggering the object invitation control 243, the object A can invite friends to perform scenario production for the multi-view video together.After the virtual reality device 21 responds to the scene editing operation by object A on the virtual video space scene 2000 (i.e., an operation to trigger a control on the scene creation bar 24 described above), object A can perform corresponding scene creation in the virtual video space scene 2000, such as adding sound and performing, and the virtual reality device 21 can obtain object data of object A in the virtual video space scene 2000. Then, the virtual reality device 21 merges the obtained object data with the virtual video space scene 2000 to obtain a created video. For the process of acquiring object data and implementing object data merging, please refer to the specific description of the subsequent embodiments.

[0027] As referenced in conjunction with FIG. 2b, FIG. 2b is a comparative schematic diagram of the video provided in the embodiment of the present application. As shown in FIG. 2b, in the virtual video screen 202 at time A in the multi-view video, the video object B corresponding to character A and the video object C corresponding to character B are talking at this time. Assuming that object A likes video object B very much and hopes that it is itself talking to video object B at this time, in this case, object A replaces video object C and then can perform as character B. Object A triggers the performance control 242 in the scene shown in FIG. 2a above and then performs its performance in the virtual video space scene 2000, that is, can talk to video object B in the virtual video space scene 2000. The virtual reality device 21 can obtain the object data when object A is talking to the virtual reality object B. The object data may be data such as the skeleton, movement, expression, image of object A, and its position in the virtual video space scene. The virtual reality device 21 sets up the video object A with the object data, then cancels the display of the video object C in the multi-view video, and then fuses the video object A into the multi-view video with the video object C filtered to obtain a production video. As can be seen from the above, through the production in the virtual video space scene 2000 by object A, object A becomes the performer of character B.

[0028] As can be understood, in the specific embodiments of the present application, when the relevant object data is utilized in the above embodiments of the present application for specific products or technologies, it is necessary to obtain the permission or consent of the user, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0029] In the embodiments of the present application, after object A enters the virtual video space scene by the virtual reality device, the virtual reality device displays a virtual video space scene corresponding to the multi-view video, and immerses object A in the virtual video space scene. Also, in response to the production operation of object A on the virtual video space scene, object data such as sound insertion and performance performed by object A in the virtual video space scene is acquired, and then the object data is continuously fused with the virtual video space scene to obtain a production video. Therefore, the secondary production of the multi-view video by object A can be efficiently realized, and the display method and interaction method of the multi-view video can be enriched.

[0030] As shown in FIG. 3, FIG. 3 is a flowchart of a data processing method provided in the embodiments of the present application. The data processing method may be executed by a virtual reality device. For ease of understanding, the embodiments of the present application will be described by taking the method being executed by a virtual reality device as an example. The data processing method may include at least the following steps S101 to S103.

[0031] Step S101: In response to a trigger operation on the multi-view video, display a virtual video space scene corresponding to the multi-view video.

[0032] The multi-view video may be a produced video such as a movie, television drama, or musical drama with a 6DOF browsing method. The multi-view video may include a virtual video screen in at least one specific view of a real-space scene. The virtual reality device displays a target virtual video screen in a target view in a 3D format in a virtual display interface. The target view is a view obtained according to the real-time position and real-time pose of the first object. For example, if the first object is currently standing facing south and looking up, the view indicated when the first object is standing facing south and looking up is the target view. The target view may also be obtained in response to a view selection operation of the first object, and the first object wearing the virtual reality device can view the virtual video space scene in the target view. When the virtual reality device switches between and displays the virtual video screen in different views using the virtual display interface, the first object can view the virtual video space scene in the different views. The virtual video space scene is not a real space scene, but a simulation of a real space scene. The objects, environment, sounds, etc. in the virtual video space scene may all be generated by a virtual reality device, or may be a combination of virtual and real. That is, some of the objects, environment, sounds, etc. may be generated by the virtual reality device, or some of the objects, environment, sounds, etc. may actually exist in the real space scene where the first object is located. When the first object wears the virtual reality device, the virtual reality device connects it to various output devices through electronic signals generated by computer technology, and enters into a virtual video space scene corresponding to multi-view video. The first object can view the virtual video space scene in the view where it is located. When the first object senses the virtual video space scene, it is the same as the real space scene, that is, the first object feels as if it is in the real space scene.For example, in the virtual video space scene 2000 shown in FIG. 2a above, in the perception of object A, object A is located in the virtual video space scene 2000, and object A can experience the development of a multi-view video scenario in a first view. It is necessary to explain that although the first object perceives itself as being in the virtual video space scene, what the first object "sees" is the virtual video space scene in the view in which it is located, and in fact, what the first object sees comes from the virtual video screen in the virtual display interface of the virtual reality device.

[0033] A possible process for shooting and creating a multi-view video may be as follows. Use a panoramic camera to shoot the real environment (i.e., the real space scene) corresponding to the multi-view video, obtain the point cloud data of the complete scene, and then perform fitting modeling using the point cloud data. Based on the fitting modeling results, perform fusion optimization of the multi-scene model. At the same time, in a special studio, perform complete human scanning modeling on the actor and complete the creation of virtual objects (i.e., video objects corresponding to the actor). Note that in the process of shooting the multi-view video, the actor needs to wear green clothes, attach marks and capture points on the body, and perform in the real environment corresponding to the multi-view video. Subsequently, use the main camera of the film director and a group of cameras including multiple cameras to perform complete recording capture on the actor. By combining the captured multi-angle data, the real-time motion data and expression data of each actor can be obtained. Place the created video object in the optimized multi-scene model, and then continue to drive the body skeleton and expression of the video object with the real-time motion data and expression data, add light and shadow effects in the modeled real environment, and finally obtain the final multi-view video. It should be noted that before performing complete human scanning and complete recording capture on the actor, the permission of the actor himself / herself needs to be obtained, and the obtained motion data and expression data are only used for the creation of the multi-view video and not for other purposes such as data analysis.

[0034] Step S102: In response to a scene editing operation on the virtual video space scene, obtain the object data of the first object in the virtual video space scene.

[0035] The first object refers to an object that initiates a trigger operation on the multi-view video. During the process of playing the multi-view video, the first object can trigger a scene editing operation on the virtual video space scene. After responding to the scene editing operation, the virtual reality device can acquire object data when the first object creates and performs a scene in the virtual video space scene. The scene creation may include sound recording, performance, and object placement. The object data may include data such as audio data, posture data, image data, and position data. The audio data corresponds to the audio of the first object, the posture data corresponds to the motion of the first object, the image data corresponds to the appearance of the first object, and the position data corresponds to the position of the first object in the virtual video space scene. One possible embodiment of how the virtual reality device acquires object data is as follows: The virtual reality device may include multiple real-time capture cameras and audio recording controls, and the cameras can capture images of the first object from different views in real time to obtain captured images from the different views. The virtual reality device can fuse multiple screens and then calculate real-time object data of the first object, including posture data and image data of the first object. The posture data may also include limb skeleton data and facial expression data of the first object. The virtual reality device can determine an image of a performance object associated with the first object according to the image data, determine a real-time movement of the performance object using the limb skeleton data, and determine a real-time facial expression of the performance object using the facial expression data. The recording control can acquire audio data of the first object in real time. When acquiring the object data of the first object, the virtual reality device needs to obtain permission from the first object, and the acquired object data is used only for creating the produced video.For example, among the object data, the audio data is only used to display the audio of the first object in the production video, the pose data is only used to display the actions and expressions of the video object corresponding to the first object in the production video, and the image data is only used to display the image and clothing of the video object corresponding to the first object in the production video. These data are not used for other purposes such as data analysis, and the same applies to the object data obtained in the subsequent embodiments of this application, which will not be described in detail again.

[0036] In the perception of the first object, the first object is located in the virtual video space scene. The first object can perform actions such as walking, talking, laughing, and crying in the virtual video space scene. Therefore, the first object can experience the scene production performance for the multi-view video scenario in the virtual video space scene. The scene production performance refers to the ability of the first object to perform a performance display for a character it likes, that is, to replace the video object corresponding to the character in the virtual video space scene and perform a performance of lines, actions, and expressions, to advance the scenario together with other video objects in the virtual video space scene, or to co-star and interact with other video objects of interest in the scenario. The first object may further perform in the virtual video space scene as a newly added character, and the first object may further invite the second object to perform together in the virtual video space scene.

[0037] In response to a scene editing operation for sound insertion, the virtual reality device acquires the sound data of the first object as object data. For example, in the scene shown in FIG. 2a above, after object A (first object) triggers the sound insertion control 271, the virtual reality device 21 can acquire sound data corresponding to the sound emitted by object A in the virtual video space scene 2000. In response to a scene editing operation for performance, the virtual reality device can acquire the posture data and image data of the first object as object data. For example, in the scene shown in FIG. 2a above, after object A triggers the performance control 272, the virtual reality device 21 can acquire posture data corresponding to the motion form of object A in the virtual video space scene 2000 and image data corresponding to the object image of object A. In response to a scene editing operation for object invitation, the virtual reality device may change the secondary production of the multi-view video from a single production to a collaborative production by multiple people. In other words, the first object and a second object invited by the first object can enter a virtual video space scene corresponding to the same multi-view video. In this case, the virtual reality device can obtain object data when the first object performs a scene-making performance in the virtual video space scene, and can also obtain target object data when the second object performs a scene-making performance in the virtual video space scene. The first object can simultaneously perform sound input and scene-making for the performance, and in this case, the virtual reality device can simultaneously obtain the voice data, posture data, and image data of the first object as object data.

[0038] A possible implementation process for a virtual reality device to obtain object data of an object in a virtual video space scene in response to a scene editing operation on the virtual video space scene may be as follows. The virtual reality device responds to a scene editing operation on the virtual video space scene by the first object, and then displays a video clip input control for the virtual video space scene on the virtual display interface. The virtual reality device obtains clip progress information for the input multi-view video in response to an input operation on the video clip input control, and may use the video clip indicated by the clip progress information as the video clip to be produced. In the process of playing the video clip to be produced, the virtual reality device can obtain the object data of the object in the virtual video space scene where the first object is located. According to the embodiments of the present application, since the video clip to be produced can be directly obtained, it is easy to directly produce the video clip to be produced later, and the waiting time for the first object to wait until the video clip to be produced is played can be saved, thereby improving the man-machine interaction efficiency.

[0039] The first object can select a video clip from the multi-view video as the video clip to be produced, and then only perform the performance on the video clip to be produced. For example, the playback period of the multi-view video is 2 hours, and the character that the first object attempts to replace appears only from the 50th minute to the 55th minute of the multi-view video. If the virtual reality device plays the multi-view video, the first object needs to wait for 50 minutes to perform. Therefore, the first object may select the video clip from the 50th minute to the 55th minute as the video clip to be produced. At this time, the virtual reality device directly plays the video clip to be produced, and the first object can perform the performance of the character to be replaced in the virtual video space scene, and the virtual reality device can obtain the object data when the first object performs. The real space scenes corresponding to different playback times of the multi-view video may be different. For example, when the playback time of the multi-view video is the 10th minute, the corresponding real space scene is a scene where object D is sleeping, and when the playback time of the multi-view video is the 20th minute, the corresponding real space scene is a scene where object D is singing. The virtual video space scene perceived by the first object can change based on the real space scene corresponding to the playback time of the multi-view video. Therefore, when the playback time of the multi-view video is the 10th minute, video object D (video object D is generated based on object D) in the virtual video space scene perceived by object A is sleeping, and when the playback time of the multi-view video is the 20th minute, video object D in the virtual video space scene perceived by object A is singing.

[0040] During the process of playing the video clip to be produced, the virtual reality device may display a playback progress control bar on the virtual display interface. The playback progress control bar may include a pause control, a play control, and a multiple control. In response to a trigger operation on the pause control, the virtual reality device can pause the playback of the video clip to be produced. In response to a trigger operation on the play control, the virtual reality device can continue to play the video clip to be produced. In response to a selection operation on the multiple control, the virtual reality device can adjust the playback speed of the video clip to be produced according to the selected playback multiple. According to the embodiments of the present application, operations such as pausing, resuming playback, and variable-speed playback can be performed on the video clip to be produced, and the video clip to be produced can be flexibly adjusted, so that the performance needs of the first object can be met in real time.

[0041] Step S103: Play the production video associated with the multi-view video on the virtual display interface.

[0042] The production video is obtained by performing an editing process on the virtual video space scene based on the object data.

[0043] When the acquisition of object data is completed, the virtual reality device can obtain a production video by fusing the object data and the virtual video space scene. When the object data is voice data acquired by the virtual reality device in response to a scene editing operation for dubbing, the virtual reality device uses the voice data based on the scene editing operation of the first object to replace the video voice data of a certain video object in the virtual video space scene, or can superimpose the voice data on the virtual video space scene. When the object data is pose data and image data acquired by the virtual reality device in response to a scene editing operation for performance, the virtual reality device generates a performance object having the pose data and the image data, and uses the performance object to replace a certain video object in the virtual video space scene, or can directly add the performance object to the virtual video space scene. The object data is object data acquired when the virtual reality device responds to a scene editing operation for object invitation, and when the virtual reality device acquires the target object data of the second object invited by the first object, the virtual reality device can fuse the object data and the target object data together into the virtual video space scene.

[0044] The produced video can correspond to two data forms. One data form is the video mode, for example, a file of the Moving Picture Experts Group 4 (MP4). At this time, the produced video can not only be played on the virtual reality device, but also be played on other terminal devices with video playback functions. Another data form is the recording file for the virtual video space scene. The recording file is a file in a specific format and contains all the data recorded this time. The first object can open the file in this specific format using an editor. That is, on the computer side, the digital scene (the digital scene includes the virtual video space scene that the first object hopes to hold in reserve and the performance object with object data) can be browsed and edited. The browsing and editing of the digital scene on the computer side is similar to the real-time editing operation of a game engine. In the editor, the user can perform processing processes such as speed, sound, light, screen, filter, and special styling on the entire content, and can also perform beautification and tone processing on the video object corresponding to the character or the performance object. For the performance, several preset effects and props can also be added. Finally, a new produced video can be generated and obtained. The first object may save the produced video in the local / cloud space of the machine, or may send it to other objects through VR social applications, non-VR social applications, and social short video applications, etc.

[0045] In the embodiment of the present application, the first object can be sensed by a virtual reality device as being located in a virtual video space scene corresponding to a multi-view video. The virtual video space scene simulates a real space scene corresponding to the multi-view video. The first object is located in the virtual video space scene corresponding to the multi-view video, and can feel the emotional expression of the character in the multi-view video in the first view and deeply experience the scenario. The first object can further perform scene production in the virtual video space scene. The virtual reality device acquires object data of the first object when the first object performs scene production, and subsequently, can fuse the object data and the virtual video space scene to obtain a production video. As can be seen from the above, since the first object can break through physical limitations, there is no need to arrange a real space scene corresponding to the multi-view video at the cost of time and cost. Secondary production for the multi-view video can be performed in the virtual video space scene corresponding to the real space scene, enriching the display method and interaction method of the multi-view video. At the same time, arranging a real space scene corresponding to the multi-view video at the cost of time and cost can be avoided, saving production costs.

[0046] In order to better understand the process in FIG. 3 above, where the virtual reality device acquires the audio data of the first object as object data in response to a scene editing operation for dubbing and obtains a production video according to the object data, please refer to FIG. 4. FIG. 4 is a flowchart of a data processing method for dubbing based on virtual reality provided in the embodiment of the present application. The data processing method may be executed by a virtual reality device. For ease of understanding, the embodiment of the present application will be described by taking the example that the method is executed by a virtual reality device. The data processing method may include at least the following steps S201 to S203.

[0047] Step S201: In response to a trigger operation on the multi-view video, display a virtual video space scene corresponding to the multi-view video.

[0048] For implementing step S201, reference can be made to step S101 in the embodiment corresponding to FIG.

[0049] In some embodiments, the scene editing operation includes a trigger operation on a sound input control in the virtual display interface, and in step 102, in response to the scene editing operation on the virtual video space scene, obtaining object data of a first object in the virtual video space scene may be realized by the following technical solution: playing a multi-view video, and in response to the trigger operation on the sound input control in the virtual display interface, obtaining audio data of the first object in the virtual video space scene, and determining the audio data of the first object as object data applied to the multi-view video. The embodiments of the present application can obtain the audio data of the first object as object data by triggering the sound input control, thereby efficiently obtaining object data and improving human-machine interaction efficiency.

[0050] The audio data of the first object includes at least one of object audio data and background audio data. The steps of acquiring the audio data of the first object in the virtual video space scene in response to a trigger operation on a sound input control in the virtual display interface and determining the audio data of the first object as object data to be applied to the multi-view video may be realized by the following steps S202 to S206.

[0051] Step S202: In response to a trigger operation on the sound insertion control in the virtual display interface, a sound insertion mode list is displayed.

[0052] The voice dubbing mode list includes an object voice dubbing control and a background voice dubbing control.

[0053] The virtual reality device may independently display the voice dubbing control, for example, the voice dubbing control 241 shown in FIG. 2a above, in the virtual display interface. At this time, the scene editing operation for voice dubbing may be a trigger operation for the voice dubbing control 241.

[0054] The object voice dubbing control and the background voice dubbing control respectively correspond to two types of voice dubbing modes. The object voice dubbing control corresponds to the object voice dubbing mode. At this time, the first object can perform voice dubbing on at least one voice-dubbable video object of the multi-view video, that is, the voice of the first object can replace the original voice of the voice-dubbable video object. The background voice dubbing control corresponds to the background voice dubbing mode, and the first object can perform voice dubbing on the entire multi-view video. That is, in a situation where the original character's voice and background sound coexist in the multi-view video, the virtual reality device can record the additional sound of the first object, and use the additional sound of the first object obtained by recording as narration, sound effects, etc.

[0055] As shown in conjunction with FIG. 5a, FIG. 5a is a schematic diagram of a scene for displaying a voice recording mode list provided in an embodiment of the present application. Based on the scene shown in FIG. 2a above, in response to a trigger operation in which object A selects the voice recording control 241 by the virtual reality handle 22, the virtual reality device 21 can display the voice recording mode list 51 on the virtual display interface 23 in response to the trigger operation on the voice recording control 241. The voice recording mode list 51 may include an object voice recording control 511 and a background voice recording control 512, and the voice recording mode list 51 may be independently displayed on the virtual video screen 500. The virtual video screen 500 is used to display a virtual video space scene in the view where object A is located.

[0056] Step S203: In response to a selection operation on the object voice recording control, display a voice-recordable video object, and in response to a selection operation on the voice-recordable video object, set the selected voice-recordable video object as the object to be voice-recorded.

[0057] The voice-recordable video objects belong to the video objects displayed in the multi-view video.

[0058] The virtual reality device may set only the currently displayed video object on the virtual display interface as the voice-recordable video object, or may set all the video objects displayed in the multi-view video as the voice-recordable video objects. The first object can select one or more video objects from the voice-recordable video objects and perform voice recording as the object to be voice-recorded, and the virtual reality device may highlight the object to be voice-recorded.

[0059] An example will be described in which the virtual reality device sets the video object currently displayed in the virtual display interface as a video object that can be inserted with sound. As shown in FIG. 5b, FIG. 5b is a schematic diagram of a scene displaying a video object that can be inserted with sound provided in an embodiment of the present application. Assuming the scene shown in FIG. 5a above, the virtual reality device 21 can set the video object B and the video object C currently displayed in the virtual display interface 23 as video objects that can be inserted with sound after responding to the trigger operation of the object insert control 511 by the object A, and highlight the video object that can be inserted with sound. As shown in FIG. 5b, the virtual reality device 21 may highlight the video object B by a dashed box 521 and highlight the video object C by a dashed box 522 in the virtual display interface 23. At this time, the object A can know that the video object B and the video object C are video objects that can be inserted with sound, and then use the virtual reality handle 22 to determine the object to be inserted with sound.

[0060] Step S204: In the process of playing the multi-view video, object audio data of a first object is obtained according to an object to which audio is to be added.

[0061] By the above steps S202 to S204, the object voice data can be flexibly acquired and the object voice data can be determined as the object data according to the user needs, thereby improving the efficiency of man-machine interaction in the voice dimension.

[0062] In a process of playing multi-view video, a process of acquiring object audio data of a first object based on an object into which sound is to be added may be as follows: During a process of playing multi-view video, a virtual reality device performs a mute process on video audio data corresponding to the object into which sound is to be added. When the object into which sound is to be added is in a speaking state, the virtual reality device displays text information and soundtrack information corresponding to the video audio data, acquires object audio data of the first object, and then determines the object audio data as object data. The text information and soundtrack information can be used to notify the first object of the lines and sound intensity corresponding to the original sound of the object into which sound is to be added. The speaking state refers to a state in which the object into which sound is to be added speaks in the multi-view video. According to an embodiment of the present application, the text information and soundtrack information can assist the user in adding sound. Furthermore, since the object into which sound is to be added is in a speaking state, the audio data can be ensured to match the screen, thereby improving the success rate of sound addition. This avoids repeated acquisition of object audio data, thereby improving resource utilization and man-machine interaction efficiency.

[0063] As referred to in conjunction with FIG. 5c, FIG. 5c is a schematic diagram of a scene of object sound insertion provided in an embodiment of the present application. On the premise of the scene shown in FIG. 5b above, object A can select a video object that can be sound-inserted by the virtual reality handle 22. Assuming that object A selects video object C as the object to be sound-inserted, when the virtual reality device plays the multi-view video, it continuously highlights the object to be sound-inserted, that is, video object C, with the dashed box 522, and can perform a muting process on the video audio data corresponding to video object C. In this way, when video object C is in a speaking state, object A cannot hear the original sound of video object C. The virtual reality device 21 may further display notification information 53 corresponding to the video audio data of video object C when video object C is in a speaking state. The notification information 53 includes text information and sound track information. Object A can know that it should perform sound insertion on video object C at this time according to the displayed notification information 53. When the virtual reality device displays the notification information 53, it can collect sound from object A and obtain the object audio data of object A.

[0064] Step S205: In the multi-view video, use the object audio data to perform a replacement process on the video audio data corresponding to the object to be sound-inserted in the multi-view video to obtain a production video.

[0065] After the virtual reality device replaces the video audio data of the object to be sound-inserted with the object audio data of the first object in the multi-view video, it can obtain the production video after the first object performs sound insertion on the multi-view video. In the production video, when the video object to be sound-inserted is in a speaking state, the sound corresponding to the object audio data of the first object can be played.

[0066] Step S206: In the process of playing a multi-view video in response to a selection operation on the background sound insertion control, obtain the background sound data of the first object and determine the background sound data as object data. Through the above step S206, the background sound data can be flexibly obtained and determined as object data according to user needs. Thereby, the man-machine interaction efficiency can be improved in the audio dimension.

[0067] Specifically, taking the above Figure 5a as an example, when object A attempts to insert background sound, object A can trigger the background sound insertion control 512. Thereafter, in the process of playing the multi-view video, the virtual reality device 21 can obtain the background sound data corresponding to the voice of object A.

[0068] Step S207: In the multi-view video, superimpose the background sound data on the multi-view video to obtain a production video.

[0069] The virtual reality device may add the voice of object A to the multi-view video based on the background sound data, thereby obtaining a production video.

[0070] By adopting the data processing method provided in the embodiments of the present application, the first object can insert sound for the characters in the multi-view video in the virtual video space scene corresponding to the multi-view video, or add background sound or narration to obtain a production video, and fuse the viewer with the video content in a lightweight manner with relatively few data processing resources, improve the man-machine interaction efficiency, and enrich the video display method and interaction method.

[0071] To better understand the process in which the virtual reality device in FIG. 3 above obtains the pose data and image data of the first object as object data in response to a scene editing operation on performance, and then obtains a production video according to the object data, please refer to FIG. 6. FIG. 6 is a flowchart of a data processing method for performing performance based on virtual reality provided in an embodiment of the present application. The data processing method may be executed by a virtual reality device. For ease of understanding, the embodiment of the present application will be described by taking the example that the method is executed by a virtual reality device. The data processing method may include at least the following steps S301 to S303.

[0072] Step S301: In response to a trigger operation on the multi-view video, display a virtual video space scene corresponding to the multi-view video.

[0073] Specifically, for the implementation of step S301, reference can be made to step S101 in the embodiment corresponding to FIG. 3 above.

[0074] In some embodiments, the scene editing operation includes a trigger operation for performance control in the virtual display interface. In step 102, in response to the scene editing operation on the virtual video space scene, obtaining the object data of the first object in the virtual video space scene may be realized by the following technical solutions. Play the multi-view video, and in response to the trigger operation for performance control in the virtual display interface, obtain the pose data and image data of the first object in the virtual video space scene, and determine the pose data and image data as the object data applied to the multi-view video. In the production video, a performance object associated with the first object is included, and the performance object in the production video is displayed based on the pose data and image data. The embodiments of the present application can efficiently obtain object data and improve the man-machine interaction efficiency by triggering performance control to obtain the voice data of the first object and use it as object data.

[0075] In response to the trigger operation for performance control in the virtual display interface, the above steps of obtaining the pose data and image data of the first object in the virtual video space scene and determining the pose data and image data as object data may be realized by the following steps S302 to S304 and step 306.

[0076] Step S302: Display a performance mode list in response to the trigger operation for performance control in the virtual display interface.

[0077] The performance mode list includes character replacement control and character setup control.

[0078] For a scenario displayed in a multi-view video, the first object can perform in a virtual video space scene corresponding to the multi-view video according to its own production idea, and obtain a production video after participating in the multi-view video scenario. The virtual reality device may independently display performance control, for example, the performance control 242 shown in FIG. 2b above, in the virtual display interface. At this time, the scene editing operation for performance may be a trigger operation for the performance control in the virtual display interface.

[0079] The character replacement control and the character setup control respectively correspond to two types of performance modes. The character replacement control corresponds to the character replacement mode. At this time, the first object can perform character performance by selecting and replacing a video object corresponding to any character in the multi-view video. The character setup control corresponds to the character production mode. At this time, the first object can perform by newly adding one character in the virtual video space scene corresponding to the target time of the multi-view video. The character may be a character that has appeared in the virtual video space scene corresponding to other times of the multi-view video, or may be a completely new character customized by the first object.

[0080] As shown in conjunction with FIG. 7a, FIG. 7a is a schematic diagram of a scene for displaying a performance mode list provided in an embodiment of the present application. Assuming that, based on the scene shown in FIG. 2a above, object A selects performance control 242 by virtual reality handle 22, virtual reality device 21 can display performance mode list 71 on virtual display interface 23 in response to a trigger operation on performance control 242. The performance mode list 71 may include character replacement control 711 and character setup control 712, and the voice dubbing mode list 51 may be independently displayed on virtual video screen 700. Virtual video screen 700 is used to display a virtual video space scene in the view where object A is located.

[0081] Step S303: In response to a trigger operation on the character replacement control, display replaceable video objects, and in response to a selection operation on the replaceable video objects, set the selected replaceable video object as the character replacement object.

[0082] The replaceable video objects belong to the video objects displayed in the multi-view video.

[0083] In response to a trigger operation on the character replacement control, the virtual reality device displays a replaceable video object, and in response to a selection operation on the replaceable video object, one feasible implementation process of using the selected replaceable video object as a character replacement object may be as follows. In response to a trigger operation on the character replacement control, the virtual reality device determines the video object currently displayed on the object virtual video screen as a replaceable video object, and in response to a marking operation on the replaceable video object, it displays the marked replaceable video object according to the first display method. Here, the first display method is different from the display methods of other video objects other than the replaceable video object, and the first display method may be a highlighting display method. For example, a filter or the like is added to the marked replaceable video object, and the marked replaceable video object is used as the character replacement object. The object virtual video screen is used to display the virtual video space scene in the current view. According to the embodiments of the present application, the marked replaceable video object can be highlighted, thereby playing a notification role in the user's performance process and improving the man-machine interaction efficiency.

[0084] As also referred to in FIG. 7b, FIG. 7b is a schematic diagram of a scene for selecting a replaceable video object provided in an embodiment of the present application. Given the scene shown in FIG. 7a above, assuming that object A triggers character replacement control 711 by virtual reality handle 22, virtual reality device 21 can set the video object currently displayed in virtual video screen 700 (i.e., the above object virtual video screen) as a replaceable video object. As shown in FIG. 7b, video object B and video object C can be replaceable video objects. Virtual reality device 21 can highlight replaceable video objects, for example, virtual reality device 21 may highlight video object B by dashed box 721 and video object C by dashed box 722 in virtual display interface 23. At this time, object A can know that video object B and video object C are replaceable video objects. Object A can mark replaceable video objects by virtual reality handle 22. Assuming that object A marks video object C, the virtual reality device 21 may newly highlight the marked video object C. For example, the solid box 723 may replace the dashed box 722 to highlight the video object C, and then the virtual reality device 21 may make the video object C a character replacement object. The virtual reality device 21 may then stop highlighting and then remove the display of the video object C in the virtual display interface 23. As shown in FIG. 7b, the virtual reality device 21 may switch from displaying a virtual video screen 700 including the video object C to displaying a virtual video screen 701 not including the video object C.

[0085] In response to a trigger operation on the character replacement control, the virtual reality device displays replaceable video objects, and in response to a selection operation on the replaceable video objects, one feasible implementation process of using the selected replaceable video object as a character replacement object may be as follows. In response to a trigger operation on the character replacement control, the virtual reality device displays at least one video clip corresponding to a multi-view video, and in response to a selection operation on the at least one video clip, it displays the video objects included in the selected video clip, determines the video objects included in the selected video clip as replaceable video objects, and in response to a marking operation on the replaceable video objects, it displays the marked replaceable video objects according to a first display method. Here, the first display method is different from the display methods of other video objects except for the replaceable video objects, and the first display method may be a highlighting display method. For example, add a filter or the like to the marked replaceable video object, and use the marked replaceable video object as a character replacement object. According to the embodiments of the present application, the marked replaceable video object can be highlighted, thereby playing a notification role in the user's performance process and improving the man-machine interaction efficiency.

[0086] As shown in conjunction with FIG. 7c, FIG. 7c is a schematic diagram of another scene for selecting a replaceable video object provided in an embodiment of the present application. Assuming that, on the premise of the scene shown in FIG. 7a above, when object A triggers the character replacement control 711 by the virtual reality handle 22, the virtual reality device may display at least one video clip corresponding to the multi-view video. As shown in FIG. 7c, after the virtual reality device 21 responds to the trigger operation on the character replacement control 711, it displays the video clip 731, the video clip 732, and the video clip 733, and the video clip 731, the video clip 732, and the video clip 733 all belong to the multi-view video. The video clip 731 includes the video object B and the video object C, the video clip 732 includes the video object D, and the video clip 733 includes the video object E. The virtual reality device 21 may highlight the selected video clip with a black frame. As can be seen from FIG. 7c, object A selects the video clip 731 and the video clip 732. After the selection by object A is completed, the virtual reality device 21 may use the video object B, the video object C, and the video object D as replaceable video objects, and then display the replaceable video objects. Object A may select at least one replaceable video object from the replaceable video objects as a character replacement object. As shown in FIG. 7c, object A may mark the replaceable video object with the virtual reality handle 22. Assuming that object A marks the video object B, the virtual reality device 21 can highlight the marked video object B, for example, surround the video object B with a solid line box 724. When object A confirms the end of the marking, the virtual reality device 21 can use the video object B as a character replacement object.Subsequently, the virtual reality device 21 can stop the highlighted display and then cancel the display of the video object B on the virtual display interface 23. As shown in FIG. 7c, the virtual reality device 21 can switch the virtual video screen 700 to a virtual video screen 702 that does not include the video object B and display it.

[0087] Step S304: In the process of playing the multi-view video, based on the character replacement object, obtain the pose data and image data of the first object, and determine the pose data and image data as object data.

[0088] After the first object determines the character replacement object, the virtual reality device can cancel the display of the character replacement object. Therefore, in the virtual video space scene sensed by the first object, the character replacement object does not appear, but other space scenes, that is, other video objects, props, and backgrounds still exist. As can be seen from the above, if the first object attempts to perform the performance of the character corresponding to the character replacement object, it is not necessary to fully arrange the real space scene corresponding to the multi-view video. The first object can perform the performance in the sensed virtual video space scene, and the virtual reality device can capture the pose data and image data of the first object.

[0089] As shown in conjunction with FIG. 7d, FIG. 7d is a schematic diagram of a virtual reality-based object performance scene provided in an embodiment of the present application. Based on the scene shown in FIG. 7b above, the virtual reality device 21 is displaying a virtual video screen 701 that does not include the video object C in the virtual display interface. The virtual video space scene sensed by the object A at this time is shown in FIG. 7d, and the object A can think that it is located in the virtual video space scene 7000. As can be seen from the above, in the perception of the object A, the virtual video space scene 7000 is a three-dimensional spatial scene, but the object A can only see the virtual video space scene 7000 in the view corresponding to the virtual video screen 701 at this time. Taking the video object B as an example, in the virtual video screen 701, only the front of the video object B is displayed, and the object A can also only see the front of the video object B, and the object A can think that the video object B stands in front of itself and faces itself. When the object A walks, the virtual reality device 21 can always obtain the virtual video screen in the view where the object A is located, and is displayed on the virtual display interface, so that the object A can always see the virtual video space scene 7000 in the view where it is located. Therefore, in the perception of the object A, it can walk as it pleases in the virtual video space scene 700, and of course, it can also perform corresponding performances in the virtual video space scene 7000. When the object A performs, both the posture data and the image data of the object A are acquired by the virtual reality device 21.

[0090] In the process of acquiring the posture data and image data of the first object, the virtual reality device can display a replacement transparency input control for the character replacement object in the virtual display interface. Then, the virtual reality device can acquire transparency information for the input character replacement object in response to an input operation on the transparency input control, perform transparency update display for the character replacement object in the virtual display interface according to the transparency information, and display a position cursor in the virtual video space scene of the character replacement object after the transparency update.

[0091] In the virtual video space scene 7000 shown in FIG. 7d above, object A cannot see video object C. Object A can function freely, but if object A wishes to mimic the motion form of video object C, object A can input transparency information for the character replacement object, for example, 60% transparency, through the replacement transparency input control 741. At this time, the virtual reality device 21 can redisplay video object C at 60% transparency in the virtual display interface 23 and can display a position cursor in the virtual display interface. The position cursor is used to notify object A of the position of video object C in the virtual video space scene 7000. For ease of understanding, as also referenced in FIG. 7e, FIG. 7e is a schematic diagram of a scene of transparent display of a virtual reality-based object provided in an embodiment of the present application. As shown in FIG. 7e, after the virtual reality device 21 obtains transparency information in response to an input operation on the replacement transparency input control 741, in the virtual display interface 23, it can switch and display the virtual video screen 701 that does not include video object C to the virtual video screen 703 that includes video object C with a transparency of 60%, and a position cursor 74 is displayed on the virtual video screen 703. At this time, object A can see that in the perceived virtual video space scene 7000, a video object C with a transparency of 60% appears in the front, and a position cursor 742 is displayed at the feet of the video object C with a transparency of 60%. Object A can determine the standing position of character B corresponding to video object C in the virtual video space scene 7000 based on the position cursor 742. At the same time, object A can learn the limb form, motion, and performance rhythm of video object C according to the video object C with a transparency of 60% and perform the performance of character B. It should be noted that the video object C with a transparency of 60% cannot appear in the production video.

[0092] Through the above steps S302 to S304, the posture data and image data applied to the character-replaced object can be flexibly obtained, thereby improving the efficiency of man-machine interaction in the image dimension.

[0093] Step S305: Unhide the character replacement objects in the performance view video, and merge the performance objects that match the performance object data into the performance view video to obtain a production video.

[0094] The performance view video is obtained by capturing the virtual video space scene using the performance view during the process of playing the multi-view video after triggering the performance control, and the performance object data refers to the data that the object data displays in the performance view.

[0095] In the produced video, the first object includes an associated performance object, and the performance object in the produced video is displayed based on the pose data and image data in the object data.

[0096] When playing a multi-view video with a virtual reality device, it is possible to view virtual video space scenes corresponding to the multi-view video with different views, and the first object can freely walk in the virtual video space scene according to its own preferences to adjust the viewing view of the virtual video space scene. However, for terminal devices that only have a single-view video playback function (such as mobile phones, tablets, and computers, etc.), when playing a certain video, only the video screen in a certain view of the corresponding real space scene can be displayed at any time. Therefore, when playing a multi-view video with a terminal device, the terminal device can also only display the video screen in the main lens view of the real space scene corresponding to the target multi-view. The main lens view may also be called the director view, that is, the original movie focus lens of the director. When the first object performs scene production for the multi-view video in the virtual video space scene, it is also possible to set the single view when the production video is played on the terminal device, that is, the performance view. When the multi-view video is not being played after the first object triggers the performance control, the virtual reality device can display the virtual camera control on the virtual display interface. The virtual reality device can set up the virtual camera in the performance view of the virtual video space scene in response to the setup operation for the virtual camera control by the first object. The virtual camera can be used to output the corresponding video screen in the performance view of the virtual video space scene. During the process of playing the multi-view video after triggering the performance control, the virtual reality device can perform shooting and recording on the virtual video space scene with the virtual camera, and obtain the corresponding performance view video in the performance view of the virtual video space scene. Subsequently, the virtual reality device can cancel the display of the character replacement object in the performance view video.At the same time, the virtual reality device can obtain object data, that is, performance object data, which is the data to be displayed in the performance view. The performance data is used to display the performance object associated with the first object in the performance view. In the process of playing the multi-view video, the virtual camera is movable, that is, the performance view is changeable. That is, the performance view corresponding to the video screen output by the virtual camera at time A may be different from the performance view corresponding to the video screen output by the virtual camera at time B. The virtual reality device can set up at least one virtual camera in the virtual video space scene, and the corresponding performance views at the same time for each virtual camera may be different. Each virtual camera can obtain a performance view video by shooting and recording, that is to say, the virtual reality device can obtain at least one performance view video at the same time. The performance views corresponding to each performance view video may be different. For the scene where the virtual reality device shoots the virtual video space scene with the virtual camera, reference can be made to the schematic diagram of the scene shown in FIG. 11c later.

[0097] The position of the virtual camera in the virtual video space scene determines the performance view, and the method of selecting the position of the virtual camera may include lens following, on-site shooting, and free movement. Here, lens following means that the view of the virtual camera follows the director's original movie focus lens view. Therefore, the position of the virtual camera in the virtual video space scene follows the position of the director's lens in the real space scene. On-site shooting means that in the process of playing the multi-view video, the virtual camera can perform shooting and recording at a fixed position in the virtual video space scene, and the position of the virtual camera does not change. The fixed position may be selected by the first object. Free movement means that in the process of playing the multi-view video, the first object can adjust the position of the virtual camera at any time, thereby changing the shooting view.

[0098] Step S306: In response to a trigger operation on the character setup control, during the process of playing the multi-view video, obtain the pose data and image data of the first object, and determine the pose data and image data as object data.

[0099] After the virtual reality device responds to the character setup control, there is no need to perform other processing on the video object in the virtual video space scene corresponding to the multi-view video. During the process of playing the multi-view video, the pose data and image data of the first object can be directly obtained and used as object data.

[0100] Step S307: Merge the performance object that matches the performance object data into the performance view video to obtain the production video.

[0101] The performance view video is the video obtained by shooting the virtual video space scene with the virtual camera in step S305 above. The virtual reality device does not need to process other video objects in the performance view video, and only needs to fuse the performance object with performance object data into the performance view video.

[0102] As an example, in the process of obtaining the pose data and image data of the first object, the virtual reality device can display a mirror preview control in the virtual display interface. In response to a trigger operation on the mirror preview control, the virtual reality device can display a performance preview area in the virtual display interface to display a performance virtual video screen. The performance virtual video screen includes a performance object fused into the virtual video space scene. According to the embodiments of the present application, the first object can adjust its own performance by viewing the performance virtual video screen, avoiding multiple changes due to performance mistakes of the first object, and improving the man-machine interaction efficiency.

[0103] As shown in conjunction with FIG. 7f, FIG. 7f is a schematic diagram of a virtual reality-based mirror preview scene provided in an embodiment of the present application. Based on the virtual reality-based scene shown in FIG. 7d above, in a virtual video space scene 7000 that does not include video object C, object A can perform a performance by playing a character corresponding to video object C. Assuming that object A walks beside video object B at the position of object A in the virtual video space scene 7000 of the previous video object C, the virtual video space scene 7000 in the view where object A is located is displayed by a virtual video screen 704 displayed in a virtual display interface. As shown in FIG. 7f, at this time, object A faces forward and cannot see video object B. Object A cannot even see its own performance in its virtual video space scene 2000, and it is not known whether the generated production video can achieve the expected effect. Therefore, in the process of acquiring the pose data and image data of object A, virtual reality device 21 may display a mirror preview control 75 in virtual display interface 23, and when object A triggers mirror preview control 75 with virtual reality handle 22, virtual reality device 21 can display a performance virtual video screen 705 in performance preview area 76 in virtual display interface 23. Here, performance virtual video screen 705 is a mirror screen in the view where object A is located of virtual video space scene 7000 that does not include video object C and object A, and performance object A is generated based on the pose data and image data of object A. As an example, the virtual reality device may further display a performance virtual video screen 706 in comparison area 77, and performance virtual video screen 706 is a mirror screen in the view where object A is located of virtual video space scene 7000 that includes video object C.Object A can adjust its own performance by viewing the performance virtual video screen 705 and the performance virtual video screen 706 .

[0104] For example, the virtual reality device may display an image customization list on a virtual display interface, and, in response to completion of a configuration operation on the image customization list, update the image data according to the configured image data to obtain configured image data. The configured image data may include clothing data, body type data, voice data, and appearance data. The virtual reality device may then display a performance object with a performance movement and a performance image in a production video. Here, the performance movement is determined based on the posture data of the first object, and the performance image is determined based on at least one of the clothing data, body type data, voice data, and appearance data. The present embodiment enables the first object to customize the image of a performance object corresponding to a target performance character in a multi-view video, thereby improving the degree of simulation of the performance object, efficiently improving the success rate of production, and thereby improving man-machine interaction efficiency.

[0105] The clothing data is used to display the appearance of the performance object. For example, the performance object displayed according to the clothing data can wear a short-sleeved shirt, a T-shirt, trousers, or a one-piece dress. The body shape data is used to display the body shape of the performance object. For example, the performance object displayed according to the body shape data may be described as having a large head and a small body, a small head and a large body, a tall and thin body, or a short and fat body. The voice data is used to display the voice of the performance object. For example, the performance object displayed according to the voice data may have a child's voice or a young person's voice. The facial appearance data is used to display the facial appearance of the performance object. The first object can customize the performance image in the production video of the performance object associated with itself.

[0106] The image customization list includes a first image customization list and a second image customization list. The first image customization list may include a character image, an object image, and a custom image. The character image is an image of a video object corresponding to a character, the object image is an image of the first object, and the custom image is some general-purpose image provided by the virtual reality device. The second image customization list may include an object image and a custom image. When the first object performs as a target performance character appearing in the multi-view video, the first object may select a performance image of the performance object that completely or partially replaces the image of the video object corresponding to the target performance character according to the first image customization list, for example, the appearance, the body type, the appearance, and the voice. Here, the target performance character may be a character that appeared at the target playback time in the multi-view video that the first object is to replace when the first object performs in the virtual video space scene corresponding to the target playback time of the multi-view video, or a character that does not appear at the target playback time in the multi-view video that the first object is to set up. When the first object performs as a newly added character, i.e., a character that has never appeared in the multi-view video, the first object can also customize the appearance, body type, appearance, and voice of the performance object corresponding to the newly added character through the image customization list, but there is no option for the character image.

[0107] To facilitate understanding of the process of customizing an image for a performance object corresponding to a target performance character in a multi-view video, the first object will be described as an example in which the target performance character is Character B. Referring also to FIG. 7g, FIG. 7g is a schematic diagram of a first image customization list provided in an embodiment of the present application. In the scene of selecting a replaceable video object shown in FIG. 7b above, after the virtual reality device 21 determines that video object C corresponding to Character B is the character replacement object, and before the virtual reality device 21 switches the virtual video screen 700 including video object C to a virtual video screen 701 not including video object C, the virtual reality device 21 can first switch the virtual video screen 700 to display an image customization interface 707 for Character B. As shown in FIG. 7g, the image customization interface 707 includes a first image customization list 78 and an image preview area 79. Here, a preview performance object 710 is displayed in the image preview area 79, and the initial image of the preview performance object 710 may match the image of video object C or the image of object A. Thereafter, object A can adjust the image of preview performance object 710 based on first image customization list 78. When object A completes the configuration operation on first image customization list 78, the image of preview performance object 710 becomes the performance image of the performance object corresponding to character B in the produced video. As shown in FIG. 7g, first image customization list 78 includes appearance control 781, body type control 782, appearance control 783, and voice control 784. After object A triggers appearance control 781 with virtual reality handle 22, object A may select "character outfit."At this time, the outfit of the preview performance object 710 is displayed as the outfit of the video object C corresponding to character B, and object A may select "actual human outfit." At this time, the outfit of the preview performance object 710 is displayed as the outfit worn by object A himself. Object A may further select "custom" and select various pre-set outfits. At this time, the outfit of the preview performance object 710 is displayed as the pre-set outfit selected by object A. Object A may freely combine outfits, i.e., some outfits may select "character outfit," some outfits may select "actual human outfit," and some outfits may select "custom." After object A triggers the outfit control 782 with the virtual reality handle 22, object A may select "character body type." At this time, the body type of the preview performance object 710 is displayed as the body type of video object C, and object A may select "actual human body type." At this time, the body type of the preview performance object 710 is displayed as the body type worn by object A himself. Object A may further select "custom" and select various pre-set outfits. At this time, the body type of the preview performance object 710 is displayed as the preset body type selected by object A. Object A can perform local or global deformation and height adjustment on the selected body type, and the virtual reality device 21 may further provide a recommended body type suitable for character B. After object A triggers the appearance control 783 with the virtual reality handle 22, object A may select "Character Appearance." At this time, the appearance of the preview performance object 710 is displayed as the facial features of video object C. Object A may also select "Real Human Appearance," and at this time, the appearance of the preview performance object 710 is displayed to maintain the facial features of object A himself.Object A may further select "Custom" and select and combine various preset features such as "face shape, eyes, nose, mouth, and ears". At this time, the appearance of the preview performance object 710 is displayed as the combined face features selected by Object A. For the "actual human appearance" and the appearance selected as "Custom", Object A can further perform adjustments such as local deformation, color, gloss, and makeup. After Object A triggers the voice control 784 by the virtual reality handle 22, Object A may select a "character voice". At this time, the voice feature of the preview performance object 710 is the same as the voice feature of the video object C, and Object A may select an "actual human voice". At this time, the voice feature of the preview performance object 710 is the same as the voice feature of Object A himself. Object A may further select "voice conversion" and select various preset voice conversion types to perform voice conversion, and the voice feature of the preview performance object 710 will change to the selected voice feature.

[0108] To facilitate the understanding of the process of customizing an image for a performance object corresponding to a newly added character in a multi-view video, an example where the newly added character is Ding will be described. As referred to in conjunction with FIG. 7h, FIG. 7h is a schematic diagram of a second image customization list provided in an embodiment of the present application. Based on the scene shown in FIG. 7a above, object A is attempting to newly add character Ding in the multi-view video. After the virtual reality device 21 responds to a trigger operation on the character setup control 712 by object A, as shown in FIG. 7h, the virtual reality device 21 can display an image customization interface 708 for character Ding in the virtual display interface. As shown in FIG. 7g, the image customization interface 708 includes a second image customization list 711 and an image preview area 712. Here, a preview performance object 713 is displayed in the image preview area 712, and the initial image of the preview performance object 713 may match the image of object A. Subsequently, object A can adjust the image of the preview performance object 713 based on the second image customization list 711. As shown in FIG. 7h, the second image customization list 711 includes an outfit control 7111, a body type control 7112, a facial features control 7113, and a voice control 7114. The difference between the second image customization list 711 and the first image customization list 78 is that in the options corresponding to each control, there are no options related to character images, and the composition of the remaining options is the same as that of the first image customization list 78, which will not be described in detail again here.

[0109] As an example, the virtual reality device may display shopping controls in the virtual display interface and, in response to a trigger operation on the shopping controls, display virtual items that can be purchased according to the second display method. Here, the second display method is different from the display method of the virtual items that can be purchased before triggering the shopping controls. The virtual items that can be purchased belong to the items displayed in the virtual video space scene. Subsequently, the virtual reality device can, in response to a selection operation on the virtual item that can be purchased, set the selected virtual item that can be purchased as a purchased item and display purchase information corresponding to the purchased item in the virtual display interface. According to the embodiments of the present application, the virtual items that can be purchased can be highlighted, thereby playing a notification role in the purchase process and improving the man-machine interaction efficiency.

[0110] As shown in conjunction with FIG. 7i, FIG. 7i is a schematic diagram of a scene for displaying virtual items that can be purchased provided in the embodiments of the present application. As shown in FIG. 7i, on the premise of the scene shown in FIG. 2a above, the virtual reality device 21 may further independently display a shopping control 714 in the virtual display interface 23. When the object A triggers the shopping control 714 by the virtual reality handle 22, the virtual reality device 21 can highlight the virtual items that can be purchased (for example, surround it with a dashed box). As shown in FIG. 7i, the virtual items that can be purchased include a virtual hat 715 and the like. The object A can select the virtual item that can be purchased to be understood. Assuming that the object A selects the virtual hat 715, the virtual reality device 21 can display the purchase information 716 and inform the object A of the price of the real hat corresponding to the virtual hat 715 and the purchase method and the like.

[0111] By adopting the method provided in the embodiments of the present application, the first object can play a character in the multi-view video in a virtual video space scene corresponding to the multi-view video, or add a new character to obtain a produced video, thereby enriching the display method of the multi-view video.

[0112] For a better understanding of the process of the virtual reality device acquiring object data of a first object in a virtual video space scene and then acquiring object data of a second object when responding to a scene editing operation for an object invitation, as described in FIG. 3 above, and then obtaining a produced video according to the object data of the first object and the object data of the second object, please refer to FIG. 8. FIG. 8 is a flowchart of a data processing method for performing multi-object video production based on virtual reality provided in an embodiment of the present application. The data processing method may be performed by a virtual reality device. For ease of understanding, the embodiment of the present application will be described using an example in which the method is performed by a virtual reality device. The data processing method may include at least the following steps S401 to S403.

[0113] Step S401: In response to a trigger operation on the multi-view video, a virtual video space scene corresponding to the multi-view video is displayed.

[0114] To implement step S401, reference can be made to step S101 in the embodiment corresponding to FIG.

[0115] Step S402: Display an object invite control in the virtual display interface, display an object list in response to a trigger operation on the object invite control, and send an invite request to a target virtual reality device associated with a second object in response to a selection operation on the object list, thereby causing the target virtual reality device associated with the second object to display the virtual video space scene. The object list includes objects having an association relationship with the first object. According to the embodiment of the present application, the second object can be invited to enter the virtual video space scene, thereby improving the interactive efficiency.

[0116] The first object can choose to invite at least one second object to join in creating a scene for the multi-view video, and a low-latency always-on network is established between the virtual reality devices of the first object and each of the second objects.

[0117] One possible implementation process for the virtual reality device to respond to a selection operation on the object list and send an invitation request to the target virtual reality device associated with the second object, causing the target virtual reality device associated with the second object to display the virtual video space scene is as follows. In response to a selection operation on the object list, start an invitation request for the second object to the server, so that the server sends an invitation request to the target virtual reality device associated with the second object. Here, when the target virtual reality device receives the invitation request, it displays the virtual video space scene on the target virtual reality device. When the target virtual reality device receives the invitation request and displays the virtual video space scene, a target virtual object is displayed on the object virtual video screen. Here, the second object enters the virtual video space scene by the target virtual object, the target virtual object is associated with the image data of the second object, and the object virtual video screen is used to display the virtual video space scene in the view where the first object is located.

[0118] As shown in conjunction with FIG. 9a, FIG. 9a is a schematic diagram of a virtual reality-based object invitation scene provided in an embodiment of the present application. Assuming that, on the premise of the scene shown in FIG. 2a above, object A selects the object invitation control 243 by the virtual reality handle 22, the virtual reality device 21 can display the object list 91 on the virtual display interface 23 in response to the trigger operation on the object invitation control 243. In the object list 91, objects having a friendship relationship with object A, such as object aaa and object aab, are included, and object A can select the second object to be invited, such as object aaa, by the virtual reality handle 22. The virtual reality device 21 can start an invitation request for object aaa to the server in response to the selection operation on object aaa.

[0119] Referring to FIG. 9b, FIG. 9b is a schematic diagram of a second object display scene based on virtual reality provided in an embodiment of the present application. As shown in FIG. 9b, object aaa may wear a target virtual reality device and a target virtual reality handle, and object aaa may accept an invitation request from object A through the target virtual reality handle. In response to the acceptance operation for the invitation request from object aaa, the target virtual reality device may obtain image data of object aaa and subsequently enter a virtual video space scene 2000 corresponding to the multi-view video, which corresponds to displaying the virtual video space scene 2000 corresponding to the multi-view video on the target virtual reality device. The target virtual reality device may share the image data of object aaa with the virtual reality device 21. Therefore, the virtual reality device 21 may generate a target virtual object 92 identical to the image of object aaa according to the image data of object aaa. Then, the virtual reality device 21 may display the target virtual object 92 in a view where object A is located on the virtual video screen 201. The target virtual reality device also acquires posture data, voice data, etc. of object aaa in real time, which are used to display the movement, facial expression, and voice of object aaa, as target object data. The acquired target object data can then be shared in real time with virtual reality device 21. Virtual reality device 21 can simulate and display the movement, facial expression, and voice of object aaa in real time using target virtual object 92 in accordance with the target object data. It should be understood that virtual reality device 21 can also acquire object data of object A and share it with the target virtual reality device, so that the target virtual reality device displays the virtual reality object associated with object A in the virtual display interface and simulates and displays the movement, facial expression, and voice of object A using the virtual reality object associated with object A.

[0120] As an example, the first object and the second object can perform an instant session in a manner such as voice and text by means of virtual reality objects associated with each in a virtual video space scene. When the second object speaks, the virtual reality device having a binding relationship with the second object can acquire the instant voice data of the second object and subsequently share it with the virtual reality device. Subsequently, the virtual reality device can reproduce the voice of the second object according to the instant voice data of the second object. In addition, the virtual reality device can further display a session message corresponding to the target virtual object on the object virtual video screen. Here, the session message is generated based on the instant voice data of the second object.

[0121] Step S403: When the second object has already triggered performance control, in response to the trigger operation on the performance control in the virtual display interface by the first object, display the target object data corresponding to the target virtual object on the object virtual video screen, and at the same time, the first object acquires the object data of the object in the virtual video space scene.

[0122] The target virtual object is associated with the image data of the second object.

[0123] In a virtual video space scene corresponding to multi-view video, the first object and the second object can perform performances simultaneously. The virtual reality device can acquire the object data of the first object, and the virtual reality device having a binding relationship with the second object can acquire the target object data of the second object, and the virtual reality device having a binding relationship with the second object can share the target object data with the virtual reality device. The acquisition of the object data and the target object data can refer to the descriptions in the embodiments corresponding to FIGS. 4 and 6 above, and will not be described in detail again here.

[0124] Step S404: Play the production video associated with the multi-view video in the virtual display interface. The production video includes a performance object associated with the first object and a target virtual object. The performance object in the production video is displayed based on the object data, and the target virtual object in the production video is displayed based on the target object data.

[0125] The virtual reality device having a binding relationship with the first object can obtain a production video obtained by cooperation of multiple objects by fusing the object data, the target object data, and the virtual video space scene with each other. The virtual reality device having a binding relationship with the first object can obtain a production video by shooting the virtual video space scene by a cooperative performance view in the process of obtaining the object data of the first object in the virtual video space scene. Then, the performance object having the performance object data and the target virtual object having the cooperative performance object data are fused into the cooperative performance video to obtain a production video. Here, the performance object data refers to data displayed by the object data in the cooperative performance view, and the cooperative performance object data refers to data displayed by the target object data in the performance view.

[0126] By adopting the method provided in the embodiments of the present application, a first object can invite a second object to perform scene creation in the same virtual video space scene, further enriching the display and interaction methods of multi-view video.

[0127] Referring to FIG. 10, FIG. 10 is a flowchart of a data processing method for performing video recording based on virtual reality provided in an embodiment of the present application. The data processing method may be executed by a virtual reality device. For ease of understanding, the embodiment of the present application will be described taking the method as being executed by a virtual reality device as an example. The data processing method may include at least the following steps S501 to S503.

[0128] Step S501: In response to a trigger operation on the multi-view video, display a virtual video space scene corresponding to the multi-view video.

[0129] When a virtual reality device displays a virtual video space scene corresponding to a multi-view video, the virtual reality device can default to the main lens virtual video screen in the virtual display interface. At this time, the first object wearing the virtual reality device senses the virtual video space scene corresponding to the multi-view video and views the virtual video space scene in the main lens view. The virtual reality device can use the view where the first object is located at this time as the main lens view. The first object can switch views at any time by the virtual reality device to view the virtual video space scene from different views.

[0130] As an example, the virtual reality device may display a moving view switching control in the virtual display interface. Then, in response to a trigger operation on the moving view switching control, the virtual reality device can obtain a view of the virtual video space scene of the first object after the movement as a moving view, and then switch and display the main lens virtual video screen to a moving virtual video screen in the moving view of the virtual video space scene. That is, after the first object triggers the moving view switching control, the first object can walk at will in the sensed virtual video space scene and view the virtual video space scene corresponding to the multi-view video in 360 degrees. For ease of understanding, as also referred to in FIG. 11a, FIG. 11a is a schematic diagram of a moving view switching scene based on virtual reality provided in an embodiment of the present application. As shown in FIG. 11a, assuming that an object A senses a virtual video space scene 1100 by a virtual reality device 1101, the virtual video space scene 1100 includes a video object G. At this time, the virtual reality device 1101 can display a virtual video screen 1104 in a virtual display interface 1103, which can be used to display the virtual video space scene 1100 in a view where object A is located, for example, object A can see the front of video object G. The virtual reality device 1101 can also display a move view switching control 1105 in the virtual display interface 1103. When object A wants to change the viewing view of the virtual video space scene 1103 by walking, object A can trigger the move view switching control 1105 using the virtual reality handle 1102, followed by object A walking. The virtual reality device 21 can obtain the view of the virtual video space scene 1100 of the first object after movement as a move view, and then obtain a virtual video screen used to display the virtual video space scene 1100 in the move view.For example, in the virtual video space scene 1100, object A walks from in front of video object G to behind video object G. As shown in FIG. 11a, at this time, the virtual reality device 1101 can display the virtual video screen 1106 on the virtual display interface 1103. As can be seen from this, at this time, object A can only see the back of video object G.

[0131] As an example, the virtual reality device may display a pointing view switching control in the virtual display interface. The virtual reality device may display a pointing cursor in the virtual display interface in response to a trigger operation on the pointing view switching control. The virtual reality device may obtain a view of the virtual video space scene of the moved pointing cursor as a pointing view in response to a movement operation on the pointing cursor, and subsequently, switch and display the main lens virtual video screen to a pointing virtual video screen in the pointing view of the virtual video space scene. That is, the first object may adjust the viewing view of the virtual video space scene by the pointing cursor without walking. For ease of understanding, as referred to in conjunction with FIG. 11b, FIG. 11b is a schematic diagram of a scene of virtual reality-based pointing view switching provided in an embodiment of the present application. In the scene of FIG. 11a described above, the virtual reality device 1101 may further display a pointing view switching control 1107 in the virtual display interface 1103. After the virtual reality device 1101 responds to a trigger operation on the pointing view switching control 1107, the object A can see a pointing cursor 1108 in the virtual video space scene. Assuming that the object A moves the pointing cursor 1108 behind the video object G by the virtual reality handle 1102, the virtual reality device 1101 obtains a view of the virtual video space scene of the pointing cursor 1108 as a pointing view, and subsequently, can obtain a virtual video screen 1109 used to display the virtual video space scene 1100 in the pointing view. As shown in FIG. 11b, although the position of the object A does not change, the position of the virtual video space scene perceived by the object A changes, and the object A sees the virtual video space scene 1100 in the pointing view.

[0132] Step S502: Display shooting and recording control in the virtual display interface, perform shooting and recording on the virtual video space scene in response to a trigger operation on the shooting and recording control, and obtain a recorded video.

[0133] One possible implementation process for a virtual reality device to capture and record a virtual video space scene and obtain a recorded video in response to a trigger operation on the capture and record control may be as follows: The virtual reality device obtains a capture view of the virtual video space scene in response to a trigger operation on the capture and record control. The virtual reality device then displays a capture virtual video screen in the capture view of the virtual video space scene on the virtual display interface. Then, a recording screen frame is displayed on the capture virtual video screen, and the video screen in the recording screen frame of the capture virtual video screen is recorded to obtain a recorded video. Here, the determination of the capture view is the same as the determination of the performance view described above, and supports three determination methods: lens tracking, on-site capture, and free movement. For ease of understanding, please also refer to FIG. 11c, which is a schematic diagram of a virtual reality-based video capture scene provided in an embodiment of the present application. The virtual reality device 1101 can capture a screen or record a clip of the virtual video space scene 1100. As shown in FIG. 11c, the virtual reality device 1101 may display a shooting / recording control 1110 in the virtual display interface 1103. After the virtual reality device 1101 responds to a trigger operation on the shooting / recording control 1110, it may display a shooting virtual video screen 1111 in a shooting view of the virtual video space scene 1110 and a recording screen frame 1112. As can be seen, only the screen in the recording screen frame 1112 is recorded, and object A can adjust the size and position of the recording screen frame 1112. Object A can take a screen shot by triggering the shooting control 1113 and can record a clip by triggering the recording control 1114. Note that object A can also trigger a path selection control 1115. After responding to a trigger operation on the path selection control 1115, the virtual reality device may display a path selection list 1115, which includes three types of shooting path modes: lens tracking, on-site shooting, and free movement.As an example, the first object may add a plurality of lenses and a shooting path to record the screen, and they do not interfere with each other.

[0134] By adopting the method provided in the embodiments of the present application, the first object selects a shooting view to perform shooting and recording on the virtual video space scene, and obtains a production video whose main lens view is the shooting view, thereby enriching the display method and interaction method of the multi-view video.

[0135] As shown in FIG. 12, FIG. 12 is a structural schematic diagram of a data processing device provided in the embodiments of the present application. The data processing device may be a computer program (including program code) operated on computer equipment. For example, the data processing device may be an application software, and the device can be used to execute corresponding steps in the data processing method provided in the embodiments of the present application. As shown in FIG. 12, the data processing device 1 may include a first response module 101, a second response module 102, and a video playback module 103. The first response module 101 is configured to display a virtual video space scene corresponding to the multi-view video in response to a trigger operation on the multi-view video. The second response module 102 is configured to obtain object data of the first object in the virtual video space scene in response to a scene editing operation on the virtual video space scene. Here, the first object refers to the object that starts the trigger operation on the multi-view video. The video playback module 103 is configured to play a production video associated with the multi-view video on the virtual display interface. Here, the production video is obtained by performing editing processing on the virtual video space scene based on the object data.

[0136] The specific implementation manners of the first response module 101, the second response module 102, and the video playback module 103 can refer to the descriptions of steps S101 to S103 in the embodiment corresponding to FIG. 3 above, and will not be described in detail again here.

[0137] In some embodiments, the scene editing operation includes a trigger operation on the voice-over control in the virtual display interface. As referred to again in FIG. 12, the second response module 102 includes a first response unit 1021. The first response unit 1021 plays a multi-view video and, in response to a trigger operation on the voice-over control in the virtual display interface, acquires the audio data of the first object in the virtual video space scene and is configured to determine the audio data of the first object as the object data applied to the multi-view video.

[0138] The specific implementation manner of the first response unit 1021 can refer to the description of step S102 in the embodiment corresponding to FIG. 3 above, and will not be described in detail again here.

[0139] In some embodiments, the audio data of the first object includes object audio data or background audio data. Referring again to FIG. 12 , the first response unit 1021 includes a first response subunit 10211, a first selection subunit 10212, a first determination subunit 10213, and a second determination subunit 10214. The first response subunit 10211 is configured to display a sound-injection mode list, where the sound-injection mode list includes an object sound-injection control and a background sound-injection control. The first selection subunit 10212 is configured to display sound-injectable video objects in response to a selection operation on the object sound-injection control, and to designate the selected sound-injectable video object as a sound-injection target object in response to a selection operation on the sound-injectable video object. Here, the sound-injectable video object belongs to a video object displayed in the multi-view video. The first determination subunit 10213 is configured to obtain object audio data of the first object based on the object to be added sound during the process of playing the multi-view video, and determine the object audio data as object data; and the second determination subunit 10214 is configured to obtain background audio data of the first object during the process of playing the multi-view video, and determine the background audio data as object data in response to a selection operation on the background audio addition control.

[0140] For the specific implementation methods of the first response subunit 10211, the first selection subunit 10212, the first determination subunit 10213, and the second determination subunit 10214, please refer to the description of steps S201 to S207 in the embodiment corresponding to Figure 4 above, and will not be described in detail again here.

[0141] In some embodiments, the first determination subunit 10213 further performs a sound cancellation process on the video and audio data corresponding to the object to be dubbed. When the object to be dubbed is in a speaking state, the text information and the sound track information corresponding to the video and audio data are displayed, and the object audio data of the first object is acquired and configured to determine the object audio data as object data.

[0142] In some embodiments, as again referred to in FIG. 12, the above data processing apparatus 1 further includes an audio replacement module 104 and an audio overlay module 105. The audio replacement module 104 is configured to perform a replacement process on the video and audio data corresponding to the object to be dubbed in the multi-view video by using the object audio data when the object data is object audio data, so as to obtain a production video. The audio overlay module 105 is configured to overlay the background audio data on the multi-view video to obtain a production video when the object data is background audio data.

[0143] For the specific implementation manners of the audio replacement module 104 and the audio overlay module 105, reference may be made to the descriptions of S201 to step S207 in the embodiment corresponding to FIG. 4 above, and details are not described again here.

[0144] In some embodiments, the scene editing operation includes a trigger operation on a performance control in the virtual display interface. Referring again to FIG. 12 , the second response module 102 includes a second response unit 1022. The second response unit 1022 is configured to play the multi-view video and, in response to the trigger operation on the performance control in the virtual display interface, acquire pose data and image data of a first object in the virtual video space scene and determine the pose data and image data as object data to be applied to the multi-view video. Here, the production video includes a performance object associated with the first object, and the performance object in the production video is to be displayed based on the pose data and image data.

[0145] For the specific implementation of the second response unit 1022, please refer to the description of step S102 in the embodiment corresponding to FIG. 3 above, and will not be described in detail again here.

[0146] In some embodiments, as again referred to FIG. 12, the second response unit 1022 includes a second response subunit 10221, a second selection subunit 10222, a third determination subunit 10223, and a fourth determination subunit 10224. The second response subunit 10221 is configured to display a performance mode list. Here, the performance mode list includes a character replacement control and a character set-up control. The second selection subunit 10222 is configured to display a replaceable video object in response to a trigger operation on the character replacement control, and to set the selected replaceable video object as a character replacement object in response to a selection operation on the replaceable video object. Here, the replaceable video object belongs to the video object displayed in the multi-view video. The third determination subunit 10223 is configured to obtain the pose data and the image data of the first object based on the character replacement object and determine the pose data and the image data as object data in the process of playing the multi-view video. The fourth determination subunit 10224 is configured to obtain the pose data and the image data of the first object and determine the pose data and the image data as object data in the process of playing the multi-view video in response to a trigger operation on the character set-up control.

[0147] For the specific implementation manners of the second response subunit 10221, the second selection subunit 10222, the third determination subunit 10223, and the fourth determination subunit 10224, reference may be made to the descriptions of steps S301 to S307 in the embodiment corresponding to FIG. 6 above, and details are not described herein again.

[0148] In some embodiments, referring again to FIG. 12 , the data processing device 1 further includes a performance substitution module 106 and a performance blending module 107. The performance substitution module 106 is configured to undisplay a character substitution object in the performance view video when object data is acquired by triggering a character substitution control, and blend a performance object matching the performance object data into the performance view video to obtain a production video. Here, the performance view video is obtained by capturing the virtual video space scene using the performance view during the process of playing the multi-view video after triggering the performance control. The performance object data refers to data displayed by the object data in the performance view. The performance blending module 107 is configured to blend a performance object matching the performance object data into the performance view video to obtain a production video when object data is acquired by triggering a character setup control.

[0149] For specific implementation methods of the performance substitution module 106 and the performance fusion module 107, please refer to the description of steps S301 to S307 in the embodiment corresponding to FIG. 6 above, and will not be described in detail again here.

[0150] In some embodiments, the second selection subunit 10222 is further configured to determine a video object currently displayed on the object virtual video screen as a replaceable video object in response to a trigger operation on the character replacement control, and to display the marked replaceable video object according to a first display manner in response to a marking operation on the replaceable video object, where the first display manner is different from the display manner of other video objects other than the replaceable video object, and the marked replaceable video object is the character replacement object, where the object virtual video screen is used to display the virtual video space scene in a view where the first object is located.

[0151] In some embodiments, the second selection subunit 10222 is further configured to: display at least one video clip corresponding to the multi-view video in response to a trigger operation on the character replacement control; display video objects included in the selected video clip in response to a selection operation on the at least one video clip; determine the video objects included in the selected video clip as replaceable video objects; and display the marked replaceable video objects according to a first display manner in response to a marking operation on the replaceable video objects, where the first display manner is different from the display manner of other video objects other than the replaceable video objects, and determine the marked replaceable video objects as character replacement objects.

[0152] In some embodiments, as again referred to FIG. 12, the above data processing apparatus 1 further includes a mirror image display module 108. The mirror image display module 108 is used to display a mirror preview control on a virtual display interface in the process of obtaining the pose data of the first object and the image data. The mirror image display module 108 is further configured to display a performance virtual video screen in a performance preview area on the virtual display interface in response to a trigger operation on the mirror preview control. Here, the performance virtual video screen includes a performance object fused into a virtual video space scene.

[0153] For the specific implementation manner of the mirror image display module 108, reference can be made to the description of the selectable embodiments in the embodiments corresponding to FIG. 6 above, and details are not described herein again.

[0154] In some embodiments, as again referred to FIG. 12, the above data processing apparatus 1 further includes an image customization module 109. The image customization module 109 is configured to display an image customization list on a virtual display interface. The image customization module 109 is further configured to update the image data according to the configured image data and obtain configured image data in response to the completion of a configuration operation on the image customization list. Here, the configured image data includes clothing data, body type data, and facial data. The image customization module 109 is further configured to display a performance object by performance actions and performance images in a production video. Here, the performance actions are determined based on the pose data of the first object, and the performance images are determined based on at least one of clothing data, body type data, voice data, and facial data.

[0155] For a specific implementation of the image customization module 109, please refer to the description of the alternative embodiment in the embodiment corresponding to FIG. 6 above, and it will not be described in detail again here.

[0156] 12 , in some embodiments, the above-mentioned data processing device 1 further includes a transparent display module 110. The transparent display module 110 is configured to display a replacement transparency input control for the character replacement object in the virtual display interface. The transparent display module 110 is further configured to, in response to an input operation on the transparency input control, obtain transparency information for the input character replacement object, update the transparency display of the character replacement object in the virtual display interface according to the transparency information, and display a position cursor in the virtual video space scene of the character replacement object after the transparency update.

[0157] For a specific implementation of the transparent display module 110, please refer to the description of the optional embodiment of step S304 in the embodiment corresponding to FIG. 6 above, and will not be described in detail again here.

[0158] 12 , in some embodiments, the above-mentioned data processing device 1 further includes an object invite module 111 and a third response module 112. The object invite module 111 is configured to display an object invite control in the virtual display interface, and to display an object list in response to a trigger operation on the object invite control, where the object list includes objects having an association relationship with the first object, and the third response module 112 is configured to send an invite request to a target virtual reality device associated with the second object in response to a selection operation on the object list, thereby causing the target virtual reality device associated with the second object to display the virtual video space scene.

[0159] For the specific implementation manners of the object invitation module 111 and the third response module 112, reference may be made to the description of step S402 in the embodiment corresponding to FIG. 8 above, and details will not be described again here.

[0160] In some embodiments, as again referred to FIG. 12, the third response module 112 includes an invitation unit 1121 and a display unit 1122. The invitation unit 1121 starts an invitation request for the second object to the server, so that the server is configured to send the invitation request to the target virtual reality device associated with the second object. Here, when the target virtual reality device receives the invitation request, the target virtual reality device displays a virtual video space scene. The display unit 1122 is configured to display a target virtual object on the object virtual video screen when the target virtual reality device receives the invitation request and displays the virtual video space scene. Here, the second object enters the virtual video space scene by the target virtual object. The target virtual object is associated with the image data of the second object, and the object virtual video screen is used to display the virtual video space scene in the view where the first object is located.

[0161] For the specific implementation manners of the invitation unit 1121 and the display unit 1122, reference may be made to the description of step S302 in the embodiment corresponding to FIG. 8 above, and details will not be described again here.

[0162] In some embodiments, when the second object has already triggered performance control, the scene editing operation includes a trigger operation for the performance control in the virtual display interface by the first object. As again referred to in FIG. 12, the second response module 102 includes a determination unit 1023. The determination unit 1023 is configured to display target object data corresponding to the target virtual object on the object virtual video screen in response to a trigger operation for the performance control in the virtual display interface by the first object, and to obtain object data of the first object in the virtual video space scene. Here, in the production video, a performance object associated with the first object and a target virtual object are included. The performance object in the production video is displayed based on the object data, and the target virtual object in the production video is displayed based on the target object data.

[0163] For the specific implementation manner of the determination unit 1023, reference may be made to the description of step S403 in the embodiment corresponding to FIG. 8 above, and details are not described again here.

[0164] 12 , in some embodiments, the above-mentioned data processing device 1 further includes a photographing module 113 and a fusing module 114. The photographing module 113 is configured to photograph the virtual video space scene using a collaborative performance view in the process of acquiring object data of a first object in the virtual video space scene to obtain a collaborative performance video. The fusing module 114 is configured to fuse the performance object having the performance object data and the target virtual object having the collaborative performance object data into the collaborative performance video to obtain a produced video. Here, the performance object data refers to the data displayed by the object data in the collaborative performance view, and the collaborative performance object data refers to the data displayed by the target object data in the performance view.

[0165] For the specific implementation of the photographing module 113 and the fusion module 114, please refer to the description of step S403 in the embodiment corresponding to FIG. 8 above, and will not be described in detail again here.

[0166] 12, in some embodiments, the above-mentioned data processing device 1 further includes a session display module 115. The session display module 115 is configured to display a session message corresponding to the target virtual object on the object virtual video screen, where the session message is generated based on the instant audio data of the second object.

[0167] Here, for the specific implementation of the session display module 115, please refer to the description of the optional embodiment of step S403 in the embodiment corresponding to FIG. 8 above, and will not be described in detail again here.

[0168] In some embodiments, as again referred to FIG. 12, the above data processing apparatus 1 further includes a shopping module 116. The shopping module 116 is configured to display shopping controls in a virtual display interface and, in response to a trigger operation on the shopping controls, display purchasable virtual items according to a second display mode. Here, the second display mode is different from the display mode of the purchasable virtual items before triggering the shopping controls, and the purchasable virtual items belong to the items displayed in the virtual video space scene. The shopping module 116 is configured to, in response to a selection operation on a purchasable virtual item, set the selected purchasable virtual item as a purchased item and display purchase information corresponding to the purchased item in the virtual display interface.

[0169] For the specific implementation manner of the shopping module 116, reference can be made to the description of the selectable embodiments in the embodiments corresponding to FIG. 6 above, and details are not described again here.

[0170] In some embodiments, as again referred to FIG. 12, the above data processing apparatus 1 further includes a first screen display module 117 and a first view switching module 118. The first screen display module 117 is configured to display a main lens virtual video screen in the main lens view of the virtual video space scene in the virtual display interface. The first view switching module 118 is configured to display a movement view switching control in the virtual display interface. The first view switching module 118 is further configured to, in response to a trigger operation on the movement view switching control, obtain a view of the virtual video space scene of the first object after movement as a movement view. The first view switching module 118 is further used to switch and display the main lens virtual video screen to a movement virtual video screen in the movement view of the virtual video space scene.

[0171] For the specific implementation manners of the first screen display module 117 and the first view switching module 118, reference can be made to the description of step S501 in the embodiment corresponding to FIG. 10 above, and details will not be described again here.

[0172] In some embodiments, as again referred to FIG. 12, the above data processing apparatus 1 further includes a second screen display module 119 and a second view switching module 120. The second screen display module 119 is configured to display a main lens virtual video screen in the main lens view of the virtual video space scene on the virtual display interface. The second view switching module 120 is configured to display a pointing view switching control on the virtual display interface. The second view switching module 120 is further configured to display a pointing cursor on the virtual display interface in response to a trigger operation on the pointing view switching control. The second view switching module 120 is further configured to obtain a view of the virtual video space scene of the moved pointing cursor as a pointing view in response to a movement operation on the pointing cursor. The second view switching module 120 is further configured to switch and display the main lens virtual video screen to a pointing virtual video screen in the pointing view of the virtual video space scene.

[0173] For the specific implementation manners of the second screen display module 119 and the second view switching module 120, reference can be made to the description of step S501 in the embodiment corresponding to FIG. 10 above, and details will not be described again here.

[0174] In some embodiments, as again referring to FIG. 12, the above data processing apparatus 1 further includes a first display module 121 and a fourth response module 122. The first display module 121 is configured to display shooting and recording control in a virtual display interface. The fourth response module 122 is configured to perform shooting and recording on a virtual video space scene in response to a trigger operation on the shooting and recording control to obtain a recorded video.

[0175] For the specific implementation manners of the first display module 121 and the fourth response module 122, reference can be made to the description of step S502 in the embodiment corresponding to FIG. 10 above, and details are not described again here.

[0176] In some embodiments, as again referring to FIG. 12, the fourth response module 122 includes a view acquisition unit 1221, a first display unit 1222, a second display unit 1223, and a recording unit 1224. The view acquisition unit 1221 is configured to acquire a shooting view of a virtual video space scene in response to a trigger operation on the shooting and recording control. The first display unit 1222 is configured to display a shooting virtual video screen in the shooting view of the virtual video space scene in a virtual display interface. The second display unit 1223 is configured to display a recording screen frame in the shooting virtual video screen. The recording unit 1224 is configured to record a video screen in the recording screen frame of the shooting virtual video screen to obtain a recorded video.

[0177] For the specific implementation manners of the view acquisition unit 1221, the first display unit 1222, the second display unit 1223, and the recording unit 1224, reference can be made to the description of step S502 in the embodiment corresponding to FIG. 10 above, and details are not described again here.

[0178] In some embodiments, as again referred to FIG. 12, the second response module 102 includes a clip input unit 1024, a clip acquisition unit 1025, and an acquisition unit 1026. The clip input unit 1024 is configured to display a video clip input control for the virtual video space scene on the virtual display interface in response to a scene editing operation on the virtual video space scene by the first object. The clip acquisition unit 1025 is configured to obtain clip progress information for the input multi-view video in response to an input operation on the video clip input control, and use the video clip indicated by the clip progress information as the video clip to be produced. The acquisition unit 1026 is configured to obtain object data of the first object in the virtual video space scene during the process of playing the video clip to be produced.

[0179] For the specific implementation manners of the clip input unit 1024, the clip acquisition unit 1025, and the acquisition unit 1026, reference may be made to the description of step S103 in the embodiment corresponding to FIG. 3 above, and details are not described herein again.

[0180] In some embodiments, as again referenced in FIG. 12, the above data processing apparatus 1 further includes a second display module 123 and a fifth response module. The second display module 123 is configured to display a playback progress control bar on the virtual display interface during the process of playing the video clip to be produced. The playback progress control bar includes a pause control, a start control, and a multiple control. The fifth response module 124 is configured to pause the playback of the video clip to be produced in response to a trigger operation on the pause control. The fifth response module is further configured to continue playing the video clip to be produced in response to a trigger operation on the start control. The fifth response module is further configured to adjust the playback speed of the video clip to be produced according to the selected playback multiple in response to a selection operation on the multiple control.

[0181] For the specific implementation manners of the second display module 123 and the fifth response module, reference may be made to the description of step S103 in the embodiment corresponding to FIG. 3 above, and details are not described again here.

[0182] FIG. 13 is a structural diagram of a computer device provided in an embodiment of the present application. As shown in FIG. 13, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. The computer device 1000 may further include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to realize communication between these components. The user interface 1003 may include a display and a keyboard. The selectable user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may include a standard wired interface and a wireless interface (e.g., a Wi-Fi interface), for example. The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one magnetic disk memory. For example, the memory 1005 may further include at least one storage device located remotely from the processor 1001. As shown in Fig. 13, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0183] In the computer device 1000 shown in FIG. 13, the network interface 1004 can provide network communication network elements. The user interface 1003 is mainly used to provide an interface for user input. The processor 1001 can be used to call the device control application program stored in the memory 1005, thereby realizing the following steps. In response to a trigger operation for a multi-view video, a virtual video space scene corresponding to the multi-view video is displayed, and the multi-view video is played in the virtual video space scene. In response to a scene editing operation on the virtual video space scene, the first object obtains object data of the object in the virtual video space scene. Here, the first object refers to the object that starts the trigger operation for the multi-view video, and plays the production video associated with the multi-view video in the virtual display interface. Here, the production video is obtained by performing an editing process on the virtual video space scene based on the object data.

[0184] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method in any of the foregoing corresponding embodiments, and will not be described in detail here again. Regarding the description of the beneficial effects obtained by adopting the same method, it will not be described in detail again.

[0185] It should also be noted that the embodiments of the present application further provide a computer-readable storage medium, in which a computer program executed by the data processing device 1 is stored. The computer program includes program instructions. When the processor executes the program instructions, it can execute the data processing method described in any of the corresponding embodiments, and therefore, a detailed description thereof will not be repeated here. Furthermore, the beneficial effects of employing the same method will also not be repeated here. For technical details not disclosed in the embodiments of the computer-readable storage medium of the present application, please refer to the description of the method embodiments of the present application.

[0186] The computer-readable storage medium may be an internal storage unit of the data processing device provided in any of the above embodiments or the computer device, such as a hard disk or internal memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, or a flash card. The computer-readable storage medium may include not only the internal storage unit of the computer device but also an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0187] Also, what needs to be pointed out here is that the embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to execute the method provided in any of the corresponding embodiments described above.

[0188] In the description of the embodiments, claims, and drawings of the present application, terms such as "first" and "second" are used to distinguish different objects and are not used to describe a specific order. Also, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device including a series of steps or units is not limited to the listed steps or modules, but may further optionally include steps or modules not listed, or other step units specific to these processes, methods, apparatus, products, or devices.

[0189] As can be recognized by those skilled in the art, the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this specification can be realized by electronic hardware, computer software, or a combination of both. To clearly illustrate the compatibility between hardware and software, the configurations and steps of each example are generally described according to network elements in the above description. Whether these network elements are executed in hardware or in a software manner is determined by the specific application of the technical solution and the design constraints. Those skilled in the art can realize the network elements described in different ways for each specific application, but such realization should not be regarded as exceeding the scope of the present application.

[0190] The above disclosure is merely a preferred embodiment of the present application, and of course, it does not limit the scope of the present application. Therefore, any equivalent modifications made in accordance with the claims of the present application still fall within the scope of the present application.

Claims

1. 1. A data processing method implemented by a computing device, comprising: displaying a virtual video space scene corresponding to the multi-view video in response to a trigger operation on the multi-view video; a step of acquiring object data of a first object in the virtual video space scene in response to a scene editing operation on the virtual video space scene, the first object being an object that initiates a trigger operation on the multi-view video, the scene editing operation including a trigger operation on a sound input control in a virtual display interface; playing the multi-view video; acquiring audio data of a first object in the virtual video space scene in response to a trigger operation on the sound input control, and determining the audio data of the first object as object data to be applied to the multi-view video; and playing a production video associated with the multi-view video on the virtual display interface, the production video being obtained by performing an editing process on the virtual video space scene based on the object data; A method comprising:

2. the audio data of the first object includes at least one of object audio data and background audio data; a step of acquiring audio data of a first object in the virtual video space scene in response to a trigger operation on the sound input control in the virtual display interface, and determining the audio data of the first object as object data to be applied to the multi-view video, a step of displaying a sound adding mode list, the sound adding mode list including an object sound adding control and a background sound adding control; displaying a video object that can be sounded in response to a selection operation on the object sound-insertion control; a step of setting the selected audio-enabled video object as an audio-enabled object in response to a selection operation on the audio-enabled video object, the audio-enabled video object belonging to video objects displayed in the multi-view video; In the process of playing the multi-view video, obtaining object audio data of the first object based on the object to which sound is to be added, and determining the object audio data as the object data; In response to a selection operation on the background sound input control, during the process of playing the multi-view video, obtaining background sound data of the first object and determining the background sound data as the object data; The method of claim 1 , comprising:

3. In the process of playing the multi-view video, the step of acquiring object audio data of the first object based on the object to which sound is to be added and determining the object audio data as the object data includes: a step of performing a muting process on the video and audio data corresponding to the object to which the sound is to be added; When the object to which the sound is to be added is in a speaking state, displaying text information and soundtrack information corresponding to the video and audio data, and acquiring object audio data of the first object; determining the object audio data as the object data; The method of claim 2 , comprising:

4. The method comprises: When the object data is the object audio data, performing a replacement process on video audio data corresponding to the object to be sounded in the multi-view video using the object audio data to obtain a production video; when the object data is the background audio data, overlaying the background audio data onto the multi-view video to obtain a production video; The method of claim 2 further comprising:

5. A data processing method executed by a computer device, comprising: displaying a virtual video space scene corresponding to the multi-view video in response to a trigger operation on the multi-view video; In response to a scene editing operation on the virtual video space scene, a step of acquiring object data of an object in the virtual video space scene by a first object, where the first object refers to an object that starts a trigger operation on the multi-view video, and the scene editing operation includes a trigger operation on performance control in a virtual display interface, a step of playing the multi-view video, in response to a trigger operation on the performance control in the virtual display interface, acquiring pose data and image data of the first object in the virtual video space scene, and determining the pose data and the image data as object data to be applied to the multi-view video, including steps, a step of playing a production video associated with the multi-view video in the virtual display interface, where the production video is obtained by performing an editing process on the virtual video space scene based on the object data, the production video includes a performance object associated with the first object, and the performance object in the production video is displayed based on the pose data and the image data, steps, including methods.

6. The step of acquiring the pose data and the image data of the first object in the virtual video space scene and determining the pose data and the image data as object data to be applied to the multi-view video includes: a step of displaying a performance mode list, where the performance mode list includes a character replacement control and a character setup control, a step of displaying a replaceable video object in response to a trigger operation on the character replacement control, and setting the selected replaceable video object as a character replacement object in response to a selection operation on the replaceable video object, the replaceable video object belonging to video objects displayed in the multi-view video, and during a process of playing the multi-view video, obtaining posture data and image data of the first object based on the character replacement object, and determining the posture data and the image data as the object data; acquiring posture data and image data of the first object in response to a trigger operation on the character setup control during the process of playing the multi-view video, and determining the posture data and the image data as the object data; The method of claim 5 , comprising:

7. The method comprises: when the object data is acquired by triggering the character replacement control, releasing the display of the character replacement object in the performance view video, and merging a performance object that matches the performance object data into the performance view video to obtain a production video, wherein the performance view video is obtained by shooting the virtual video space scene using a performance view during the process of playing the multi-view video after triggering the performance control, and the performance object data refers to the data that the object data displays in the performance view; when the object data is acquired by triggering the character setup control, merging the performance object that matches the performance object data into the performance view video to obtain a production video; The method of claim 6 further comprising:

8. In response to a trigger operation on the character replacement control, displaying a replaceable video object, and in response to a selection operation on the replaceable video object, setting the selected replaceable video object as a character replacement object, the steps are: In response to a trigger operation on the character replacement control, determining the currently displayed video object on the object virtual video screen as a replaceable video object; In response to a marking operation on the replaceable video object, displaying the marked replaceable video object according to a first display method, wherein the first display method is different from the display methods of other video objects other than the replaceable video object, setting the marked replaceable video object as a character replacement object, and the object virtual video screen is used to display a virtual video space scene in the view where the first object is located. The method according to claim 6, comprising the above steps.

9. In response to a trigger operation on the character replacement control, displaying a replaceable video object, and in response to a selection operation on the replaceable video object, setting the selected replaceable video object as a character replacement object, the steps are: In response to a trigger operation on the character replacement control, displaying at least one video clip corresponding to the multi-view video; In response to a selection operation on the at least one video clip, displaying the video object included in the selected video clip and determining the video object included in the selected video clip as a replaceable video object; In response to a marking operation on the replaceable video object, displaying the marked replaceable video object according to a first display method, wherein the first display method is different from the display methods of other video objects other than the replaceable video object, and setting the marked replaceable video object as a character replacement object. The method according to claim 6, comprising the above steps.

10. The method includes: In the process of obtaining the pose data and image data of the first object, displaying a mirror preview control on the virtual display interface; Responding to a trigger operation on the mirror preview control, displaying a performance virtual video screen in a performance preview area on the virtual display interface, where the performance virtual video screen includes the performance object fused into the virtual video space scene; The method according to claim 5, further comprising the above steps.

11. The method includes: Displaying an image customization list on the virtual display interface; Responding to the completion of a configuration operation on the image customization list, updating the image data according to the configured image data to obtain configured image data, where the configured image data includes clothing data, body type data, voice data, and facial data; Displaying the performance object by a performance action and a performance image in the production video, where the performance action is determined based on the pose data of the first object, and the performance image is determined based on at least one of the clothing data, the body type data, the voice data, and the facial data; The method according to claim 5, further comprising the above steps.

12. Displaying a replacement transparency input control for the character replacement object on the virtual display interface; Responding to an input operation on the replacement transparency input control, obtaining transparency information for the input character replacement object; Performing an updated display of the transparency for the character replacement object on the virtual display interface according to the transparency information, and displaying a position cursor of the character replacement object in the virtual video space scene after the transparency update; The method according to claim 6, further comprising the above steps.

13. A data processing method executed by a computer device, comprising: Displaying an object invitation control on the virtual display interface; In response to a trigger operation on the object invitation control, a step of displaying an object list, wherein the object list includes an object having an association relationship with a first object; In response to a selection operation on the object list, a step of causing a target virtual reality device associated with a second object to display a virtual video space scene by sending an invitation request to the target virtual reality device associated with the second object; In response to a trigger operation on the multi-view video, a step of displaying the virtual video space scene corresponding to the multi-view video; In response to a scene editing operation on the virtual video space scene, a step of obtaining object data of the first object in the virtual video space scene, wherein the first object refers to an object that starts a trigger operation on the multi-view video; In the virtual display interface, a step of playing a production video associated with the multi-view video, wherein the production video is obtained by performing an editing process on the virtual video space scene based on the object data; A method comprising.

14. The step of causing a target virtual reality device associated with the second object to display the virtual video space scene by sending an invitation request to the target virtual reality device associated with the second object in response to a selection operation on the object list includes: Starting an invitation request for the second object to a server, and including a step of the server sending an invitation request to a target virtual reality device associated with the second object; When the target virtual reality device receives the invitation request, displaying the virtual video space scene on the target virtual reality device; The method is 14. The method of claim 13, further comprising the step of displaying a target virtual object on an object virtual video screen when the target virtual reality device accepts the invitation request and displays the virtual video space scene, wherein the second object enters the virtual video space scene through the target virtual object, the target virtual object is associated with image data of the second object, and the object virtual video screen is used to display the virtual video space scene in a view where the first object is located.

15. the scene editing operation includes a trigger operation for a performance control in the virtual display interface by the first object; When the second object has already triggered the performance control, The step of acquiring object data of a first object in the virtual video space scene in response to a scene editing operation on the virtual video space scene includes: and displaying target object data corresponding to the target virtual object on the object virtual video screen in response to a trigger operation of the performance control in the virtual display interface by the first object to obtain object data of the first object in the virtual video space scene, wherein the production video includes a performance object associated with the first object and the target virtual object, the performance object in the production video is displayed based on the object data, and the target virtual object in the production video is displayed based on the target object data.

15. The method of claim 14.

16. A data processing method executed by a computer device, comprising: displaying a virtual video space scene corresponding to the multi-view video in response to a trigger operation on the multi-view video; In response to a scene editing operation on the virtual video space scene, a step of obtaining object data of an object in the virtual video space scene by a first object, where the first object refers to an object that starts a trigger operation on the multi-view video, A step of displaying a shopping control in a virtual display interface; In response to a trigger operation on the shopping control, a step of displaying virtual items that can be purchased according to a second display method, where the second display method is different from the display method of virtual items that can be purchased before triggering the shopping control, and the virtual items that can be purchased belong to the items displayed in the virtual video space scene, In response to a selection operation on the purchasable virtual item, a step of setting the selected purchasable virtual item as a purchased item and displaying purchase information corresponding to the purchased item in the virtual display interface; A step of playing a production video associated with the multi-view video in the virtual display interface, where the production video is obtained by performing an editing process on the virtual video space scene based on the object data, A method comprising the above steps.

17. A data processing method executed by a computer device, comprising: In response to a trigger operation on a multi-view video, a step of displaying a virtual video space scene corresponding to the multi-view video; In response to a scene editing operation on the virtual video space scene, a step of obtaining object data of an object in the virtual video space scene by a first object, where the first object refers to an object that starts a trigger operation on the multi-view video, A step of displaying a main lens virtual video screen in a virtual display interface, where the main lens virtual video screen is used to display the virtual video space scene in the main lens view; A step of displaying a moving view switching control in the virtual display interface; In response to a trigger operation on the moving view switching control, obtaining a view of the virtual video space scene of the first object after movement as a moving view; Switching and displaying the main lens virtual video screen to a moving virtual video screen in the moving view of the virtual video space scene; Playing a production video associated with the multi-view video in the virtual display interface, wherein the production video is obtained by performing an editing process on the virtual video space scene based on the object data; A method comprising. **Claim 18**: A data processing method executed by a computer device, In response to a trigger operation on a multi-view video, displaying a virtual video space scene corresponding to the multi-view video; In response to a scene editing operation on the virtual video space scene, obtaining object data of a first object in the virtual video space scene, wherein the first object refers to an object that starts a trigger operation on the multi-view video; Displaying a main lens virtual video screen in the main lens view of the virtual video space scene in the virtual display interface; Displaying a pointing view switching control in the virtual display interface; In response to a trigger operation on the pointing view switching control, displaying a pointing cursor in the virtual display interface; In response to a movement operation on the pointing cursor, obtaining a view of the virtual video space scene of the moved pointing cursor as a pointing view; Switching and displaying the main lens virtual video screen to a pointing virtual video screen in the pointing view of the virtual video space scene; Playing a production video associated with the multi-view video in the virtual display interface, wherein the production video is obtained by performing an editing process on the virtual video space scene based on the object data; A method comprising.

19. A data processing apparatus, comprising: a first response module configured to display a virtual video space scene corresponding to a multi-view video in response to a trigger operation on the multi-view video; a second response module configured to acquire object data of a first object in the virtual video space scene in response to a scene editing operation on the virtual video space scene, wherein the first object refers to an object that starts a trigger operation on the multi-view video, and the scene editing operation includes a trigger operation on a voiceover control in a virtual display interface; a first response unit, which plays the multi-view video, acquires voice data of the first object in the virtual video space scene in response to a trigger operation on the voiceover control, and determines the voice data of the first object as object data to be applied to the multi-view video; configured as such; comprising the first response unit; the second response module; and a video playback module configured to play a production video associated with the multi-view video in a virtual display interface, wherein the production video is obtained by performing an editing process on the virtual video space scene based on the object data; an apparatus.

20. A computer device, comprising a processor, a memory, and a network interface, wherein the processor is connected to the memory and the network interface, the network interface is used to provide a data communication function, the memory is used to store program code, and the processor is used to execute the data processing method according to any one of Claims 1 to 18 by calling the program code.

21. A computer program that, when executed by a processor, realizes the data processing method according to any one of Claims 1 to 18.

Citation Information

Patent Citations

  • UE engine-based performance capturing system

    CN108564643A

  • Video processing method and device based on virtual reality, electronic equipment and medium

    CN113709543A

  • Authoring and presenting 3D presentations in augmented reality

    US10438414B2

  • Augmenting a physical object with virtual components

    US11024098B1

  • Facilitating editing of virtual-reality content using a virtual-reality headset

    US20180121069A1