Video processing method and apparatus, and storage medium and electronic device
By combining mobile terminals and the cloud, and using cloud processing technology to replace proprietary hardware equipment, low-cost and low-barrier video processing in virtual studios has been achieved, solving the problem of high costs in virtual studios and improving their accessibility and flexibility.
Patent Information
- Application Number
- PCT/CN2025/080786
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-11
- Filing Date
- 2025-03-05
- Publication Date
- 2026-01-15
AI Technical Summary
The implementation of virtual studios requires high-end hardware and software, resulting in high costs and high barriers to entry, making large-scale adoption difficult.
Raw video stream data is collected by mobile terminals and scene replacement is performed in the cloud to generate virtual scene video stream data, reducing reliance on proprietary hardware equipment and replacing proprietary camera equipment with cloud processing technology.
It reduces the cost of video acquisition, simplifies hardware requirements, lowers the barrier to entry, and improves the accessibility and flexibility of virtual studio technology.
Smart Images

Figure CN2025080786_15012026_PF_FP_ABST
Abstract
Description
A video processing method, apparatus, storage medium, and electronic device
[0001] This application claims priority to Chinese Patent Application No. 202410931451.0, filed on July 11, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure relates to a video processing method, apparatus, storage medium, and electronic device. Background Technology
[0003] A virtual studio is a method of generating video by embedding the foreground image from video data captured by a real camera into a virtual background and using a combination of virtual and real elements.
[0004] Currently, the implementation of virtual studios requires the construction of a real green screen environment, which incurs costs. In addition, the high requirements for hardware computing power and software functionality also lead to high costs, resulting in a high overall cost for the implementation of virtual studios. Furthermore, the high cost of upgrading both hardware and software makes the implementation of virtual studios a high-barrier process. Summary of the Invention
[0005] This disclosure provides a video processing method, apparatus, storage medium, and electronic device to reduce the cost and difficulty of implementing virtual studios.
[0006] In a first aspect, embodiments of this disclosure provide a video processing method applied to a client, the method comprising:
[0007] A video display page is provided, which includes a first display area and a second display area.
[0008] The system acquires the original video stream data from the client, transmits the original video stream data to the cloud, and receives virtual scene video stream data from the cloud. The virtual scene video stream data is obtained by the cloud replacing the original video stream data with a scene. The original video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene.
[0009] The original video stream data is displayed in the first display area; and the virtual scene video stream data is displayed in the second display area.
[0010] Secondly, this disclosure also provides a video processing method applied in the cloud, the method comprising:
[0011] Receive raw video stream data transmitted by at least one client, and extract the target object from the raw video stream data;
[0012] The target object is then integrated into the virtual scene.
[0013] The target object in the virtual scene is captured by a virtual camera to form virtual scene video stream data, which is then transmitted to the client.
[0014] Thirdly, embodiments of this disclosure also provide a video processing apparatus integrated into a client, the apparatus comprising:
[0015] A page display module is used to display a video page, which includes a first display area and a second display area;
[0016] The video stream data processing module is used to acquire the original video stream data of the client, transmit the original video stream data to the cloud, and receive virtual scene video stream data fed back by the cloud. The virtual scene video stream data is obtained by the cloud replacing the original video stream data with a scene. The original video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene.
[0017] The video stream data display module is used to display the original video stream data in the first display area; and to display the virtual scene video stream data in the second display area.
[0018] Fourthly, embodiments of this disclosure also provide a video processing apparatus integrated in the cloud, the apparatus comprising:
[0019] The target object extraction module is used to receive raw video stream data transmitted by at least one client and extract the target object from the raw video stream data;
[0020] A scene fusion module is used to fuse the target object into a virtual scene;
[0021] The video stream data generation module is used to collect video stream data of the target object in the virtual scene based on the virtual camera, form virtual scene video stream data, and transmit the virtual scene video stream data to the client.
[0022] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0023] One or more processors;
[0024] Storage device for storing one or more programs.
[0025] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method provided in any embodiment of this disclosure.
[0026] Sixthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video processing method provided in any embodiment of this disclosure. Attached Figure Description
[0027] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0028] Figure 1 is a schematic diagram of the implementation process of a virtual studio provided in an embodiment of this disclosure;
[0029] Figure 2 is a schematic flowchart of a video processing method provided in an embodiment of this disclosure;
[0030] Figure 3 is a schematic diagram of a video page provided in an embodiment of this disclosure;
[0031] Figure 4 is a schematic diagram of a video page provided in an embodiment of this disclosure;
[0032] Figure 5 is a flowchart of a video processing method provided in an embodiment of this disclosure;
[0033] Figure 6 is a schematic diagram of a processing task allocation provided in an embodiment of this disclosure;
[0034] Figure 7 is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this disclosure;
[0035] Figure 8 is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this disclosure;
[0036] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0037] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0038] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0039] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0040] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0041] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0042] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0043] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0044] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0045] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0046] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0047] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0048] For example, referring to Figure 1, which is a schematic diagram of the implementation process of a virtual studio provided by an embodiment of this disclosure. By constructing a realistic green screen environment, video data from the green screen environment is captured using dedicated camera equipment. Simultaneously, camera parameters are recorded during the video data acquisition process. The video data and camera parameters are transmitted to a powerful local video processing engine via a physical interface and low-latency transmission protocols such as a local area network. This video processing engine can be a workstation or dedicated hardware. The video data is embedded into a virtual background using hardware devices. During this process, the video data undergoes processes such as keying, AI inference, noise reduction, and virtual scene fusion. The video processing engine provides a preview of the virtual production effects and can stream or push the generated virtual scene video stream.
[0049] The implementation of the aforementioned virtual studio requires the construction of a specific green screen environment and the configuration of dedicated camera and hardware equipment, resulting in high costs and a high barrier to entry, hindering large-scale adoption. To address these technical issues, this disclosure provides a video processing method. Figure 2 is a flowchart illustrating a video processing method provided in an embodiment of this disclosure. This embodiment is applicable to the implementation of a virtual studio combining mobile terminals and the cloud. The method involves collecting raw video stream data via a mobile terminal and then replacing the scene in the raw video stream data via the cloud to obtain virtual scene video stream data. This method can be executed by a video processing device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal like a smartphone or tablet. In this embodiment, the video processing scenario can optionally be a live streaming scenario, for example, where the client is a broadcaster equipped with a camera. The broadcaster collects raw video stream data via the camera, processes it in the cloud to obtain virtual scene video stream data, which is then pushed into the stream. Optionally, the video processing scenario could be a short video production scenario, where the client is equipped with a camera to capture raw video stream data, or the client reads raw video stream data already stored in the client, processes it in the cloud to obtain virtual scene video stream data, and then exports, stores, or uploads the virtual scene video stream data to a server. In this embodiment, virtual studio technology is implemented through a mobile terminal such as a smartphone, replacing dedicated camera equipment and reducing costs during video acquisition.
[0050] As shown in Figure 2, the method includes:
[0051] S110. Display a video page, which includes a first display area and a second display area.
[0052] S120. Obtain the original video stream data from the client, transmit the original video stream data to the cloud, and receive the virtual scene video stream data fed back from the cloud. The virtual scene video stream data is obtained by the cloud replacing the original video stream data with a scene. The original video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene.
[0053] S130. Display the original video stream data in the first display area; and display the virtual scene video stream data in the second display area.
[0054] In this embodiment, the raw video stream data is the video stream data to be processed. For example, it can be obtained in real time through a camera configured on the client, or it can be read from the client's storage space. The raw video stream data includes the target object. The acquisition environment of the raw video stream data is not limited here. The raw video stream data can be video stream data captured in any environment through video capture such as a camera, reducing the limitation on the shooting environment of the raw video stream data, eliminating the need to build a specific shooting environment, and reducing the cost of the video shooting process.
[0055] Optionally, before transmitting the raw video stream data to the cloud, the client may further perform preprocessing on the raw video stream data, including but not limited to noise reduction and beautification. Correspondingly, the client transmits the preprocessed raw video stream data to the cloud, simplifying the cloud's processing flow for the raw video stream data and reducing the processing load on the cloud.
[0056] The client transmits the raw video stream data to the cloud. The cloud is configured with processing algorithms for the raw video stream data. The cloud uses these algorithms to perform scene replacement on the raw video stream data transmitted by the client, resulting in virtual scene video stream data. The virtual scene video stream data differs from the original video stream data in that the target object is located in a real scene, while the virtual scene video stream data uses a virtual scene. For example, the raw video stream data might be a video of user A in a room, captured by a camera. Accordingly, the target object in the raw video stream data is user A, and the real scene is the room. After processing the raw video stream data in the cloud, the resulting virtual scene video stream data retains user A as the target object, but the scene is a virtual scene distinct from the real scene. For example, the virtual scene could be, but is not limited to, a square, a beach, or an ancient temple. The virtual scene can be pre-selected based on user needs.
[0057] In some embodiments, before acquiring the client's raw video stream data, the process may further include selecting a virtual scene and previewing the virtual scene. Optionally, in response to a trigger operation on the video page, a list of virtual scenes is displayed. This list may be a list of virtual scene names or a list of virtual scene images. In response to a selection operation in the virtual scene list, the selected virtual scene is determined. A virtual scene instruction is generated based on the selected virtual scene and transmitted to the cloud. Alternatively, virtual scene information may be carried during the transmission of the raw video stream data. This virtual scene information may be the name or identifier of the virtual scene. The cloud determines the virtual scene selected by the client based on the virtual scene information and performs scene replacement on the raw video stream data based on this virtual scene.
[0058] Optionally, the system receives virtual scene description information input by the user, such as text or voice input. This virtual scene description information is then transmitted to the cloud, enabling the cloud to generate a virtual scene matching the description. The cloud may be configured with a first virtual scene generation model, which can be a neural network model, used to construct the virtual scene. The virtual scene description information is input into the first virtual scene generation model, which outputs a virtual scene matching the description.
[0059] Optionally, in response to an image selection operation, at least one image selected by the user is determined and uploaded to the cloud, so that the cloud can determine a virtual scene matching the at least one image. The at least one image may include the same scene, for example, at least one image taken of the same scene, or the at least one image may be taken from different angles of the same scene. For example, the cloud may be configured with a second virtual scene generation model, which inputs the at least one image and outputs a virtual scene matching the scene in the at least one image. For example, the cloud may perform similarity calculations between the at least one image and stored virtual scenes, and determine the virtual scene with the highest similarity as the virtual scene matching the scene in the at least one image.
[0060] The video page is used to display video data. For example, see Figures 3 and 4, which are schematic diagrams of a video page provided in an embodiment of this disclosure. The video page includes a first display area and a second display area, and before processing the original video stream data, it also includes previewing at least one of a real scene and a virtual scene through the video page. For example, a real scene is displayed in the first display area, and / or a virtual scene is displayed in the second display area. In response to gesture operations within the second display area, the display position and / or display perspective of the virtual scene are switched, facilitating the display of the virtual scene to the user from different display positions and / or display perspectives.
[0061] Optionally, the cloud-based processing of the raw video stream data may include: performing image matting on each video frame in the virtual scene video stream data to obtain the target object image; rendering the target object image in the virtual scene; and capturing video stream data of the target object in the virtual scene using a virtual camera to form the virtual scene video stream data. The cloud then transmits the generated virtual scene video stream data to the client, which displays the virtual scene video stream data.
[0062] Optionally, the target object can be a foreground object in the original video stream data, such as a human figure. Optionally, the target object can be determined based on an object selection operation. For example, in response to a setting operation on the video page, multiple object types are displayed, including but not limited to human figures, tables, and food. By selecting at least one object type, the type of the target object is determined, and the type of the target object is transmitted to the cloud so that the cloud can extract the target object from the original video stream data based on the target object type. Optionally, the description information of the target object is received and transmitted to the cloud so that the cloud can extract the target object from the original video stream data based on the target object description information, and the extracted target object matches the target object description information. During the cloud's extraction of the target object from the original video stream data, the type or description information of the target object can be used as prior information to assist in the extraction. Optionally, the cloud is configured with a matting model used to extract the target object from the original video stream data. Specifically, each video frame in the original video stream data can be sequentially input into the matting model to obtain human figures in each video frame of the original video stream data, which can then be used as the target object. Alternatively, the type or description information of the target object can be converted into vector features. The vector features and each video frame in the original video stream data are then sequentially input into the matting model to obtain the target object in the original video stream data. This target object is matched with the type or description information of the target object.
[0063] The client receives virtual scene video stream data transmitted from the cloud and displays both the raw video stream data and the virtual scene video stream data on the video page. The first display area of the video page shows the raw video stream data captured by the client, while the second display area shows the virtual scene video stream data processed by the cloud. The raw video stream data and the virtual scene video stream data are displayed simultaneously on the video page.
[0064] The arrangement of the first and second display areas on the video page is not limited. The first and second display areas can be arranged vertically, as shown in Figure 3; alternatively, they can be arranged horizontally; or they can be displayed in different windows, as shown in Figure 4, where the first display area is displayed above the second display area in a floating window. Optionally, the different arrangements of the first and second display areas can be switched, and the arrangement of the first and second display areas on the video page will be updated in response to the switching operation. The specific method of the switching operation is not limited here.
[0065] Optionally, the shooting attribute information of the first and second display areas in the video page can be adjusted. The shooting attribute information of the display areas includes position information, size information, etc. In response to the adjustment operation of the first or second display area, the shooting attribute information of the first and / or second display areas is adjusted. For example, in the video page shown in Figure 3, the adjustment operation can be a drag operation on the common edge of the first and second display areas, or a scaling operation on the first or second display area. In response to the above adjustment operations, one display area can be enlarged, and the other display area can be shrunk. In the video page shown in Figure 4, the adjustment operation can be a scaling operation on the first display area to adjust its size; the adjustment operation can also be a drag operation on the first display area to adjust its position. In Figure 3 or Figure 4, the adjustment operation can also be a position swap operation. In response to the position swap operation, virtual scene video stream data is displayed in the first display area, and original video stream data is displayed in the second display area.
[0066] The technical solution provided in this embodiment acquires raw video stream data through an image acquisition component configured on the client side, replacing dedicated camera equipment and reducing equipment costs in the video stream data acquisition process. The client transmits the raw video stream data to the cloud, where the cloud performs scene replacement on the raw video stream data to obtain virtual scene video stream data, which is then transmitted back to the client for display. This eliminates the need to deploy dedicated hardware locally, further reducing equipment costs in the video processing process. Simultaneously, the cloud can process raw video stream data transmitted from different clients, lowering the cost and barrier to entry for a large number of users to implement virtual studio technology.
[0067] Building upon the above embodiments, during the display of raw video stream data and virtual scene video stream data, users can interact with the video page. For example, users can input interactive operations on the video page, and the virtual scene video stream data corresponding to the interactive operation will be displayed on the video page. Alternatively, the user can be a target object in the raw video stream data; if the target object constitutes a set action, the virtual scene video stream data corresponding to the set action will be displayed on the video page. Through user interaction with the video page, the flexibility and diversity of virtual scene video stream data generation and display processes are improved.
[0068] Optionally, in response to a first trigger operation within the second display area, the shooting attribute information of the target object within the second display area in the virtual scene is adjusted, wherein the shooting attribute information includes one or more of the following: position information, size information, and angle information. Specifically, the position information characterizes the position of the target object in the virtual scene, the angle information characterizes the orientation of the target in the virtual scene, and the size information characterizes the size of the target object in the virtual scene.
[0069] In this embodiment, the shooting attribute information of the target object in the virtual scene can be updated through a first trigger operation during the preview stage or the video display stage. Taking a live streaming scene as an example, the preview stage is the stage before entering the live streaming state, and the video display stage is the stage after entering the live streaming state. During the preview stage, a preview screen is displayed in the second display area. This preview screen can be partial virtual scene video stream data obtained by replacing the scene of partial video stream data in the original video stream data in the cloud. During the video display stage, the virtual scene video stream data obtained by cloud processing is displayed in the second display area.
[0070] During the cloud processing of raw video stream data, target objects from the raw video stream data are integrated into a virtual scene. The initial shooting attributes of the target object in the virtual scene can be pre-set. These initial shooting attributes can include the target object's initial position, initial size, and initial angle. For example, the initial position of the target object in the virtual scene could be the center of the virtual scene or the most frequently used position. The initial size of the target object can be determined based on the scale of the virtual scene and the actual size of the target object. The initial angle for shooting the target object can be a preset angle or the most frequently used angle. It is understood that different shooting attribute information will result in different virtual scene video streams.
[0071] The first triggering operation can be an operation used to adjust the attribute information of the target object. For example, the first triggering operation can be a position adjustment operation, a size adjustment operation, and an angle adjustment operation, respectively used to adjust the position information, size information, and angle information of the target object. Optionally, the first triggering operation can include a selection operation and an adjustment operation. The selection operation can include, but is not limited to, clicking or long-pressing the target object in the second display area, or selecting the target object. In response to the selection operation, the outline of the target object can be marked in the second display area, such as thickening the outline of the target object or displaying the outline of the target object in a set color, to indicate that the target object has been selected and is in an adjustable state. The adjustment operation can include, but is not limited to, dragging, scaling, and rotating the target object, used to adjust the position information, size information, and angle information. For example, the position information can be adjusted by dragging, and the position information of the target object after adjustment is determined based on the release position of the drag operation; the size information can be adjusted by scaling, increasing the size information of the target object in response to zooming in, and decreasing the size information of the target object in response to zooming out; the angle information can be adjusted by rotating, and the angle information of the target object is adjusted according to the angle change of the rotation operation.
[0072] Optionally, the virtual scene may include multiple sub-scenes. In the above embodiments, adjusting the position information of the target object may further include adjusting the sub-scene where the target object is located. For example, virtual scene 1 may include a pavilion sub-scene and a plaza sub-scene. When the first triggering operation is a position adjustment operation, the target sub-scene corresponding to the position adjustment operation is determined, and the sub-scene where the target object is located is updated to that target sub-scene. For example, by dragging the target object from the pavilion sub-scene to the plaza sub-scene, the position information of the target object is adjusted from the pavilion sub-scene to the plaza sub-scene, thus achieving sub-scene switching.
[0073] By adjusting the attribute information of the target object in the second display area through the first trigger operation, the configuration of the target object in the process of generating virtual scene video stream data is simplified. In response to the adjustment of the attribute information of the target object by different users, different virtual scene video stream data can be obtained, which can meet the generation needs of different users for virtual scene video stream data and realize the diversity of virtual scene video stream data.
[0074] The client determines to send the adjusted shooting attribute information to the cloud based on the first trigger operation, so that the cloud can generate the virtual scene video stream data based on the adjusted shooting attribute information. During the preview stage, when the client sends the shooting attribute information to the cloud, the cloud processes the original video stream data based on the received shooting attribute information to obtain virtual scene video stream data that matches the shooting attribute information. Specifically, based on the received shooting attribute information, the target object is merged into the virtual scene, and video is captured from the target object in the virtual scene using a virtual camera to obtain virtual scene video stream data. During the video display stage, when the client sends the shooting attribute information to the cloud, the cloud determines the received timestamp, processes the original video stream data after the received timestamp based on the received shooting attribute information to obtain virtual scene video stream data that matches the shooting attribute information. Specifically, based on the received shooting attribute information, at least one of the target object's position, size, and angle information in the virtual scene is updated, and video is captured from the updated target object in the virtual scene using a virtual camera to obtain virtual scene video stream data.
[0075] Optionally, in response to a set action of the target object in the first display area, the shooting attribute information of the target object in the virtual scene in the second display area is adjusted, wherein the set action is detected by the cloud and the shooting attribute information of the target object in the virtual scene is determined based on the set action.
[0076] The system pre-sets the correspondence between preset actions and shooting attribute information. The cloud detects the actions of target objects in the raw video stream data. Upon detecting a preset action, it determines the corresponding shooting attribute information and adjusts it accordingly, thereby regulating the shooting attribute information of the target object in the virtual scene video stream data. This adjustment can include adjusting position information, size information, and angle information. Different shooting attribute information corresponds to different preset actions. For example, a preset action might be walking or jumping to adjust position information; a first hand movement to adjust size information; and a second hand movement to adjust angle information. The preset actions are not limited and can be pre-set by the user. Optionally, the preset actions corresponding to position information can include a first preset action corresponding to sub-scene switching and a second preset action corresponding to position information adjustment within the same sub-scene.
[0077] In this embodiment, by setting the actions of the target object being photographed, the adjustment of the shooting attribute information in the virtual scene is controlled, which enhances the interactivity with the target object, simplifies the way to adjust the shooting attribute information of the target object in the virtual scene, and achieves the goal of adjusting the shooting attribute information without manual operation on the client.
[0078] Based on the above implementation, the virtual camera acquires video stream data of the target object in the virtual scene during camera movement, forming virtual scene video stream data; the camera movement mode of the virtual camera affects the video frames in the virtual scene video stream data. The camera movement mode of the virtual camera refers to the movement method of the virtual camera during the shooting process of the target object in the virtual scene. The virtual camera acquires video stream data of the target object in the virtual scene based on the target camera movement mode.
[0079] Optionally, the target camera movement mode includes one or more of the following: a follow mode of the virtual camera to the actual camera in the client; a follow mode of the virtual camera to the target object; and an automatic camera movement mode in which the virtual camera sets a trajectory in the virtual scene.
[0080] The virtual camera's following mode on the client can be understood as follows: during the process of acquiring raw video stream data on the client, the camera parameters of the actual camera on the client are recorded and transmitted to the cloud in real time. The cloud adjusts the shooting parameters of the virtual camera according to the real-time received camera parameters so that the shooting parameters of the virtual camera are consistent with the camera parameters of the actual camera on the client, thereby realizing the virtual camera's following mode on the client.
[0081] The virtual camera's following mode for the target object can be understood as follows: when the target object is in motion in the original video stream data, the target object is also in motion within the virtual scene. The virtual camera follows the target object to capture footage, ensuring that the target object is centered in the video frame captured by the virtual camera. For example, the method of capturing the original video stream data where the target object is in motion can be: the client's actual camera maintains a relative shooting position to the target object, for example, the client's actual camera maintains the same speed of movement as the target object; or, the client's actual camera is positioned at a fixed location, and the target object moves relative to the client's actual camera.
[0082] The automatic camera movement mode, where a virtual camera is positioned along a set trajectory within the virtual scene, can be understood as follows: the virtual camera moves along a pre-defined trajectory. For example, if the target object is stationary or moving within the virtual scene, the virtual camera moves along the pre-defined trajectory and captures video frames during the movement. Exemplary trajectories may include, but are not limited to, top-to-bottom, bottom-to-top, left-to-right, right-to-left, circular, rotating, or spiral trajectories, etc., and are not limited here. In automatic camera movement mode, the target object may be at any position within a video frame or may be outside of a video frame.
[0083] Optionally, in response to a second trigger operation on the video page, virtual scene video stream data generated based on a target camera movement mode is displayed in the second display area, wherein the target camera movement mode is determined based on the second trigger operation. The second trigger operation may be an operation for switching camera movement modes, a pre-set gesture operation, or a trigger operation on a camera movement mode switching control on the video page.
[0084] The client generates camera movement mode switching information based on a second trigger operation. This information can be the switched-off camera movement mode, such as its name or identifier. The client sends this information to the cloud, enabling the cloud to switch the camera movement mode in the virtual scene to the target camera movement mode and generate a video stream of the virtual scene based on the target camera movement mode. Correspondingly, the client receives the virtual scene video stream data generated by the cloud based on the switched target camera movement mode and displays it in the second display area.
[0085] In this embodiment, the camera movement mode of the virtual camera is switched through the second trigger operation, which enhances the interactivity of the generation process, increases the diversity of virtual scene video stream data, and generates virtual scene video stream data under different camera movement modes by switching camera movement modes in the virtual scene, thereby reducing the difficulty of camera movement in the process of shooting the original video stream data.
[0086] Figure 5 is a flowchart of a video processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations where virtual scene replacement is performed on raw video stream data transmitted from at least one client in the cloud to generate virtual scene video stream data. This method can be executed by a video processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a cloud server or a cloud server cluster. Referring to Figure 5, the method specifically includes:
[0087] S210. Receive raw video stream data transmitted by at least one client, and extract the target object from the raw video stream data.
[0088] S220. Integrate the target object into the virtual scene.
[0089] S230. Based on the virtual camera, collect video stream data of the target object in the virtual scene to form virtual scene video stream data, and transmit the virtual scene video stream data to the client.
[0090] In this embodiment, the target object can be a human image; alternatively, the target object can be determined based on the client's object selection operation in the original video stream data. For example, during the preview stage, the client selects the target object within a first display area based on the object selection operation. This object selection operation could be a box selection within the first display area or a click operation on the target object. For example, during the preview stage, the object selection operation could be a selection of the object type or an input of descriptive information about the target object. There can be one or more target objects in the original video stream data. The cloud determines the target object feature information based on the client's object selection operation to assist in the extraction of the target object.
[0091] Here, the extraction of target objects can be achieved through a pre-trained machine learning model. Optionally, the extraction of target objects from the original video stream data includes: inputting the video frames of the original video stream data into a matting model to obtain the target objects in each video frame. The matting model is a pre-trained machine learning model, such as a neural network model; the specific structure and type of the matting model are not limited here. In the above embodiment where the target object is determined based on an object selection operation, target object feature information is generated based on any one of the bounding box operation, object type selection operation, or target object description information. For example, the target object feature information can be in the form of a feature vector. For example, any one of the bounding box operation, object type selection operation, or target object description information can be input into a feature extraction model to obtain the target object feature information. The feature extraction model can be a convolutional neural network model or an embedding model, etc. Using the target object feature information as reference information facilitates the extraction of target objects that match the target object feature information. Specifically, the target object feature information and each video frame of the original video stream data are input into the matting model to obtain the target object corresponding to each video frame.
[0092] By performing image matting on the target object from the original video stream data to remove the background, scene replacement is achieved by merging the target object with the virtual scene. Optionally, the target object extracted from the original video stream data is in the form of a two-dimensional image. A three-dimensional target object model is constructed based on the two-dimensional image of the target object, and the three-dimensional target object model is rendered in the virtual scene to achieve the fusion of the target object and the virtual scene. For example, the two-dimensional image of the target object is input into the three-dimensional construction model to obtain the three-dimensional target object model. This three-dimensional construction model can be a pre-trained neural network model used for three-dimensional reconstruction of two-dimensional images. For example, it can be trained using a three-dimensional sample model and its two-dimensional image. The two-dimensional image of the three-dimensional sample model can be used as sample data, and the three-dimensional sample model can be used as a label model; the two-dimensional image of the three-dimensional sample model can be a two-dimensional image captured from any angle of the three-dimensional sample model. The two-dimensional image of the three-dimensional sample model is input into the three-dimensional construction model to be trained to obtain the predicted three-dimensional model. A loss function is generated based on the predicted three-dimensional model and the three-dimensional sample model. The three-dimensional construction model to be trained is trained based on the loss function. The above iterative training process is performed until a good three-dimensional construction model to be trained is obtained.
[0093] It is understandable that the cloud replaces the original video stream data with a virtual scene, which is different from the real scene of the original video stream data. Consequently, the lighting and shadow information in the original video stream data may not match the virtual scene, resulting in distortion of the generated virtual scene video stream data. For example, if the original video stream data was shot in a room scene, and the replaced virtual scene is a mountain forest scene, the lighting and shadow information of the room scene and the mountain forest scene will not match.
[0094] Optionally, after integrating the target object into the virtual scene, lighting and shadow information is set in the virtual scene to match the virtual scene. Specifically, a pre-defined correspondence between lighting and shadow information and the virtual scene is established. The target object in the virtual scene is then re-lit based on the corresponding lighting and shadow information. For example, a shadow processing algorithm is pre-set in the cloud, and the lighting and shadow information in the virtual scene is set based on this algorithm. By re-lighting the virtual scene and setting lighting and shadow information that matches the virtual scene, the realism of the virtual scene video stream data is improved.
[0095] A virtual camera captures video of target objects in a virtual environment, and the captured video frames form a virtual scene video stream. This virtual scene video stream data is then transmitted to the client for display. Optionally, in a live streaming scenario, the virtual scene video stream data is the video stream to be streamed, and can be processed for streaming.
[0096] The technical solution provided in this embodiment extracts the target object from the raw video stream data transmitted by the client, merges the target object into a virtual scene, and captures video of the target object in the virtual scene using a virtual camera, obtaining virtual scene video stream data of the target object within the virtual scene, thus replacing the scene of the raw video stream data. The virtual scene video stream data is then transmitted to the client, converting the raw video stream data captured by the client into virtual scene video stream data specific to the virtual scene. This assists the client in implementing virtual studio technology, reducing the hardware requirements and equipment costs for implementing virtual studio technology on the client side.
[0097] The cloud can communicate with multiple clients and receive raw video stream data transmitted by multiple clients. In some embodiments, the cloud processes the raw video stream data of each client independently to obtain virtual scene video stream data corresponding to that client. In some embodiments, the cloud extracts target objects from the raw video stream data of multiple clients, merges the target objects corresponding to different clients into the same virtual scene, and obtains virtual scene video stream data corresponding to multiple clients in the same virtual scene. Each client corresponds to at least one target object. For example, if multiple clients select the same virtual scene, merging the target objects corresponding to multiple clients into the same virtual scene can be achieved by using one or more virtual cameras to capture video in the virtual scene, obtaining one or more virtual scene video stream data. For example, each client can correspond to one virtual camera, and video capture is performed on the target objects extracted from the client's raw video stream data using that virtual camera to obtain the virtual scene video stream data corresponding to that client. Accordingly, different clients have different virtual scene video stream data. For example, video capture is performed in the virtual scene using one virtual camera to obtain one virtual scene video stream data, which serves as the virtual scene video stream data corresponding to multiple clients. In some embodiments, there is no interaction between different clients. To avoid interference between target objects corresponding to different clients in the same virtual scene, while the virtual camera corresponding to any client is capturing video of the target object corresponding to that client, the target objects corresponding to other clients are set to be invisible. That is, the virtual camera corresponding to any client is only used to capture the video stream data of the target object corresponding to that client. When the target objects corresponding to other clients enter the shooting range of the virtual camera corresponding to that client, the target objects corresponding to other clients are invisible relative to the virtual camera corresponding to that client, and the target objects of different clients do not interfere with each other. In this embodiment, by merging the target objects of different clients into the same virtual scene, the repeated rendering of the same virtual scene is reduced, and the reusability of the virtual scene is improved.
[0098] When different clients interact, the target objects corresponding to the clients with interactive relationships are visible to each other in the virtual scene. For example, in a live streaming scenario, different clients located in the same virtual room can interact, while clients located in different virtual rooms cannot. For instance, different clients can request to establish interactive relationships. For example, clients A, B, and C select the same virtual scene. Clients A and B have an interactive relationship, but neither has an interactive relationship with client C. The target objects corresponding to clients A, B, and C are merged into the same virtual scene. In this virtual scene, the target objects corresponding to clients A and B are visible to each other, while the target objects corresponding to clients A and B are not visible to the target object corresponding to client C.
[0099] In some embodiments, target objects corresponding to clients with interactive relationships can be merged into the same virtual scene, while target objects corresponding to clients without interactive relationships can be merged into independently rendered virtual scenes to avoid mutual interference between target objects.
[0100] Optionally, the raw video stream data is live video stream data, that is, video stream data collected by the client in a live streaming scenario; the raw video stream data sent by the client carries virtual room information, and the virtual room information carried in the raw video stream data sent by different clients may be the same or different. In a live streaming scenario with multiple people interacting, the virtual room information carried in the raw video stream data sent by multiple clients interacting is the same.
[0101] Optionally, fusing the target object into the virtual scene includes: fusing target objects from different original video streams carrying the same virtual room information into the same virtual scene. Additionally, forming virtual scene video stream data by collecting video stream data of the target object in the virtual scene using virtual cameras includes: collecting video stream data of the target object corresponding to different clients in the virtual scene using different virtual cameras, forming virtual scene video stream data corresponding to each client. The target object extracted from the original video stream data of each client corresponds to one virtual camera; here, the target object extracted from the original video stream data of each client can be one or more.
[0102] By fusing target objects from the raw video streams of multiple clients into the same virtual scene, the interactivity and scene uniformity of the live streaming scenario are improved. It is understood that the target objects corresponding to the multiple clients participating in the live stream have interactive relationships. For example, client A corresponds to target object 1, client B corresponds to target object 2, and clients A and B are in a live streaming co-hosting state. Clients A and B have the same virtual room information. Target object 1 and target object 2 are fused into the same virtual scene. The virtual scene video stream data of target object 1 is collected based on a first virtual camera, and the virtual scene video stream data of target object 2 is collected based on a second virtual camera. Target object 1 is visible relative to both the first and second virtual cameras, and target object 2 is also visible relative to both the first and second virtual cameras.
[0103] Based on the above embodiments, the cloud receives first control information transmitted by at least one of the clients and renders virtual effect data in the virtual scene. This first control information can be generated based on client interaction operations, including but not limited to triggering interactive controls on the client's video page and inputting preset interactive gestures on the interactive page. The first control information is used to control the triggering of virtual effect display. The cloud determines the virtual effect data corresponding to the first control information and renders the virtual effect data in the virtual scene. The rendering position of the virtual effect data can be determined based on the position of the target object corresponding to the virtual effect data. For example, the first control information transmitted by client A determines virtual effect data 1, and the target object corresponding to virtual effect data 1 is target object 1 corresponding to client A. Correspondingly, the rendering position of virtual effect data 1 is determined based on the position information of target object 1. For example, in a live-streaming scenario, target object 1 corresponding to client A and target object 2 corresponding to client B are in the same virtual scene. The first control information transmitted by client A determines virtual effect data 1, and the target object corresponding to virtual effect data 1 is target object 2 corresponding to client B. Correspondingly, the rendering position of virtual effect data 1 is determined based on the position information of target object 2. Optionally, the rendering location of the virtual effect data is within a preset range of the location of the target object corresponding to the virtual effect data, so that the virtual camera of the target object corresponding to the virtual effect data can capture the aforementioned virtual effect data.
[0104] For example, the first control information can be based on a trigger operation on a gift control. Correspondingly, the virtual effect data can be gift display data. The gift display data is rendered in the virtual scene, and the virtual camera can capture the aforementioned virtual effect data to form virtual scene video stream data including the virtual effect data. It is understood that the specific form of the virtual effect data is not limited here, and may include, but is not limited to, gift-giving effect data or other special effects data.
[0105] Optionally, in a live streaming scenario, the cloud combines the interactive information transmitted from the viewer's device to render virtual effect data corresponding to the interactive information in a virtual scene. The interactive information transmitted from the viewer's device includes, but is not limited to, information about gift-giving, and the corresponding virtual effect data can be gift display effect data, etc.
[0106] In this embodiment, virtual effect data is rendered in the virtual scene in response to the first control information transmitted by the client, thereby enhancing the interactivity between the client and the virtual scene.
[0107] Based on the above embodiments, a virtual camera is used to acquire video stream data of a target object in a virtual scene. The virtual camera can move and change angles within the virtual scene, which constitutes the camera movement process. The virtual camera can acquire video stream data of the target object through different camera movement modes. Optionally, receiving second control information from any of the clients, the target camera movement mode of the virtual camera corresponding to the client is switched, and the virtual camera corresponding to the client is controlled to form the virtual scene video stream data corresponding to the client based on the switched target camera movement mode. The second control information can be generated based on a second trigger operation of the client, used to switch the target camera movement mode of the virtual camera corresponding to the client. The target camera movement mode includes one or more of the following: a follow mode of the virtual camera to the actual camera in the client; a follow mode of the virtual camera to the target object; and an automatic camera movement mode with a set trajectory in the virtual scene, which will not be elaborated further here.
[0108] In a live streaming scenario with multiple users interacting, the second control information sent by each client is used to control the target camera movement mode of the virtual camera corresponding to that client. Consequently, the target camera movement mode of the virtual camera corresponding to different clients can be different, and the virtual scene video stream data corresponding to different clients will also be different, which improves the diversity of virtual scene video stream data and meets the virtual scene video stream data generation needs of different clients.
[0109] Based on the above embodiments, the raw video stream data collected by the client is configured with corresponding audio data. For example, the client merges the raw video stream data and audio data to obtain raw audio-video data, and transmits the raw audio data to the cloud. The cloud parses the raw audio-video data to obtain the raw video stream data and audio data, performs scene replacement on the raw video stream data to obtain virtual scene video stream data, merges the virtual scene video stream data with the audio data to obtain virtual scene merged audio-video data, and transmits the virtual scene merged audio-video data to the client.
[0110] Understandably, the original video stream data and audio data are matched. During the recording of the original video stream data, the camera movement of the client's actual camera matches the volume of the audio data, and the volume of the audio data is positively correlated with the distance between the client's actual camera and the target object. However, during the generation of the virtual scene video stream data, the camera movement mode of the virtual camera may differ from that of the client's actual camera. Consequently, the resulting virtual scene video stream data and audio data will not match, affecting the realism of the merged audio and video data of the virtual scene.
[0111] Optionally, audio data corresponding to the original video stream data is acquired; during the process of the virtual camera acquiring the virtual scene video stream data, the distance between the virtual camera and the target object is determined, and the volume of the audio data is adjusted based on the distance to obtain adjusted audio data; the virtual scene video stream data and the adjusted audio data are then merged. The volume of the audio data is adjusted according to the distance between the virtual camera and the target object, and the volume of the audio data is positively correlated with the aforementioned distance. The volume of the adjusted audio data is matched with the virtual scene video stream data, and by merging the virtual scene video stream data and the adjusted audio data, merged virtual scene audio and video data is obtained.
[0112] Based on the above embodiments, the target object is integrated into the virtual scene. The shooting attribute information of the target object in the virtual scene includes one or more of position information, size information, and angle information, which will not be elaborated here. The cloud is also used to receive third control information from the client. The third control information is used to adjust the shooting attribute information of the target object in the virtual scene and is generated based on the client's first trigger operation. The shooting attribute information of the target object in the virtual scene is adjusted based on the third control information. Optionally, the third control information may be the target shooting attribute information of the target object, and the target object is adjusted based on the target shooting attribute information. For example, the current position information of the target object is position 1, and the third control information includes the target position information as position 2. In response to the third control information, the position information of the target object is adjusted from position 1 to position 2. Optionally, the third control information may be the shooting attribute information adjustment value of the target object, and the target object is adjusted based on the shooting attribute information adjustment value. For example, the current size information of the target object is size 1, and the third control information includes a size increase ratio of 5%. In response to the third control information, the size information of the target object is increased by 5% based on size 1.
[0113] In some embodiments, when a predetermined action of the target object is detected, the shooting attribute information of the target object in the virtual scene is adjusted. After extracting the target object from the original video stream data, the method further includes matching the target object's actions with predetermined actions. When a predetermined action is detected in the target object's actions, the shooting attribute information of the target object in the virtual scene is adjusted based on the predetermined action. The predetermined action may correspond to a specific adjustment method for the shooting attribute information; for example, predetermined action 1 corresponds to increasing size information, predetermined action 3 corresponds to decreasing size information; predetermined action 3 corresponds to adjusting angle information in a predetermined direction, predetermined action 4 corresponds to adjusting position information, etc.
[0114] In some embodiments, the virtual scene includes multiple sub-scenes. Correspondingly, the adjustment of position information includes sub-scene switching and adjustment of position information within a sub-scene. Different position information adjustment methods correspond to different third control information or different set actions. For example, the first set action is walking, corresponding to adjustment of position information within a sub-scene; the second set action is jumping, corresponding to sub-scene switching. For example, the target position information in the third control information can be other position information within the sub-scene where the target object is located, corresponding to adjustment of position information within a sub-scene; or the target position information in the third control information can be position information within a sub-scene other than the sub-scene where the target object is located, corresponding to sub-scene switching.
[0115] Optionally, adjusting the shooting attribute information of the target object in the virtual scene includes: switching the sub-scene where the target object is located when the third control information or the detected set action is used to control the sub-scene switching; and controlling the virtual camera corresponding to the target object to collect video stream data of the target object in the switched sub-scene. Controlling the target object to switch sub-scenes in the virtual scene through client-triggered operations or set actions of the target object enhances the interactivity of the target object in the virtual scene and the diversity of the virtual scene video stream data.
[0116] Based on the above embodiments, the video processing method provided in this disclosure is executed using at least one GPU. The video processing method includes multiple processing steps, each corresponding to at least one processing task. The processing steps include extracting a target object, fusing the target object with a virtual scene, setting lighting and shadow information for the virtual scene, and controlling the camera movement mode of a virtual camera. Specifically, the processing tasks corresponding to extracting the target object include portrait recognition and portrait matting tasks; the processing tasks corresponding to fusing the target object with the virtual scene include 3D model construction and scene fusion tasks; the processing tasks corresponding to setting lighting and shadow information for the virtual scene include lighting and shadow setting tasks; and the processing tasks corresponding to controlling the camera movement mode of the virtual camera include camera movement mode switching tasks and virtual camera movement tasks.
[0117] Different processing tasks consume different amounts of resources during execution. The total resource consumption of video processing for raw video stream data from different clients also varies. The resource consumption for the same processing task can differ between clients, and even for the same client, the resource consumption for the same processing task can vary at different times. For example, taking portrait matting as an example, different clients have different numbers of target objects, resulting in different resource consumption for portrait matting tasks on different clients. Similarly, for the same client, the number of target objects in video frames corresponding to different timestamps can differ, consequently leading to different resource consumption for portrait matting tasks on the same client at different times.
[0118] The amount of available resources can vary between different GPUs, and the amount of available resources of a single GPU can also change in real time. For example, a single GPU can support video processing of raw video stream data from one or more clients, or multiple GPUs can support video processing of raw video stream data from one client.
[0119] In this embodiment, during the execution of the video processing method, the resource usage of each processing task and the available resources of each GPU are detected in real time, and each processing task is allocated to the at least one GPU based on the resource usage of each processing task.
[0120] If the available resources of a single GPU are sufficient to meet the resource requirements of multiple processing tasks corresponding to a video processing method, then the multiple processing tasks corresponding to the video processing method will be allocated to that single GPU. If the available resources of a single GPU are insufficient to meet the resource requirements of multiple processing tasks corresponding to a video processing method, then the multiple processing tasks corresponding to the video processing method will be allocated to multiple GPUs.
[0121] During the execution of the video processing method, the resource usage of each processing task can change dynamically. The GPUs corresponding to each processing task can be dynamically adjusted based on the resource usage of the processing tasks at each moment. Each GPU corresponds to at least one processing task, which includes a first processing task. If the resource usage of the first processing task increases during processing, leading to an increase in the total resource usage of the at least one processing task, and this increase exceeds the available resources of the GPUs, the first processing task is allocated to other GPUs. The remaining resources of the other GPUs are determined. If the remaining resources of any other GPU are greater than the resource usage of the first processing task, the first processing task is allocated to that other GPU.
[0122] A single GPU can include multiple GPU nodes, which can be used to execute different processing tasks. For example, a GPU can perform both facial recognition and facial matting tasks. The facial recognition task can be performed by a first GPU node, and the task by a second GPU node. These GPU nodes can be virtual computing nodes, created and released according to the processing needs of the tasks being performed by the GPU.
[0123] Multiple worker threads are created in the cloud to control data processing and flow. Queues are used to decouple data waiting dependencies between different worker threads, enabling parallel processing. For example, see Figure 6, which is a schematic diagram of task allocation according to an embodiment of this disclosure. In Figure 6, two GPUs execute the video processing method. GPU-0 executes video hardware decoding, human detection, automatic virtual camera movement, XR rendering, and video hardware encoding tasks, while GPU-1 executes portrait matting. The available resources of GPU-0 in Figure 6 cannot meet the resource usage of all processing tasks in the video processing method. By allocating the portrait matting task to GPU-1, the available resources of GPU-0 can meet the resource usage of other processing tasks, and the available resources of GPU-1 can meet the resource usage of the portrait matting task. Optionally, during the execution of the video processing method, if the resource usage of the portrait matting task decreases, and the reduced resource usage of the portrait matting task is less than the remaining resources of GPU-0, the portrait matting task can be allocated to GPU-0 for execution, thereby freeing up the computing power of GPU-1.
[0124] In Figure 6, four worker threads are set up in the cloud. The data flow between different processing tasks is controlled by multiple worker threads, so as to realize the transmission of data corresponding to multiple processing tasks between different GPUs and ensure the execution of video processing methods.
[0125] In this embodiment, by dynamically allocating multiple processing tasks during the execution of the video processing method, the computing resources of the GPU are fully utilized, thus avoiding waste of computing resources.
[0126] Figure 7 is a schematic diagram of a video processing device provided in an embodiment of the present disclosure. As shown in Figure 7, the device includes: a page display module 310, a video stream data processing module 320, and a video stream data display module 330.
[0127] Page display module 310 is used to display a video page, the video page including a first display area and a second display area;
[0128] The video stream data processing module 320 is used to acquire the original video stream data of the client, transmit the original video stream data to the cloud, and receive virtual scene video stream data fed back by the cloud. The virtual scene video stream data is obtained by the cloud replacing the original video stream data with a scene. The original video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene.
[0129] The video stream data display module 330 is used to display the original video stream data in the first display area; and to display the virtual scene video stream data in the second display area.
[0130] The technical solution provided in this disclosure acquires raw video stream data through an image acquisition component configured on the client side, replacing dedicated camera equipment and reducing equipment costs in the video stream data acquisition process. The client transmits the raw video stream data to the cloud, where the cloud performs scene replacement on the raw video stream data to obtain virtual scene video stream data, which is then transmitted back to the client for display. This eliminates the need to deploy dedicated hardware locally, further reducing equipment costs in the video processing process. Simultaneously, the cloud can process raw video stream data transmitted from different clients, lowering the cost and barrier to entry for a large number of users to implement virtual studio technology.
[0131] Optionally, based on the above embodiments, the device further includes:
[0132] An attribute adjustment module is used to adjust the shooting attribute information of the target object in the virtual scene in response to a first trigger operation in the second display area, wherein the shooting attribute information includes one or more of the following: position information, size information, and angle information.
[0133] Optionally, the attribute adjustment module is further configured to: determine, based on the first trigger operation, to send the adjusted shooting attribute information to the cloud, so that the cloud generates the virtual scene video stream data based on the adjusted shooting attribute information.
[0134] Optionally, the attribute adjustment module is further configured to: adjust the shooting attribute information of the target object in the virtual scene in the second display area in response to a setting action of the target object in the first display area, wherein the setting action is detected by the cloud and the shooting attribute information of the target object in the virtual scene is determined based on the setting action.
[0135] Optionally, based on the above embodiments, the device further includes:
[0136] A camera movement mode adjustment module is used to display virtual scene video stream data generated based on a target camera movement mode in the second display area in response to a second trigger operation on the video page, wherein the target camera movement mode is determined based on the second trigger operation.
[0137] Optionally, the camera movement mode adjustment module is also used for:
[0138] Based on the second trigger operation, camera movement mode switching information is generated and sent to the cloud, so that the cloud switches the camera movement mode in the virtual scene to the target camera movement mode based on the camera movement mode switching information, and generates the virtual scene video stream data based on the target camera movement mode.
[0139] Optionally, the target camera movement mode is the camera movement mode of a virtual camera in a virtual scene. During the camera movement, the virtual camera collects video stream data of the target object in the virtual scene to form virtual scene video stream data.
[0140] The target camera movement mode includes one or more of the following:
[0141] The virtual camera follows the actual camera in the client.
[0142] The virtual camera's follow mode for the target object;
[0143] The virtual camera uses an automatic camera movement mode with a set trajectory in the virtual scene.
[0144] The video processing apparatus provided in this disclosure can execute the video processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0145] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0146] Figure 8 is a schematic diagram of a video processing device provided in an embodiment of this disclosure. As shown in Figure 8, the device includes: a target object extraction module 410, a scene fusion module 420, and a video stream data generation module 430.
[0147] The target object extraction module 410 is used to receive raw video stream data transmitted by at least one client and extract the target object from the raw video stream data.
[0148] Scene fusion module 420 is used to fuse the target object into a virtual scene;
[0149] The video stream data generation module 430 is used to collect video stream data of the target object in the virtual scene based on the virtual camera, form virtual scene video stream data, and transmit the virtual scene video stream data to the client.
[0150] The technical solution provided in this disclosure extracts the target object from the raw video stream data transmitted by the client, merges the target object into a virtual scene, and captures video of the target object in the virtual scene using a virtual camera, obtaining virtual scene video stream data of the target object within the virtual scene, thus replacing the scene of the raw video stream data. The virtual scene video stream data is then transmitted to the client, converting the raw video stream data captured by the client into virtual scene video stream data of a specific virtual scene, assisting the client in implementing virtual studio technology and reducing the hardware requirements and equipment costs for implementing virtual studio technology on the client side.
[0151] Based on the above embodiments, optionally, the target object extraction module 410 is used to: input the video frames of the original video stream data into the matting model respectively to obtain the target object in each video frame; wherein, the target object is a human image; or, the target object is determined based on the object selection operation in the original video stream data performed by the client.
[0152] Optionally, based on the above embodiments, the device further includes:
[0153] The lighting and shadow setting module is used to set the lighting and shadow information in the virtual scene after the target object is merged into the virtual scene, and the lighting and shadow information is matched with the virtual scene.
[0154] Optionally, the raw video stream data is live video stream data; the raw video stream data sent by the client carries virtual room information;
[0155] The scene fusion module 420 is used to fuse target objects in different original video stream data carrying the same virtual room information into the same virtual scene;
[0156] The video stream data generation module 430 is used to collect video stream data of the target object corresponding to different clients in the virtual scene through different virtual cameras, and form virtual scene video stream data corresponding to each client.
[0157] Optionally, based on the above embodiments, the device further includes:
[0158] The effect rendering module is used to receive at least one first control information transmitted by the client and render virtual effect data in the virtual scene.
[0159] Optionally, based on the above embodiments, the device further includes:
[0160] The camera movement control module is used to receive second control information from any of the clients, switch the target camera movement mode of the virtual camera corresponding to the client, and control the virtual camera corresponding to the client to form virtual scene video stream data corresponding to the client based on the switched target camera movement mode;
[0161] The target camera movement mode includes one or more of the following: the virtual camera following the actual camera in the client; the virtual camera following the target object; and the virtual camera automatically moving along a trajectory set in the virtual scene.
[0162] Based on the above embodiments, optionally, the video stream data generation module 430 is further configured to:
[0163] Acquire the audio data corresponding to the original video stream data; during the process of the virtual camera acquiring the virtual scene video stream data, determine the distance between the virtual camera and the target object, adjust the volume of the audio data based on the distance, and obtain the adjusted audio data; merge the virtual scene video stream data and the adjusted audio data.
[0164] Optionally, based on the above embodiments, the device further includes:
[0165] The attribute adjustment module is used to receive third control information from the client, or, when a set action of the target object is detected, adjust the shooting attribute information of the target object in the virtual scene, wherein the shooting attribute information includes one or more of position information, size information and angle information.
[0166] Optionally, the virtual scene includes multiple sub-scenes;
[0167] The attribute adjustment module is also used to: switch the sub-scene where the target object is located, and control the virtual camera corresponding to the target object to collect video stream data of the target object in the switched sub-scene.
[0168] Based on the above embodiments, optionally, the video processing method is executed on at least one GPU; the video processing method includes multiple processing steps, each processing step corresponding to at least one processing task;
[0169] The device further includes a task allocation module, used to detect the resource usage of each processing task during the execution of the video processing method, and allocate each processing task to the at least one GPU based on the resource usage of each processing task.
[0170] The video processing apparatus provided in this disclosure can execute the video processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0171] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0172] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Referring now to Figure 9, a schematic diagram of the structure of an electronic device (e.g., the terminal device or server in Figure 9) 500 suitable for implementing embodiments of this disclosure is shown. The terminal device in embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 9 is merely an example and should not impose any limitations on the functionality and scope of use of embodiments of this disclosure.
[0173] As shown in Figure 9, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.
[0174] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG9 shows electronic device 500 with various devices, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0175] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0176] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0177] The electronic device provided in this embodiment and the video processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0178] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video processing method provided in the above embodiments.
[0179] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0180] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0181] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0182] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:
[0183] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the one or more programs, the electronic device causes the electronic device to: display a video page, the video page including a first display area and a second display area; acquire raw video stream data from the client, transmit the raw video stream data to the cloud, and receive virtual scene video stream data fed back from the cloud, wherein the virtual scene video stream data is obtained by the cloud by scene replacement of the raw video stream data; the raw video stream data is video stream data of a target object in a real scene, and the virtual scene video stream data is video stream data of the target object in a virtual scene; display the raw video stream data in the first display area; and display the virtual scene video stream data in the second display area.
[0184] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive raw video stream data transmitted by at least one client, extract a target object from the raw video stream data; merge the target object into a virtual scene; collect video stream data of the target object in the virtual scene based on a virtual camera, form virtual scene video stream data, and transmit the virtual scene video stream data to the client.
[0185] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0186] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0187] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0188] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0189] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0190] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0191] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0192] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video processing method, applied to a client, the method comprising: A video display page is provided, which includes a first display area and a second display area. The system acquires the original video stream data from the client, transmits the original video stream data to the cloud, and receives virtual scene video stream data from the cloud. The virtual scene video stream data is obtained by the cloud replacing the original video stream data with a scene. The original video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene. The original video stream data is displayed in the first display area; and the virtual scene video stream data is displayed in the second display area.
2. The method according to claim 1, further comprising: In response to a first trigger operation within the second display area, the shooting attribute information of the target object within the second display area in the virtual scene is adjusted, wherein the shooting attribute information includes one or more of the following: position information, size information, and angle information.
3. The method according to claim 2, further comprising: Based on the first triggering operation, it is determined that the adjusted shooting attribute information will be sent to the cloud, so that the cloud can generate the virtual scene video stream data based on the adjusted shooting attribute information.
4. The method according to claim 1, further comprising: In response to a set action of the target object in the first display area, the shooting attribute information of the target object in the virtual scene in the second display area is adjusted, wherein the set action is detected by the cloud and the shooting attribute information of the target object in the virtual scene is determined based on the set action.
5. The method according to any one of claims 1-4, further comprising: In response to a second triggering operation on the video page, virtual scene video stream data generated based on a target camera movement mode is displayed in the second display area, wherein the target camera movement mode is determined based on the second triggering operation.
6. The method according to claim 5, further comprising: Based on the second trigger operation, camera movement mode switching information is generated and sent to the cloud, so that the cloud switches the camera movement mode in the virtual scene to the target camera movement mode based on the camera movement mode switching information, and generates the virtual scene video stream data based on the target camera movement mode.
7. The method according to claim 5, wherein, The target camera movement mode is the camera movement mode of a virtual camera in a virtual scene. During the camera movement, the virtual camera collects video stream data of the target object in the virtual scene to form virtual scene video stream data. The target camera movement mode includes one or more of the following: The virtual camera follows the actual camera in the client. The virtual camera's follow mode for the target object; The virtual camera uses an automatic camera movement mode with a set trajectory in the virtual scene.
8. A video processing method applied in the cloud, the method comprising: Receive raw video stream data transmitted by at least one client, and extract the target object from the raw video stream data; The target object is then integrated into the virtual scene; The target object in the virtual scene is captured by a virtual camera to form virtual scene video stream data, which is then transmitted to the client.
9. The method according to claim 8, wherein, Extracting the target object from the original video stream data includes: The video frames of the original video stream data are input into the matting model to obtain the target objects in each video frame; The target object is a human image; or the target object is determined based on the client's object selection operation in the original video stream data.
10. The method according to claim 8 or 9, wherein, After integrating the target object into the virtual scene, the method further includes: The lighting and shadow information in the virtual scene is set, and the lighting and shadow information is matched with the virtual scene.
11. The method according to any one of claims 8-10, wherein, The original video stream data is live video stream data; The raw video stream data sent by the client carries virtual room information; Integrating the target object into a virtual scene includes: integrating target objects from different original video stream data carrying the same virtual room information into the same virtual scene; In addition, the virtual scene video stream data is formed by collecting video stream data of the target object in the virtual scene based on the virtual camera, including: collecting video stream data of the target object corresponding to different clients in the virtual scene through different virtual cameras, and forming virtual scene video stream data corresponding to each client.
12. The method according to any one of claims 8-11, further comprising: Receive first control information transmitted by at least one of the clients, and render virtual effect data in the virtual scene.
13. The method according to any one of claims 8-12, further comprising: Receive second control information from any of the clients, switch the target camera movement mode of the virtual camera corresponding to the client, and control the virtual camera corresponding to the client to form virtual scene video stream data corresponding to the client based on the switched target camera movement mode; The target camera movement mode includes one or more of the following: The virtual camera follows the actual camera in the client. The virtual camera's follow mode for the target object; The virtual camera uses an automatic camera movement mode with a set trajectory in the virtual scene.
14. The method of claim 13, further comprising: Obtain the audio data corresponding to the original video stream data; During the process of the virtual camera acquiring video stream data of the virtual scene, the distance between the virtual camera and the target object is determined, and the volume of the audio data is adjusted based on the distance to obtain the adjusted audio data; The virtual scene video stream data and the adjusted audio data are merged.
15. The method according to any one of claims 8-14, further comprising: The system receives third control information from the client, or, upon detecting a set action of the target object, adjusts the shooting attribute information of the target object in the virtual scene, wherein the shooting attribute information includes one or more of position information, size information, and angle information.
16. The method according to claim 15, wherein, The virtual scene includes multiple sub-scenes; The adjustment of the target object's shooting attribute information in the virtual scene includes: Switch the sub-scene where the target object is located, and control the virtual camera corresponding to the target object to collect video stream data of the target object in the switched sub-scene.
17. The method according to claim 8, wherein, The video processing method is executed on at least one GPU; the video processing method includes multiple processing steps, each processing step corresponding to at least one processing task. During the execution of the video processing method, the resource usage of each processing task is detected, and each processing task is allocated to the at least one GPU based on the resource usage of each processing task.
18. A video processing apparatus integrated into a client, the apparatus comprising: The page display module is configured to display a video page, which includes a first display area and a second display area. The video stream data processing module is configured to acquire the raw video stream data from the client, transmit the raw video stream data to the cloud, and receive virtual scene video stream data fed back from the cloud. The virtual scene video stream data is obtained by the cloud replacing the raw video stream data with a scene. The raw video stream data is the video stream data of the target object in a real scene, and the virtual scene video stream data is the video stream data of the target object in a virtual scene. The video stream data display module is configured to display the original video stream data in the first display area; and to display the virtual scene video stream data in the second display area.
19. A video processing apparatus integrated in the cloud, the apparatus comprising: The target object extraction module is configured to receive raw video stream data transmitted by at least one client and extract the target object from the raw video stream data; The scene fusion module is configured to fuse the target object into a virtual scene; The video stream data generation module is configured to collect video stream data of the target object in the virtual scene based on a virtual camera, form virtual scene video stream data, and transmit the virtual scene video stream data to the client.
20. An electronic device, comprising: One or more processors; Storage device, configured to store one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any one of claims 1-7, or the video processing method as described in any one of claims 8-17.
21. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the video processing method as claimed in any one of claims 1-7, or the video processing method as claimed in any one of claims 8-17.
Citation Information
Patent Citations
Complex background real-time alternating method based on background modeling and energy minimization
CN101777180A
Live broadcast picture processing method and device, electronic equipment and storage medium
CN116708851A
Video processing method and device, storage medium and electronic equipment
CN118870065A
Live broadcast interaction method, device, and system
WO2023142756A1