A video data processing method, device and storage medium
By responding to movement operations on the video live streaming interface of the virtual live streaming room, the live video is moved to the overlapping area and different data is displayed in the overlapping and missing areas, which solves the problem of the target object being occluded and realizes the flexible display and custom position adjustment of the target object.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-04-13
- Publication Date
- 2026-05-19
AI Technical Summary
In the video live streaming interface of a virtual live streaming room, when there is a lot of multimedia data information, the target object is easily obscured, resulting in poor display effect.
By responding to the user's movement, the live video is moved to the overlapping area on the live video interface, and different video data is displayed in the overlapping and missing areas respectively, adjusting the display position of the target object.
The display effect of target objects on the video live streaming interface has been optimized, allowing users to customize the display position and solving the problem of multimedia information obscuring target objects.
Smart Images

Figure CN115209167B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video data processing method, apparatus and storage medium. Background Technology
[0002] Currently, when a user (e.g., user A) watches a live video of a streamer in a virtual live streaming room, multimedia data information associated with the live video can be displayed on the video live streaming interface of the virtual live streaming room (e.g., widgets in the virtual live streaming room and text, pictures, etc. sent by other users). In this way, when there is a lot of multimedia data information on the video live streaming interface, some key information in the virtual live streaming room (e.g., the target object that user A is interested in) will inevitably be obscured.
[0003] For example, when there is a lot of multimedia data on the live video interface, the target object that user A is interested in (e.g., the virtual item X introduced by the host) will be obscured on the live video interface, making it difficult for user A to see the video of the virtual item X clearly and intuitively. This means that the existing target object display scheme is extremely monotonous, which reduces the display effect of the target object on the live video interface. Summary of the Invention
[0004] This application provides a video data processing method, apparatus, and storage medium, which can flexibly optimize the display effect of target objects in a live video interface.
[0005] One embodiment of this application provides a video data processing method, the method comprising:
[0006] Output the first live video associated with the target object on the video live streaming interface of the virtual live streaming room;
[0007] In response to a movement operation on the first live video, the first live video is moved to the first display area of the live display area corresponding to the live video interface, and a second display area to be filled is determined on the live video interface; the first display area is the overlapping area between the video display area of the moved first live video and the live display area; the second display area is the area in the live display area excluding the overlapping area.
[0008] Based on the first live video after the movement, a second live video consisting of first video data and second video data is determined. When the second live video is played on the live video interface, the first video data is displayed in the first display area, and the second video data is displayed in the second display area. The display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0009] One embodiment of this application provides a video data processing apparatus, the apparatus comprising:
[0010] The first video acquisition module is used to output the first live video associated with the target object on the video live streaming interface of the virtual live streaming room.
[0011] The video movement module is used to respond to the movement operation of the first live video, move the first live video to the first display area of the live display area corresponding to the video live interface, and determine the second display area to be filled on the video live interface; the first display area is the overlapping area between the video display area of the moved first live video and the live display area; the second display area is the remaining area of the live display area excluding the overlapping area.
[0012] The video display module is used to determine a second live video composed of first video data and second video data based on the first live video after the movement. When the second live video is played on the video live interface, the first video data is displayed in the first display area, and the second video data is displayed in the second display area. The display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0013] The first video acquisition module includes:
[0014] The application interface output unit is used to respond to the launch operation for the application client and output the application display interface corresponding to the application client.
[0015] The access request sending unit is used to send an access request for the virtual live room to the server in response to a trigger operation on the virtual live room in the application display interface; the access request is used to instruct the server to extract a first video sequence from the captured video sequence corresponding to the original captured video stream that matches the terminal screen information of the application client; the original captured video stream is the video stream that the server pulls from the broadcaster client and is associated with the broadcaster user in the virtual live room; the size of the video frames in the first video sequence is less than or equal to the size of the video frames in the captured video sequence;
[0016] The first video stream receiving unit is used to receive the first video stream corresponding to the first video sequence returned by the server, and to decode the first video stream to obtain the first live video of the target object associated with the broadcaster user.
[0017] The first video output unit is used to output the first live video on the video live streaming interface corresponding to the virtual live streaming room.
[0018] Wherein, when the first video acquisition module performs the step of outputting the first live video associated with the target object on the video live streaming interface of the virtual live streaming room, the apparatus further includes:
[0019] The video element output module is used to output auxiliary video elements associated with the first live video on the live video interface.
[0020] The video live streaming interface includes a first business layer and a second business layer; the first business layer is used to display the first live video played in the virtual live streaming room; the second business layer is used to display public screen message elements, interactive elements and operation elements associated with the virtual live streaming room.
[0021] The video element output module includes:
[0022] The area determination unit is used to determine the public screen message area, business interaction area, and content operation area associated with the second business level on the video live streaming interface;
[0023] The broadcast message output unit is used to output public broadcast messages associated with the target object in the virtual live broadcast room in the public screen message area, and to use the public broadcast messages as public screen message elements;
[0024] The element information output unit is used to output interactive message elements associated with the anchor users in the virtual live broadcast room in the business interaction area, and to output operation elements associated with the audience users in the virtual live broadcast room in the content operation area.
[0025] The auxiliary element determination unit is used to identify public screen message elements, interactive elements, and operation elements as video auxiliary elements associated with the first live video.
[0026] The device further includes, prior to the video motion module performing a step in response to a motion operation on the first live video, the following:
[0027] The drag condition acquisition module is used to acquire the video drag conditions corresponding to the live video interface, and to detect the video auxiliary elements at the second business level based on the video drag conditions.
[0028] The instruction information output module is used to output movement instruction information for the first live video at the first business level on the video live streaming interface when the number of video auxiliary elements is detected to meet the quantity threshold in the video dragging condition; the movement instruction information is used to instruct the audience users in the virtual live streaming room to perform movement operations on the first live video.
[0029] The video motion module includes:
[0030] A motion operation response unit is used to respond to a motion operation on the first live video, take the video frame corresponding to the motion operation as the motion video frame, and determine the motion duration corresponding to the motion operation; the motion duration includes the motion start time and the motion end time; the motion video frame corresponding to the motion start time is the first video frame, and the video frame corresponding to the motion end time is the second video frame.
[0031] The display area changing unit is used to change the video display area of the first live video from the first image area to the second image area during the movement duration; the first image area is the video display area of the first video frame at the start of the movement; the second image area is the video display area of the second video frame at the end of the movement.
[0032] The overlapping area determination unit is used to determine the overlapping area between the second image area and the live display area corresponding to the video live interface, and to determine the overlapping area as the first display area corresponding to the first live video after the movement.
[0033] The missing area determination unit is used to identify the remaining video missing areas (excluding overlapping areas) in the live broadcast display area as the second display area to be filled.
[0034] The mobile operation response unit includes:
[0035] The instruction information triggering unit is used to obtain the motion instruction information corresponding to the first live video, respond to the triggering operation for the motion execution information, determine the video display area to which the first live video belongs at the first business level of the video live interface, and adjust the business status of the video display area to which the first live video belongs to the movable state.
[0036] The video motion unit is configured to, within a video display area having a movable state, in response to a motion operation on a first live video, determine the video frame corresponding to the motion operation as a motion video frame in the first live video, and determine the motion duration corresponding to the motion operation based on the motion start time corresponding to the first video frame in the motion video frame and the motion end time corresponding to the second video frame in the motion video frame.
[0037] The video display module includes:
[0038] The starting position recording unit is used to obtain the moving video frame corresponding to the moving operation from the first live video, and record the starting position information of the first image area corresponding to the first video frame in the screen coordinate system to which the live display area belongs, based on the first image area of the first video frame at the start time of the moving operation.
[0039] The end position recording unit is used to record the end position information of the second image area corresponding to the second video frame in the screen coordinate system to which the live display area belongs, based on the second image area of the second video frame in the moving video frame at the end of the moving operation.
[0040] The location sending unit is used to send start position information and end position information to the server, so that when the server determines the movement displacement of the first live video based on the start position information and end position information, it can determine the second video sequence associated with the target object from the original captured video stream based on the movement displacement and the positioning information of the moving video frame in the captured video frame; the captured video frame is a video frame with the same video frame number as the moving video frame in the captured video sequence cached by the server, and the moving video frame is obtained by the server taking a screenshot of the captured video frame based on the terminal screen information;
[0041] The second video stream receiving unit is used to receive the second video stream corresponding to the second video sequence returned by the server, decode the second video stream to obtain the second live video corresponding to the second video sequence, and when playing the second live video on the video live streaming interface, display the first video data in the second live video in the first display area and the second video data in the second live video in the second display area.
[0042] The first video stream is obtained by the server encoding the first video sequence using the first encoder; the frame number of the video frames in the first video sequence is consistent with the frame number of the video frames in the acquired video sequence; the device further includes:
[0043] The switching request sending module is used to generate an encoding switching request for the virtual live broadcast room based on the start position information and the end position information, and send the encoding switching request to the server. The encoding switching request is used to instruct the server to encode the second video sequence through the second encoder when the second encoder is started, so as to obtain the second video stream corresponding to the second video sequence. The video frames in the second video sequence are determined by the server after taking screenshots of the video frames in the acquired video sequence based on the terminal screen information and the positioning position information, and the display position of the target object in the second video sequence is different from the display position of the target object in the first video sequence.
[0044] One embodiment of this application provides a video data processing method, the method comprising:
[0045] The application client receives an access request for accessing the virtual live room based on the terminal screen information. Based on the access request, the application client determines the first video stream that matches the terminal screen information from the original captured video stream corresponding to the virtual live room. The application client returns the first video stream to the application client so that the application client can output the first live video associated with the target object on the video live streaming interface of the virtual live room based on the first video stream.
[0046] When the application client moves the first live video to the first display area of the live display area corresponding to the live video interface, it receives the video coordinate position information sent by the application client based on the moved first live video, and determines the movement displacement of the first live video based on the video coordinate position information; the first display area is the overlapping area between the video display area of the moved first live video and the live display area.
[0047] Based on the displacement, a second video stream associated with the target object is determined from the original acquired video stream. The second video stream is returned to the application client so that when the application client obtains the second live video based on the second video stream, it outputs the first video data in the second live video in the first display area and the second video data in the second live video in the second display area. The second display area is the area in the live display area excluding the overlapping area. The display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0048] One embodiment of this application provides a video data processing apparatus, the apparatus comprising:
[0049] The first video return module is used to receive the access request sent by the application client based on the terminal screen information for accessing the virtual live room, and to determine the first video sequence that matches the terminal screen information from the original captured video stream corresponding to the virtual live room based on the access request, and to return the video stream corresponding to the first video sequence to the application client so that the application client can output the first live video associated with the target object on the video live interface of the virtual live room based on the video stream corresponding to the first video sequence.
[0050] The displacement determination module is used to receive video coordinate position information sent by the application client based on the moved first live video when the application client moves the first live video to the first display area of the live display area corresponding to the video live interface, and to determine the displacement of the first live video based on the video coordinate position information; the first display area is the overlapping area between the video display area of the moved first live video and the live display area.
[0051] The second video determination module is used to determine the second video stream associated with the target object from the original acquired video stream based on the movement displacement, and return the second video stream to the application client. This allows the application client to output the first video data in the second live video in the first display area and the second video data in the second live video in the second display area when it obtains the second live video based on the second video stream. The second display area is the area in the live display area excluding the overlapping area. The display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0052] The displacement determination module includes:
[0053] The coordinate position receiving unit is used to receive video coordinate position information sent by the application client based on the first live video after movement, and to determine the start position information and end position information associated with the moving video frame in the first live video based on the video coordinate position information; the moving video frame is the video frame corresponding to the movement operation determined by the application client in response to the movement operation on the first live video; the start position information is the coordinate position information of the first video frame in the moving video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the first image area at the start time of the movement operation; the end position information is the coordinate position information of the second video frame in the moving video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the second image area at the end time of the movement operation.
[0054] The displacement determination unit is used to determine the coordinate position change information of the video display area to which the first live video belongs, from the first image area to the second image area, based on the start position information and the end position information, and to determine the displacement of the first live video based on the coordinate position change information.
[0055] The second video determination module includes:
[0056] The positioning and location determination unit is used to determine the positioning information of the first live video after movement in the original acquired video stream based on the movement displacement, and to determine the second video sequence associated with the target object from the original acquired video stream based on the positioning and location information and terminal screen information.
[0057] The video sequence encoding unit is used to encode the second video sequence to obtain the second video stream corresponding to the second video sequence, and return the second video stream to the application client.
[0058] The positioning and location determination unit includes:
[0059] The acquisition sequence acquisition subunit is used to acquire the acquisition video sequence corresponding to the original acquisition video stream. Based on the vertex position information of the video frames in the acquisition video sequence in the positioning coordinate system, the mobile safety area corresponding to the original acquisition video stream is determined. In the acquisition video sequence, the video frame with the same video frame number as the mobile video frame is taken as the acquisition video frame.
[0060] The positioning determination subunit is used to determine the start and end coordinates of the moving video frame in the captured video frame based on the displacement in the positioning coordinate system. The end coordinates are used as the positioning position information of the first live video after movement in the original captured video stream. The positioning position information is compared with the movement safety area to obtain the comparison result. The image area formed by the start coordinates is the first positioning area of the first video frame in the moving video frame in the positioning coordinate system, and the image area formed by the end coordinates is the second positioning area of the second video frame in the moving video frame in the positioning coordinate system. The overlapping area between the second positioning area and the first positioning area is the first screenshot area, and the remaining area in the first positioning area excluding the overlapping area is the second screenshot area.
[0061] The safe zone determination sub-unit is used to determine that the first live video after the movement is located within the safe zone if the comparison result indicates that each movement end coordinate information in the positioning information is within the safe zone.
[0062] The video sequence synthesis subunit is used to, based on terminal screen information and a first screenshot method, take video data belonging to the first screenshot area in the captured video sequence as first video data to be displayed in the first display area, and take video data belonging to the second screenshot area in the captured video sequence as second video data to be displayed in the second display area, synthesize the first video data and the second video data, and use the synthesized video data as the second video sequence associated with the target object.
[0063] The positioning and location determination unit also includes:
[0064] The non-safe area determination subunit is used to determine that if the comparison result indicates that each movement end coordinate information in the positioning information has movement end coordinate information outside the movement safe area, then the first live video after movement has a local video area located outside the movement safe area; the local video area is the area in the second positioning area composed of movement frame position information outside the movement safe area.
[0065] The local video filling subunit is used to fill the image data in the local video area into the second screenshot area based on the terminal screen information and the second screenshot method. In the captured video sequence, the video data belonging to the first screenshot area is used as the first video data to be displayed in the first display area, and the video data in the second screenshot area is used as the second video data to be displayed in the second display. The first video data and the second video data are combined and processed, and the combined video data is used as the second video sequence associated with the target object.
[0066] The first video stream is obtained by encoding the first video sequence using the first encoder; the video frames in the first video sequence are determined by taking screenshots of the video frames in the captured video sequence corresponding to the original captured video stream based on the terminal screen information; the video frame number of the video frames in the first video sequence is consistent with the video frame number of the video frames in the captured video sequence.
[0067] The device also includes:
[0068] The encoder startup module is used to keep the first encoder running and start the second encoder when it receives an encoding switch request sent by the application client based on the first live video after the move; the second encoder is used to encode the second video stream corresponding to the second video sequence when the second video sequence associated with the target object is obtained.
[0069] The device also includes:
[0070] The frame number synchronization module synchronizes the frame numbers of video frames with the same video frame content in the first video sequence corresponding to the first video stream and the second video sequence corresponding to the second video stream, and uses the video frames after frame number synchronization as buffered video frames in the first video sequence.
[0071] The keyframe determination module is used to determine the video frame corresponding to the motion operation in the first video sequence as the motion video frame in the first live video, and to determine the next video frame of the motion video frame as the key video frame in the second live video in the second video sequence; the frame spacing between the video frame number of the key video frame and the video frame number of the target video frame in the buffered video frames satisfies the video encoding conditions of the second video stream.
[0072] The encoder shutdown module is used to shut down the first encoder based on a key video frame when it is detected that the application client has finished playing the moving video frame based on the target video frame.
[0073] One embodiment of this application provides a computer device, which includes a processor and a memory;
[0074] The processor is connected to a memory, wherein the memory is used to store computer programs, and the processor is used to invoke the computer programs so that the computer device performs the methods in any aspect of the embodiments of this application.
[0075] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, such that a computer device having a processor performs the method in any aspect of this application.
[0076] One aspect of this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method in any aspect of this application.
[0077] In this embodiment of the application, when outputting a first live video associated with a target object on the video live streaming interface of a virtual live streaming room, the first live video can be moved in response to a viewer's movement operation on the first live video. This allows the first live video to be moved to a first display area within the corresponding live streaming display area of the video live streaming interface, and a second display area to be filled is determined on the video live streaming interface. It should be understood that the first display area is the overlapping area between the video display area of the moved first live video and the live streaming display area; the second display area is the area within the live streaming display area excluding the overlapping area. Furthermore, this embodiment of the application can determine a second live video composed of first video data and second video data based on the moved first live video. Therefore, when playing the second live video on the video playback interface, the first video data is displayed in the first display area, and the second video data is displayed in the second display area. It is understood that the display position of the target object in the second live video differs from its display position in the first live video. Thus, this embodiment of the application can refer to the live video before the movement as the first live video, and the live video obtained after moving the first live video as the second live video. It should be understood that the second live video is the live video obtained after adjusting the display position of the target object on the live video interface. In other words, the embodiments of this application provide a solution that can flexibly control the display position of the target object on the live video interface. This solution allows viewers in the virtual live room to customize the display position of the target object while watching the live video (i.e., the aforementioned first live video). This solves the problem of multimedia information obscuring the target object on the live video interface in a live streaming scenario, thereby fundamentally optimizing the display effect of the target object on the live video interface. Attached Figure Description
[0078] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0079] Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of the present invention;
[0080] Figure 2 This is a schematic diagram illustrating a scenario of adjusting the display position of a target object in a live streaming environment, provided by an embodiment of this application.
[0081] Figure 3 This is a flowchart illustrating a video data processing method provided in an embodiment of this application;
[0082] Figure 4 This is a schematic diagram of an application display interface provided in an embodiment of this application;
[0083] Figure 5 This is a schematic diagram of a scenario for capturing video sequences based on terminal screen information, provided in an embodiment of this application;
[0084] Figure 6 This is a schematic diagram of a mobile first live video scenario provided in an embodiment of this application;
[0085] Figure 7 This is a schematic diagram of a scenario in which video coordinate location information is sent according to an embodiment of this application;
[0086] Figure 8 This is a schematic diagram of a video data processing method provided in an embodiment of this application;
[0087] Figure 9 This is a schematic diagram of a scenario involving two business layers in a video live streaming interface, provided by an embodiment of this application.
[0088] Figure 10 This is a schematic diagram of a scenario where movement indication information is output on a live video interface, as provided in an embodiment of this application.
[0089] Figure 11 This is an interactive schematic diagram of a video data processing method provided in an embodiment of this application;
[0090] Figure 12 This is a schematic diagram of a scenario where missing frames are automatically filled in within a safe area, as provided in an embodiment of this application.
[0091] Figure 13 This is a schematic diagram illustrating a scenario where missing frames are automatically filled in within a non-safe area, as provided in an embodiment of this application.
[0092] Figure 14 This is a schematic diagram of the structure of a video data processing device provided in an embodiment of this application;
[0093] Figure 15 This is a schematic diagram of the structure of a video data processing device provided in an embodiment of this application;
[0094] Figure 16 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0095] Figure 17 This is a schematic diagram of the structure of a video data processing system provided in an embodiment of this application. Detailed Implementation
[0096] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0097] For further details, please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of the present invention. Figure 1 As shown, the network architecture can be applied to data processing systems in live streaming scenarios. This data processing system can specifically include... Figure 1 The diagram shows server 2000, a broadcast terminal cluster, and a viewer terminal cluster. The broadcast terminal cluster can specifically include one or more broadcast terminals; the number of broadcast terminals in the cluster is not limited here. Figure 1 As shown, multiple broadcast terminals may specifically include broadcast terminal 3000a, broadcast terminal 3000b, broadcast terminal 3000c, ..., broadcast terminal 3000n; as... Figure 1 As shown, broadcast terminal 3000a, broadcast terminal 3000b, broadcast terminal 3000c, ..., broadcast terminal 3000n can each connect to server 2000 via the network so that each broadcast terminal can interact with server 2000 through the network connection.
[0098] For example, when a broadcaster's terminal is running a broadcaster client, and the broadcaster uses this client to record live video in a virtual live streaming room, the broadcaster's terminal can push (i.e., upload) the recorded live video stream to server 2000 in real time. This allows server 2000 to decode the received video stream and cache the resulting video sequence, which can then be distributed to the viewer's terminal within the virtual live streaming room. It should be understood that for different viewer terminals, the live video stream is determined by server 2000 based on the screen information of each viewer's terminal within the virtual live streaming room. This screen information may include, but is not limited to, the screen size and resolution. In other words, the server 2000 here needs to adapt the captured video sequence corresponding to the captured video stream according to the terminal screen information of different viewer terminals, so as to return the adapted live video stream to these viewer terminals, so that after these viewer terminals decode the received live video stream, they can display live images of different sizes or resolutions on their own live video interface.
[0099] The audience terminal cluster can specifically include one or more audience terminals; there is no limit to the number of audience terminals in the audience terminal cluster. For example... Figure 1 As shown, the multiple audience terminals may specifically include audience terminal 4000a, audience terminal 4000b, audience terminal 4000c, ..., audience terminal 4000n; as Figure 1 As shown, viewer terminals 4000a, 4000b, 4000c, ..., 4000n can each connect to server 2000 via a network, allowing each viewer terminal running a viewer client to interact with server 2000 through this network connection. For example, when these viewer terminals play live video in a virtual live streaming room, they can receive live video streams associated with the broadcaster users located in the same virtual live streaming room, sent by server 2000 in real time. It is understood that the viewer client here refers to the live streaming client running on the viewer terminal, and the broadcaster client refers to another live streaming client running on the broadcaster terminal. For ease of understanding, in this embodiment, the live streaming clients accessed by viewer users can be collectively referred to as application clients, and the live streaming clients accessed by broadcaster users can be collectively referred to as broadcaster clients.
[0100] Among them, such as Figure 1The server 2000 shown can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0101] For ease of understanding, the embodiments of this application may include... Figure 1 From the audience terminal cluster shown, one audience terminal is selected as the target audience terminal. For example, in the embodiments of this application, an audience terminal can be selected as the target audience terminal. Figure 1 The audience terminal 4000a shown serves as the target audience terminal, which may integrate an application client with video data processing capabilities (e.g., video data loading and playback). Specifically, the application client may include live streaming clients such as video streaming clients, game streaming clients, and online education clients, which have frame sequence (e.g., frame animation sequence) loading and playback capabilities.
[0102] Specifically, the target audience terminal (e.g., audience terminal 4000a) may include: smartphones, tablets, laptops, desktop computers, wearable devices, smart home devices (e.g., smart TVs) and other smart terminals with video data processing functions (e.g., video data playback functions).
[0103] For ease of understanding, in this application embodiment, the video live streaming client A running on the target viewer's terminal by a user (e.g., viewer user B) can be collectively referred to as the aforementioned application client, and the streamer user selected by viewer user B in the application client's display interface that matches their interests can be collectively referred to as the target streamer user, and the virtual room created by the target streamer user can be collectively referred to as a virtual live streaming room. Furthermore, in this application embodiment, the video data displayed on the video live streaming interface of the virtual live streaming room can also be collectively referred to as live video.
[0104] It is understood that, in this embodiment of the application, the video live streaming window used to display live video on the video live streaming interface can be collectively referred to as the video display area, and the terminal screen window corresponding to the video live streaming interface can be collectively referred to as the live streaming display area. It is understood that, when viewer user B performs a drag operation on the live video on the video display interface in the aforementioned target viewer terminal, the target viewer terminal can respond to the drag operation on the live video, thereby allowing viewer user B to drag the live video to the first display area within the live streaming display area. It should be understood that, during the process of viewer user B dragging the live video on the video display interface, a second display area to be filled can be displayed on the video live streaming interface, thereby simultaneously displaying second video data in the second display area while the first video data is displayed in the aforementioned first display area. It should be understood that both the first video data and the second video data here originate from the moved live video, and the display position of the target object in the moved live video is different from the display position of the target object in the live video before the move. To facilitate the distinction between the live stream videos before and after the move, this application embodiment may refer to the live stream videos displayed on the live stream interface before the move as the first live stream video, and the live stream videos displayed on the live stream interface after the move as the second live stream video.
[0105] It should be understood that during the process of viewer user B dragging the live video on the video display interface, the second display area here refers to the missing video area within the live video display area. For example, this missing video area could be used to display a black background as the live video is moved during dragging. Optionally, this missing video area could also be used to display a white background as the live video is moved during dragging. The default background color displayed in this missing video area will not be limited here.
[0106] For further information, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram illustrating a scenario of adjusting the display position of a target object in a live streaming environment, as provided in an embodiment of this application. Figure 2 As shown, audience users (e.g., Figure 2 User B) shown is on the video live streaming interface of the virtual live streaming room (e.g., Figure 2 When watching the first live video (i.e., the live video before dragging) on the live streaming interface 100a shown, the display position of the target object on the live streaming interface 100a can be flexibly customized. It should be understood that the target object here can be any object that user B is interested in on the video live streaming interface corresponding to the virtual live streaming room. This object can include, but is not limited to, the broadcaster user within the virtual live streaming room (e.g., [example of user]). Figure 2The items associated with the streamer A shown (e.g., streamer A) and the streamer user. Figure 2 (The product that anchor A is currently introducing is shown).
[0107] Wherein, if the target object that user B is interested in in the virtual live broadcast room is a product introduced by the currently live broadcast host A (i.e., the aforementioned target host user), for example, the product (i.e., the target object) introduced by host A in the virtual live broadcast room is... Figure 2 The "Strawberry Mousse Cake" shown is located within the current product purchase link area. Therefore, user B, while watching the first live stream video on the live stream interface 100a via their viewer terminal, can customize the display position of this product within the virtual live stream room.
[0108] For example, such as Figure 2 As shown, user B is along Figure 2 If the arrow shown (e.g., the first direction) is dragged upwards on the first live video and the user does not release the arrow, the video can be simultaneously output to the viewer's terminal screen. Figure 2 The live stream interface shown is 100b. For ease of understanding, it is referred to here as... Figure 2 The live streaming interface 100b shown is an example of the live streaming screen when the video drag ends (i.e., when the hand is about to be released), which is used to illustrate the first display area and the second display area used to compose the live streaming interface 100b.
[0109] Specifically, such as Figure 2 As shown, when user B triggers the video live streaming interface corresponding to the virtual live streaming room and drags the first live streaming video upwards, the dragging operation performed by user B on the video live streaming interface for the first live streaming video can be collectively referred to as a movement operation. Thus, when the viewer terminal corresponding to user B responds to the movement operation for the first live streaming video, it can... Figure 2 The live streaming interface shown in 100b is as follows Figure 2 The first display area and the second display area to be filled are shown. For example... Figure 2 The first display area shown can be the live streaming client corresponding to the virtual live streaming room. Based on the overlapping area between the video display area of the moved first live streaming video and the live streaming display area corresponding to the video live streaming interface, when the movement of the first live streaming video is completed (i.e., dragging the first live streaming video upwards along the first direction and releasing the drag), this overlapping area can be used for display. Figure 2 The video data 200a may contain the target object that user B is interested in. It is understood that, in this embodiment of the application, the video data 200a containing the target object that user B is interested in can be collectively referred to as the first video data.
[0110] In addition, specifically, such as Figure 2The second display area shown can be the area outside the overlapping area of the live broadcast display area (i.e., the missing video area). When moving the first live video along the first direction (i.e., dragging upwards), this missing video area (i.e., the second display area) can be used to fill in and display the video data supplemented based on the aforementioned first video data. For example, it can display... Figure 2 The video data 200b is determined based on the second live video to which the first video data in the first display area belongs. That is, the video data displayed in the second display area is associated with the target object that user B is interested in. It is understood that, in this embodiment, the video data 200b associated with the target object that user B is interested in can be collectively referred to as the second video data.
[0111] It should be understood that, Figure 2 As shown in the example where user B drags the first live video, this embodiment of the application can collectively refer to the overlapping area between the video display area and the live broadcast display area of the currently dragged first live video as the overlapping area. Thus, this embodiment of the application can further refer to the area in the live broadcast display area excluding the overlapping area as the video missing area. It should be understood that during the dragging of the first live video, the video missing area can be used to display the background color of the gap created when dragging the first live video, for example, Figure 2 The background color displayed in the second display area of the live broadcast interface 100b can be black.
[0112] It is understood that, in this application embodiment, the direction of dragging the video display area of the first live video upward along the +Y axis can be collectively referred to as the first direction in the screen coordinate system, and the direction of dragging the video display area of the first live video to the right along the +X axis can be collectively referred to as the second direction in the same screen coordinate system. Optionally, in this application embodiment, the direction of dragging the video display area of the first live video downward along the -Y axis can also be collectively referred to as the third direction in the screen coordinate system, and the direction of dragging the video display area of the first live video to the right along the -X axis can be collectively referred to as the fourth direction.
[0113] like Figure 2As shown, during the dragging of the first live video, the live streaming client corresponding to the virtual live streaming room can record the video coordinate position information of the first live video before and after the drag. For example, the live streaming client running on the viewer's terminal (i.e., the aforementioned application client) can record the video coordinate position information of the first live video from the start time to the end time of the drag within the movement duration (also known as the drag duration) in the screen coordinate system. The recorded video coordinate position information of the video display area to which the first live video belongs at the start time of the drag can be collectively referred to as the start position information, and the recorded video coordinate position information of the video display area to which the first live video belongs at the end time of the drag can be collectively referred to as the end position information. In this way, the live streaming client (i.e., the aforementioned application client) can send the recorded start and end position information to the server corresponding to the live streaming client, so that the server can determine the movement displacement of the dragged first live video based on the received start and end position information, and then, based on the determined movement displacement, move the video from the screen to the end time of the drag. Figure 2 From the original captured video stream associated with anchor A, a second video stream associated with the target object is determined. Furthermore, the server can return the second video stream to the live streaming client (i.e., the aforementioned application client), so that the live streaming client (i.e., the aforementioned application client) can output the first video data (i.e., the first video data) from the second live video in the first display area based on the second video stream. Figure 2 The video data 200a), and output the second video data in the second live video in the second display area (i.e., the .... Figure 2 (Video data 200b). This should be understood. Figure 2 The target object shown is displayed in a different position in the second live video than it was displayed in the first live video.
[0114] Optionally, if the target audience that user B is following in the virtual live stream is the currently streaming host A, then user B can customize the display position of host A within the virtual live stream on the video live stream interface corresponding to the viewer terminal used by user B while watching the first live stream video. That is, when user B finds that the host A they are following is obscured by video auxiliary elements on the video live stream interface (such as announcement information and interactive information in the virtual live stream), by flexibly adjusting the display position of the host A (i.e., another target audience) on the video live stream interface, user B can clearly see their favorite host A on the video live stream interface.
[0115] The specific implementation of the application client moving the first live video to the first display area, displaying the first video data in the first display area, and displaying the second video data in the second display area can be found below. Figures 3-13 The corresponding implementation example.
[0116] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a video data processing method provided in an embodiment of this application. Figure 3 As shown, this method can be executed by the audience terminal corresponding to the aforementioned user B (i.e., the aforementioned target audience terminal), or by the service server (e.g., the aforementioned...). Figure 1 The method can be executed by the server 2000 shown, or jointly by the viewer terminal and the service server. For ease of understanding, this embodiment uses the example of the method being executed by the viewer terminal (i.e., the target viewer terminal) corresponding to user B to illustrate the specific process of moving the first live video in the target viewer terminal, displaying the first video data in the first display area, and displaying the second video data in the second display area. The method may include at least the following steps S101-S103:
[0117] Step S101: Output the first live video associated with the target object on the video live streaming interface of the virtual live streaming room;
[0118] Specifically, the target viewer terminal can respond to the launch operation of the application client by outputting the application display interface corresponding to the application client; further, the target viewer terminal can respond to the trigger operation of the virtual live room in the application display interface by sending an access request to the server for the virtual live room; the access request is used to instruct the server to extract a first video sequence from the captured video sequence corresponding to the original captured video stream that matches the terminal screen information of the application client; the original captured video stream is the video stream associated with the broadcaster user in the virtual live room, which is pulled by the server from the broadcaster client; the size of the video frames in the first video sequence is less than or equal to the size of the video frames in the captured video sequence; further, the target viewer terminal can receive the first video stream corresponding to the first video sequence returned by the server, decode the first video stream, and obtain the first live video of the target object associated with the broadcaster user; further, the target viewer terminal can output the first live video on the video live interface corresponding to the virtual live room.
[0119] It can be understood that the application client here can be a live streaming client with audio and video playback capabilities running on the target viewer's terminal. When a viewer (e.g., user B) performs a startup operation on this application client, the application display interface corresponding to that application client can be output. This application display interface can be used to display the application client from the server (e.g., the one mentioned above). Figure 1 The server 2000 in the corresponding embodiment pulls multiple video data that match the interests of the viewer (e.g., user B), and each video data corresponds to a virtual room created by a broadcaster.
[0120] For further information, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic diagram of an application display interface provided in an embodiment of this application. Specifically, in a live streaming scenario... Figure 4 User B, as shown, can start... Figure 4 When the application client in the target viewer's terminal is shown, authorize user B to access it. Figure 4 The viewer user information 10a shown accesses the application client running on the target viewer's terminal. Then, user B can output [information] on the target viewer's terminal. Figure 4 The application display interface 300a is shown. It can be understood that, as... Figure 4 The application display interface 300a shown can be used to display the cover data of at least one virtual room that matches the interests of user B. The cover data can be the image data of the corresponding broadcaster user or the live video of the corresponding broadcaster user in each virtual room. The specific display format of the cover data of each virtual room in the application display interface 300a will not be limited here.
[0121] For example, in Figure 4 The cover data displayed in each virtual room in the application display interface 300a shown can be the live video recorded by the corresponding broadcaster. Based on this, the data displayed in... Figure 4 The live video in each virtual room in the application display interface 300a shown can contain Figure 4 The live videos 20a, 20b, 20c, and 20d shown are illustrated. It is understood that, in this embodiment of the application, the live videos 20a, 20b, 20c, and 20d dynamically played in each virtual room within the application display interface 300a can be collectively referred to as recommended live videos.
[0122] It is understandable that by dynamically playing corresponding recommended live videos in various virtual rooms on the application's display interface 300a, it can help... Figure 4As shown, User B can quickly learn about the live video content being streamed by the host in the current virtual room, and can then quickly select a virtual room that best suits their interests based on this synchronously playing video content. It is understood that, in this embodiment of the application, the virtual room selected by User B in the application's display interface can be collectively referred to as a virtual live streaming room. Thus, as... Figure 4 As shown, the target viewer's terminal can respond to user B's triggered operation (e.g., click operation) on the virtual live room and quickly access the corresponding video live interface (e.g., ...) Figure 4 The video live streaming interface 300b shown allows user B to quickly connect with the streamer in the virtual live streaming room (e.g., [example]) through the video live streaming interface 300b. Figure 4 The host A) is shown interacting with the audience during the live stream.
[0123] Among them, such as Figure 4 As shown, when user B needs to play a recommended live stream (e.g., live video 20b), the virtual room corresponding to the live video 20b selected by user B from the application display interface 300a can be used as the aforementioned virtual live stream room. Therefore, when accessing this virtual live stream room, the live video 20b played in the video live stream interface 300b of that virtual live stream room can be collectively referred to as... Figure 4 The first live stream video shown. Figure 4 The live streaming display area corresponding to the video live streaming interface 300b shown can be the screen display area corresponding to the terminal screen information of the target viewer's terminal.
[0124] What is understandable is that Figure 4 When the target viewer terminal, as shown, responds to a trigger operation for the virtual live room corresponding to video data 20b, it can send an access request for the virtual live room to the server based on its terminal screen information and the live room identifier. It is understood that this access request can instruct the server (e.g., server 2000 mentioned above) to pull a video stream matching the live room identifier from the broadcaster client in real time, and the video stream pulled from the broadcaster client (i.e., the video stream associated with the broadcaster user in the virtual live room) can be collectively referred to as the original captured video stream. It is understood that at this time, the server can decode the original captured video stream pulled from the broadcaster client according to the network protocol used for real-time data communication to obtain the captured video sequence corresponding to the original captured video stream.
[0125] Understandably, at this point, the server can extract a video sequence from the captured video sequence that matches the terminal screen information of the application client, and the extracted video sequence can be collectively referred to as the first video sequence. Understandably, the size of the video frames in this first video sequence can be less than or equal to the size of the video frames in the captured video sequence.
[0126] For further information, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating a scenario for capturing video sequences based on terminal screen information, provided in an embodiment of this application. For example... Figure 5 The target audience terminal shown can be the one described above. Figure 4 The target viewer terminal in the corresponding embodiment. Based on this, in the above... Figure 4 The first video data displayed in the live video interface 300b shown is determined by the target viewer terminal based on the first video stream returned by the server.
[0127] Among these, it is understandable that when Figure 5 The anchor A shown here passed through Figure 5 When the broadcaster terminal shown is conducting a live broadcast, it can create a live broadcast on the broadcaster terminal running the aforementioned broadcaster client. Figure 5 The virtual room K shown can be used by the broadcaster. Figure 5 The presenter A is shown. (For example...) Figure 5 As shown, the broadcaster client in the broadcaster terminal can push the video stream 400a from the virtual room K to the server based on the network protocol between the broadcaster and the server. It should be understood that, optionally, the server can also directly pull the video stream 400a associated with broadcaster A in the virtual room K from the broadcaster client based on the aforementioned network protocol. It is understood that in this embodiment, the pulled video stream 400a can be collectively referred to as the aforementioned original captured video stream.
[0128] like Figure 5 As shown, at this point, the server can further decode the original captured video stream (i.e., video stream 400a) to obtain the video sequence corresponding to video stream 400a. It is understood that in this embodiment, the video sequence corresponding to the original captured video stream can be collectively referred to as the captured video sequence. Furthermore, based on the terminal screen information (e.g., screen size and aspect ratio) of different viewers in the virtual room K, the size of the live stream image displayed on these viewers' terminals can be adapted to ensure that different live stream image sizes are presented on viewers' terminals corresponding to different terminal screen information. Figure 5 The target audience terminal shown can be a certain audience member in the virtual room K (e.g., Figure 5 The viewer terminal corresponding to user B shown.
[0129] like Figure 5 As shown, the server can obtain the captured video sequence corresponding to the original captured video stream (i.e. Figure 5 From the video sequence 400b shown, a video sequence that matches the terminal screen information of the target viewer's terminal is extracted (e.g., Figure 5 The video sequence 500b shown can be collectively referred to as the first video sequence. Figure 5 As shown, the size of the video frames in the first video sequence (i.e., video sequence 500b) can be smaller than or equal to the size of the video frames in the acquired video sequence (i.e., video sequence 400b). Here, the specific size of the target viewer's terminal screen information is not limited. Furthermore, the server can encode the first video sequence to obtain... Figure 5 The video stream shown is 500a. It is understood that embodiments of this application can use the video stream encoded by the server (e.g., Figure 5 The video stream 500a shown is collectively referred to as the first video stream, and this first video stream can be returned to... Figure 5 The target viewer terminal is shown. It is understood that this target viewer terminal receives... Figure 5 When the server returns the first video stream corresponding to the first video sequence, the first video stream can be decoded based on the network protocol between the target viewer's terminal and the server to obtain the first live video for playback in the aforementioned live video interface 300a. It can be understood that the first live video here refers to the video stream used by the broadcaster (e.g., the aforementioned...). Figure 4 The live video of the target object associated with the anchor A shown.
[0130] The network protocol mentioned here may include, but is not limited to, the Real-Time Messaging Protocol (RTMP). It should be understood that the RTMP protocol is a TCP-based protocol suite, which may specifically include the basic RTMP protocol and various variants such as RTMPT / RTMPS / RTMPE. Specifically, it can be understood that the RTMP protocol is a network protocol used for real-time data communication, enabling audio, video, and data communication between Flash / AIR platforms and streaming media / interactive servers that support the RTMP protocol. Furthermore, it can be understood that software supporting the RTMP protocol may include Adobe Media Server / Ultrant Media Server / red5, etc., and these software programs can be integrated into the aforementioned... Figure 5 In the corresponding embodiment, in the broadcaster terminal and the target audience client.
[0131] Optionally, it is understood that when the target viewer terminal performs step S101, it can also output video auxiliary elements associated with the first live video on the second service level of the video live streaming interface. Thus, when the target viewer terminal detects that the number of video auxiliary elements on the video live streaming interface reaches the video dragging condition, it can output movement instruction information for the first live video on the first service level on the video live streaming interface. The movement instruction information is used to instruct the viewer in the virtual live streaming room to move the first live video. At this time, the target viewer terminal can further execute steps S102 and S103, so that the target viewer terminal can ensure the integrity of the video auxiliary elements on the second service level without clearing them, and can flexibly adjust the display position of the target object in the virtual live streaming room. This can fundamentally solve the problem of the target object that the viewer is interested in being obscured on the first service level, thus preventing key information in the virtual live streaming room from being obscured, thereby helping the viewer improve the fun and interactivity of the live streaming viewing experience.
[0132] Step S102: In response to the movement operation of the first live video, the first live video is moved to the first display area of the live display area corresponding to the live video interface, and a second display area to be filled is determined on the live video interface; the first display area is the overlapping area between the video display area of the moved first live video and the live display area; the second display area is the area in the live display area excluding the overlapping area.
[0133] Specifically, the target viewer terminal can respond to a movement operation on the first live video by using the video frame corresponding to the movement operation as the moving video frame and determining the movement duration corresponding to the movement operation. The movement duration includes a start time and an end time. The moving video frame corresponding to the start time is the first video frame, and the moving video frame corresponding to the end time is the second video frame. Further, the target viewer terminal can change the video display area of the first live video from a first image area to a second image area within the movement duration. The first image area is the video display area of the first video frame at the start time of the movement; the second image area is the video display area of the second video frame at the end time of the movement. Further, the target viewer terminal can determine the overlapping area between the second image area and the live display area corresponding to the live video interface, and define the overlapping area as the first display area corresponding to the moved first live video. Further, the target viewer terminal can use the remaining missing video area outside the overlapping area as the second display area to be filled.
[0134] For further information, please refer to [link / reference]. Figure 6 , Figure 6 This is a schematic diagram of a mobile first live video scenario provided in an embodiment of this application. For example... Figure 6 As shown, the target viewer terminal can move in the second direction. Figure 6 The first live video shown can then have its display position customized along the movement direction (e.g., the second direction). Here, the second direction can be the +X axis direction in the screen coordinate system of the target viewer's terminal.
[0135] It should be understood that, such as Figure 6 The screen size of the live stream image 600a shown is determined by the aforementioned server based on the terminal screen information of the target viewer's terminal. Therefore, for virtual live streaming rooms (e.g., the aforementioned...), Figure 5 For different viewers (in the virtual room K accessed by user B in the corresponding embodiment), the live broadcast screen displayed on their respective terminals may not be exactly the same due to differences in the screen size and aspect ratio of their terminals. For example, when the server obtains the captured video sequence corresponding to the original captured video stream, it can traverse and capture video frames that match the screen size information of the target viewer's terminal from each captured video frame of the captured video sequence. The video sequence composed of these captured video frames can then be collectively referred to as the aforementioned first video sequence.
[0136] It is understood that, in this embodiment of the application, each video frame in the first video sequence can be divided into a grid, thereby obtaining an image block of size N*N, where N can be 3. It is understood that the reason for dividing each captured video frame in the first video sequence into a grid is that the smallest image block obtained from the division can be used to quickly locate the missing video frame in the captured video frame of the first video sequence, thereby quickly locating the image block to which the video frame that has moved out of the live video interface belongs.
[0137] For example, such as Figure 6 As shown, when user B needs to view the target object located in the live streaming interface 600a, they can follow... Figure 6 The first live video is moved along the +X axis (i.e., to the right) to customize the display position of the video display area to which the first live video belongs. For example... Figure 6As shown, when user B performs a trigger operation (e.g., a long press or other contact operation) on the live streaming interface 600a, the target viewer's terminal can adjust the service status of the video display area to a movable state in response to the trigger operation. Then, within the movable video display area, in response to the movement operation on the first live streaming video, the video frame corresponding to the movement operation in the first live streaming video can be identified as a moving video frame. This moving video frame can include a first video frame and a second video frame. Specifically, the first video frame is the moving video frame corresponding to the start time of the movement within the movement duration recorded by the target viewer's terminal, and the second video frame is the moving video frame corresponding to the end time of the movement within the movement duration recorded by the target viewer's terminal. Both the first and second video frames originate from the first live streaming video.
[0138] It should be understood that, with respect to the start time of the movement, such as Figure 6 The video frame displayed on the live streaming interface 600a can be the aforementioned first video frame. In this case, the video display area to which the first video frame belongs can be the complete image area within the live streaming display area. In this embodiment, the complete image area presented by the first video frame at the start of the movement can be collectively referred to as the first image area in the live streaming interface. At this time, the first image area has the same area size as the live streaming display area corresponding to the live streaming interface 600a (i.e., the live streaming interface).
[0139] However, as Figure 6 As shown, during the process of moving the first live video to the right, the target viewer's terminal can record in real time the video coordinate position information of the video display area to which these moving video frames belong in the first live video. For example, regarding the aforementioned end time of movement, such as Figure 6 The video frame displayed on the live streaming interface 600b can be the aforementioned second video frame. In this case, the video display area to which the second video frame belongs can be a partial image area within the aforementioned live streaming display area. In this embodiment of the application, the area containing the partial image area presented by the second video frame at the end of the movement can be collectively referred to as the second image area in the live streaming interface.
[0140] Obviously, as the target user B moves the first live video to the right, the video display area of the first live video will change from the first image area (i.e., the aforementioned complete image area) to the second image area (i.e., the aforementioned area containing the partial image area) during the movement time. Therefore, the overlapping area between the second image area and the live display area can be used as the live interface 600b. Figure 6The first display area shown should be understood to be used to display the first video data in the second live video (e.g., Figure 6 The video data shown is 601a), and in the live streaming interface 600b, the remaining video missing area outside the overlapping area in the live streaming display area can be used as... Figure 6 The second display area shown is to be filled, and this second display area can be used to display second video data in the second live video (e.g., Figure 6 The video data shown is 602a.
[0141] Step S103: Based on the first live video after the movement, determine the second live video composed of the first video data and the second video data. When playing the second live video on the live video interface, display the first video data in the first display area and the second video data in the second display area.
[0142] The first video data and the second video data both originate from the second live video, and the target object is displayed in a different position in the second live video than in the first live video.
[0143] Specifically, the target viewer terminal can obtain the moving video frame corresponding to the movement operation from the first live video, and record the starting position information of the first image area corresponding to the first video frame in the screen coordinate system of the live display area based on the first image area of the first video frame at the start time of the movement operation. Furthermore, the target viewer terminal can record the ending position information of the second image area corresponding to the second video frame in the screen coordinate system of the live display area based on the second image area of the second video frame at the end time of the movement operation. Further, the target viewer terminal can send the starting and ending position information to the server, so that the server can determine the movement of the first live video based on the starting and ending position information. During the motion displacement, based on the motion displacement and the positioning information of the moving video frame in the acquired video frame, a second video sequence associated with the target object is determined from the original acquired video stream. The acquired video frame is a video frame with the same video frame number as the moving video frame in the acquired video sequence cached by the server, and the moving video frame is obtained by the server taking a screenshot of the acquired video frame based on the terminal screen information. Furthermore, the target viewer terminal can receive the second video stream corresponding to the second video sequence returned by the server, decode the second video stream to obtain the second live video corresponding to the second video sequence, and when playing the second live video on the live video interface, the first video data in the second live video is displayed in the first display area, and the second video data in the second live video is displayed in the second display area.
[0144] For further information, please refer to [link / reference]. Figure 7 , Figure 7 This is a schematic diagram illustrating a scenario where video coordinate location information is sent according to an embodiment of this application. For example... Figure 7 The audience terminal 701a shown can be the above Figure 6 The target viewer terminal in the corresponding embodiment. For example... Figure 7 The video display area 700a shown can be the first image area of the first video frame at the start of the movement duration. Figure 7 In the screen coordinate system shown, the viewer terminal 701a can record the coordinate position information of the four corners of the first image area. Specifically, the coordinate position information of these four corners can be... Figure 7 The coordinate position information 711a, coordinate position information 712a, coordinate position information 713a, and coordinate position information 714a are shown. For ease of understanding, in this embodiment, the coordinate position information of the four corners of the first image area corresponding to the first video frame recorded in the screen coordinate system can be collectively referred to as the starting position information of the video display area to which the first live video belongs before movement.
[0145] Similarly, such as Figure 7 The video display area 700b shown is the first image area of the second video frame at the end of its movement duration. Figure 7 In the screen coordinate system shown, the viewer terminal 701a can record the coordinate position information of the four corners of the second image area. Specifically, the coordinate position information of these four corners can be... Figure 7 The coordinate position information 711b, coordinate position information 712b, coordinate position information 713b, and coordinate position information 714b are shown. For ease of understanding, in this embodiment, the coordinate position information of the four corners of the second image area corresponding to the recorded second video frame in the screen coordinate system can be collectively referred to as the end position information of the video display area to which the first live video belongs after movement.
[0146] Furthermore, it can be understood that after recording the starting position information before the first live video is moved and the ending position information after the first live video is moved, the viewer terminal 701a can encapsulate these information and send the encapsulated video coordinate position information to the camera. Figure 7 The server 702a shown is configured to parse the received video coordinate position information, thereby obtaining the starting position information before the first live video was moved and the ending position information after the first live video was moved. For example... Figure 7As shown, at this time, the server can quickly determine the coordinate position change information of the video display area to which the first live video belongs, from the first image area to the second image area, based on the parsed start position information and end position information. Thus, the server can determine the movement displacement of the first live video based on the coordinate position change information.
[0147] For example, before moving the first live video (i.e., before moving the first live video), the coordinate position information of the upper left corner of the video display area to which the first live video belongs in the screen coordinate system (e.g., Figure 7 The coordinate position information 711a shown can be (0px, 1950px). However, after moving the first live video (i.e., after moving the first live video), the coordinate position information of the upper left corner of the video display area to which the first live video belongs in the screen coordinate system (i.e. Figure 7 The coordinate position information shown in 711b can be (20px, 1950px). It should be understood that for... Figure 7 As shown in the server 702, it can quickly know that the viewer corresponding to the viewer terminal 701 dragged the video display area of the first live video horizontally to the right by 20px (that is, moved horizontally to the right by 20 pixels) during the process of watching the first live video, based on the coordinate position information of the upper left corner of the video display area recorded by the viewer terminal 701a (i.e., (0px, 1950px) and (20px, 1950px)).
[0148] Similarly, server 702 can also obtain the coordinate positions of the other three vertices before the first live video was moved (e.g., coordinate positions 712a, 713a, and 714a mentioned above), and the coordinate positions of the three vertices after the first live video was moved (e.g., coordinate positions 712b, 713b, and 714b mentioned above). Therefore, server 702 can quickly determine the displacement of the first live video based on the obtained coordinate position change information of these four vertices. Furthermore, based on this displacement (e.g., a horizontal movement of 20 pixels to the right), it can determine the second video stream associated with the target object from the original acquired video stream, and then return the second video stream. Figure 7 The viewer terminal 701a shown is configured to decode the received second video stream to obtain a second live video corresponding to the second video stream, which is synthesized from the first video data and the second video data.
[0149] At this point, the audience terminal 701a can further be displayed in the first display area (e.g., Figure 7 The first video data in the second live video is output in the overlapping area between the video display area 700b and the live broadcast display area, and can be displayed in the second broadcast area (e.g., Figure 7 The second video data in the second live video is output in area 704a); the second display area is the area in the live display area excluding the overlapping area (e.g., Figure 7 (As shown in area 704a); the target object is displayed in a different position in the second live video than it is displayed in the first live video.
[0150] Among these, it is understandable that when viewers / users... Figure 7 During the process of the viewer terminal 701a moving the first live video, the video display area to which the first live video belongs (e.g., Figure 7 The video display area 700b shown may have some missing frames. For example, the missing video frame in the second video frame at the end of the aforementioned movement is... Figure 7 The video data in region 703a, as shown, is invisible to the viewer currently moving the first live video. Simultaneously, a missing video area with the same size as region 703a appears on the live video interface. Therefore, to ensure the integrity of video playback in a live streaming scenario, this embodiment can quickly acquire data to fill the missing area as the video data in region 703 fades out of the viewer's field of vision. Figure 7 The second video data in area 704 shown can then be displayed in the second display area.
[0151] In this embodiment of the application, when outputting a first live video associated with a target object on the video live streaming interface of a virtual live streaming room, the first live video can be moved in response to a viewer's movement operation on the first live video. This allows the first live video to be moved to a first display area within the corresponding live streaming display area of the video live streaming interface, and a second display area to be filled is determined on the video live streaming interface. It should be understood that the first display area is the overlapping area between the video display area of the moved first live video and the live streaming display area; the second display area is the area within the live streaming display area excluding the overlapping area. Furthermore, this embodiment of the application can determine first video data and second video data associated with the target object based on the moved first live video. Therefore, when displaying the first video data in the first display area, the second video data can also be displayed in the second display area. It is understood that both the first and second video data originate from the second live video, and the display position of the target object in the second live video differs from its display position in the first live video. Thus, this embodiment of the application can refer to the live video before the movement as the first live video, and the live video obtained after moving the first live video as the second live video. It should be understood that the second live video is the live video obtained after adjusting the display position of the target object on the live video interface. In other words, the embodiments of this application provide a solution that can flexibly control the display position of the target object on the live video interface. This solution allows viewers in the virtual live room to customize the display position of the target object while watching the live video (i.e., the aforementioned first live video). This solves the problem of multimedia information obscuring the target object on the live video interface in a live streaming scenario, thereby fundamentally optimizing the display effect of the target object on the live video interface.
[0152] Further, please see Figure 8 , Figure 8 This is a schematic diagram of a video data processing method provided in an embodiment of this application. Figure 8 As shown, the method can be implemented by the viewer terminal (e.g., the one described above). Figure 1 The corresponding embodiment's viewer terminal 4000a) can be executed, or it can be executed by a service server (e.g., the one described above). Figure 1 The method can be executed by the server 2000 in the corresponding embodiment, or it can be jointly executed by the viewer terminal and the service server. For ease of understanding, this embodiment uses the execution of the method by the viewer terminal as an example. The viewer terminal can be the target viewer terminal mentioned above. The method can specifically include the following steps S201-S211:
[0153] Step S201: In response to the launch operation for the application client, output the application display interface corresponding to the application client;
[0154] Step S202: In response to the triggering operation of the virtual live room in the application display interface, send an access request for the virtual live room to the server.
[0155] The access request instructs the server to extract a first video sequence from the captured video sequence corresponding to the original captured video stream, which matches the terminal screen information of the application client. The original captured video stream is the video stream that the server pulls from the broadcaster client and is associated with the broadcaster user in the virtual live room. The size of the video frames in the first video sequence is less than or equal to the size of the video frames in the captured video sequence. It can be understood that the broadcaster client here is the live streaming client corresponding to the broadcaster user in the virtual live room.
[0156] Step S203: Receive the first video stream corresponding to the first video sequence returned by the server, decode the first video stream, and obtain the first live video of the target object associated with the broadcaster user.
[0157] Step S204: Output the first live video on the video live streaming interface corresponding to the virtual live streaming room;
[0158] The specific implementation methods of steps S201-S204 can be found in the above. Figure 3 The description of step S101 in the corresponding embodiments will not be repeated here.
[0159] Step S205: Output video auxiliary elements associated with the first live video on the live video interface.
[0160] It is understood that when the target audience terminal performs step S204, it can simultaneously execute step S205. For example, when a viewer accesses a virtual live streaming room created by a broadcaster through the target audience terminal, the first live video associated with the broadcaster can be output at the first business level of the video live streaming interface of the virtual live streaming room, and video auxiliary elements associated with the first live video can also be output at the second business level of the video live streaming interface. These video auxiliary elements may include, but are not limited to, public screen message elements, interactive elements, and operation elements associated with the virtual live streaming room, and then the following step S206 can be executed.
[0161] For further information, please refer to [link / reference]. Figure 9 , Figure 9 This is a schematic diagram illustrating a scenario of two business layers within a video live streaming interface, as provided in an embodiment of this application. The two business layers within the video live streaming interface can be... Figure 9 The business layers 800a and 800b are shown.
[0162] Specifically, service level 800a can be the first service level (i.e., level 1) in the video live streaming interface. This first service level can be used to display the video data obtained by the target viewer's terminal after decoding the first video stream. For example, it can be used in... Figure 9 The video data 900a obtained by decoding is displayed on the business layer 800a shown. The video data 900a here can be a pure live screen of the target object associated with the anchor user played in the video live screen interface. In this application embodiment, the live screen played in the video live screen interface can be collectively referred to as the first live video.
[0163] Specifically, business layer 800b can be the second business layer (i.e., layer 2) in the video live streaming interface. This second business layer can be used to display public screen message elements, interactive elements, and operation elements associated with the virtual live streaming room. For example, Figure 9 As shown, the target viewer terminal can determine the public screen message area associated with the second service level on the video live streaming interface (e.g., Figure 9 Area 902a shown), business interaction area (e.g., Figure 9 The area shown is 901a) and the content operation area (e.g., Figure 9 The area shown is 903a.
[0164] Among them, such as Figure 9 As shown, the target viewer's terminal can be in the public screen message area (e.g., Figure 9 In area 902a) shown, a public broadcast message associated with the target object in the virtual live stream is output, which can then be used as a public screen message element. Specifically, public screen message elements may include, but are not limited to, bullet screen messages sent by viewers in the virtual live stream (e.g., ...). Figure 9 The interactive bullet screen message 1 posted by viewer B1 ("Looks delicious") and the interactive bullet screen message 2 posted by viewer B2 ("The cake looks beautiful") are displayed in the lower left corner of area 902a in the form of bubbles, along with recommended product information and product coupons distributed in the virtual live broadcast room.
[0165] Among them, such as Figure 9 As shown, the target audience terminal can be located in the business interaction area (e.g., Figure 9 The output in area 902a) contains interactive message elements associated with the streamer in the virtual live stream room (e.g., the streamer's avatar, nickname, and popularity information; viewers can follow the streamer by triggering their avatar and thus make them a follower based on their personal interests), and in the content operation area (e.g., ...). Figure 9 The area 903a) shown outputs action elements associated with viewers in the virtual live stream room. These action elements may include entering and posting interactive bullet comments, purchasing goods, sharing the live stream, and liking the live stream.
[0166] Based on this, the target audience terminal can collectively refer to the aforementioned public screen message elements, interactive elements, and operational elements, as video auxiliary elements associated with the first live video, and thus... Figure 9 The video live streaming interface 911a shown outputs video auxiliary elements associated with the first live video at the first service level, so that the following steps S206 can be performed subsequently.
[0167] Step S206: Obtain the video dragging conditions corresponding to the live video interface, and detect the video auxiliary elements on the second business level based on the video dragging conditions.
[0168] Specifically, when the target audience terminal detects that the number of video auxiliary elements meets the quantity threshold in the video dragging condition, it can output movement instruction information for the first live video at the first business level on the live video interface; it can be understood that the movement instruction information here is used to instruct the audience users in the virtual live room to perform movement operations on the first live video.
[0169] For further information, please refer to [link / reference]. Figure 10 , Figure 10 This is a schematic diagram illustrating a scenario where movement indication information is output on a live video interface, as provided in an embodiment of this application. It can be understood that... Figure 10 The video live streaming interface 1011a shown can be the above Figure 9 The corresponding embodiment shows the video live streaming interface. It should be understood that the second service level is independent of the first service level, and in the video live streaming interface, the second service level is above the first service level. Therefore, as... Figure 10 As shown, when the target viewer's terminal detects that the number of video auxiliary elements output on the second service level of the video live streaming interface 1011a reaches the quantity threshold in the video dragging condition (i.e., when the number of video auxiliary elements on the second service level is sufficient), it can... Figure 10The live video interface 1011b shown outputs movement instructions for the first live video at the first business level. Specifically, this movement instruction can be a user-friendly guide (or simply guidance information) that allows for customizing the video display area. This guidance information may include: What to do if the video content is blocked? Long-press the video to drag the live window. It should be understood that the live window refers to the video display area used to display the first live video. Thus, when a user watching the first live video finds that an object they are watching (i.e., the target object) is obscured by video auxiliary elements displayed at the second business level, they can adaptively choose whether to move the first live video according to their needs. For example, within the indicated duration of the movement instruction, the user can switch the business status of the video display area to a movable state based on the movement instruction.
[0170] For example, the target viewer's terminal can respond to a long-press operation on the video live streaming interface 1011b, and then, based on this long-press operation, switch the service status of the video display area to which the first live video belongs to a movable state. This allows for customization of the video display position of the first live video's video display area on the target viewer's terminal's video live streaming interface 1011b, and further, the video display area to which the first live video belongs after being moved can be determined based on the customized video display position. It is understood that, in this embodiment, the overlapping area between the video display area to which the first live video belongs after being moved and the live streaming display area of the aforementioned video live streaming interface can be collectively referred to as the aforementioned first display area.
[0171] Step S207: In response to a motion operation for the first live video, the video frame corresponding to the motion operation is taken as the motion video frame, and the motion duration corresponding to the motion operation is determined.
[0172] The movement duration includes the movement start time and the movement end time; the movement video frame corresponding to the movement start time is the first video frame, and the movement video frame corresponding to the movement end time is the second video frame; specifically, when the target viewer terminal obtains the movement indication information corresponding to the first live video, it can respond to the trigger operation for the movement execution information, determine the video display area to which the first live video belongs at the first business level of the video live interface, and adjust the business status of the video display area to which the first live video belongs to the movable state; furthermore, within the video display area with the movable state, in response to the movement operation for the first live video, the target viewer terminal can determine the video frame corresponding to the movement operation in the first live video as the movement video frame, and determine the movement duration corresponding to the movement operation based on the movement start time corresponding to the first video frame and the movement end time corresponding to the second video frame in the movement video frame.
[0173] Step S208: During the movement time, change the video display area of the first live video from the first image area to the second image area;
[0174] Wherein, the first image area is the video display area of the first video frame at the start of the movement; the second image area is the video display area of the second video frame at the end of the movement;
[0175] Step S209: Determine the overlapping area between the second image area and the live display area corresponding to the live video interface, and define the overlapping area as the first display area corresponding to the first live video after the movement.
[0176] Step S210: In the live broadcast display area, the remaining video missing areas, excluding the overlapping areas, are used as the second display area to be filled.
[0177] Step S211: Based on the first live video after the movement, determine the second live video composed of the first video data and the second video data. When playing the second live video on the live video interface, display the first video data in the first display area and the second video data in the second display area.
[0178] Specifically, the target viewer terminal can obtain the moving video frame corresponding to the movement operation from the first live video. Based on the first image region of the first video frame at the start time of the movement operation, it records the start position information of the first image region corresponding to the first video frame in the screen coordinate system to which the live display area belongs. Furthermore, the target viewer terminal can record the end position information of the second image region corresponding to the second video frame in the screen coordinate system to which the live display area belongs, based on the second image region of the second video frame at the end time of the movement operation. Further, the target viewer terminal can send the start and end position information to the server, so that the server can determine the movement position of the first live video based on the start and end position information. During the shift, based on the displacement and the positioning information of the moving video frame in the acquired video frame, a second video sequence associated with the target object is determined from the original acquired video stream; the acquired video frame is a video frame with the same video frame number as the moving video frame in the acquired video sequence cached by the server, and the moving video frame is obtained by the server taking a screenshot of the acquired video frame based on the terminal screen information; furthermore, the target viewer terminal can receive the second video stream corresponding to the second video sequence returned by the server, decode the second video stream to obtain the second live video corresponding to the second video sequence, and when playing the second live video on the video live interface, the first video data in the second live video is displayed in the first display area, and the second video data in the second live video is displayed in the second display area.
[0179] It can be understood that the aforementioned first video stream is obtained by the server encoding the first video sequence using the first encoder; the frame numbers of the video frames in the first video sequence are consistent with the frame numbers of the video frames in the acquired video sequence. Therefore, during the execution of step S211, the target viewer terminal can also perform the following steps. For example, the target viewer terminal can also generate an encoding switching request for the virtual live broadcast room based on the start position information and end position information, and can send the encoding switching request to the server.
[0180] The encoding switch request is used to instruct the server to keep the first encoder running and start the second encoder. It should be understood that in this embodiment, when the second encoder is started to encode a new video stream, the video sequence in the new video stream (i.e., the second video stream) is synchronized with the video sequence in the aforementioned first video stream (i.e., the original video stream) in terms of frame number. Thus, after the server selects a keyframe (e.g., video frame I) in the video sequence of the new video stream, it can shut down the first encoder when it detects that the previous video frame of video frame I in the first video stream (i.e., the original video stream) has been transmitted, freeing up the encoder's hardware resources to prepare for receiving the next encoding switch request. It should be understood that the server can encode the second video sequence using the newly started second encoder to obtain the second video stream corresponding to the second video sequence; it is understood that the video frames in the second video sequence are determined by the server after taking screenshots of the video frames in the acquired video sequence based on terminal screen information and positioning information, and the display position of the target object in the second video sequence is different from the display position of the target object in the first video sequence. The specific implementation method of the server extracting the second video sequence based on the captured video sequence corresponding to the original captured video stream can be found in the above. Figure 3 The specific process of capturing the first video sequence described in the corresponding embodiments will not be repeated here.
[0181] It is understood that, to ensure the continuity of live video playback, the frame interval between the frame number of video frame I selected in this embodiment and the frame number of the last keyframe (e.g., video frame H) in the first video stream can be greater than 1 / 2 of the frame group length of the second video stream. This ensures the naturalness of frame switching in the live video playback interface on the target viewer's terminal. Furthermore, it is understood that, in a live streaming scenario, this embodiment allows for customization of the display position of the video display area of the first live video according to the actual needs of the viewer. Thus, when a viewer finds that their target object is obscured by video auxiliary elements displayed at the second service level, they can use the target viewer's terminal to output movement instructions on the video live streaming interface corresponding to the virtual live streaming room. This allows the viewer to drag the first live video according to the movement instructions to customize the display position of the video display area of the first live video on the video live streaming interface. Since the first business layer on the video live streaming interface is independent of the second business layer, this application embodiment can customize the display position of the video display area, thus avoiding the need to forcibly remove the video auxiliary elements displayed on the second business layer. This not only ensures the integrity of watching the live stream, but also ensures that certain key information in the virtual live streaming room (such as the target object that the viewer is interested in) is not obscured, thereby enhancing the fun of watching the live stream.
[0182] Further, please see Figure 11 , Figure 11 This is an interactive schematic diagram of a video data processing method provided in an embodiment of this application. For example... Figure 11 As shown, this method can be used by the viewer terminal (e.g., the one described above). Figure 1 The corresponding embodiment includes the audience terminal 4000a and the service server (e.g., the one described above). Figure 1 The server 2000 in the corresponding embodiment executes the method together. The viewer terminal can be the aforementioned target viewer terminal, and the service server can be the aforementioned server. Specifically, the method may include the following steps S301-S308:
[0183] In step S301, the target viewer terminal can respond to the launch operation for the application client and output the application display interface corresponding to the application client.
[0184] In step S302, the target viewer terminal can respond to the trigger operation of the virtual live room in the application display interface by sending an access request to the server for the virtual live room.
[0185] It is understood that the access request can be used to instruct the server to perform the following step S303, for example, the server can extract a first video sequence from the captured video sequence corresponding to the original captured video stream that matches the terminal screen information of the application client; the original captured video stream can be a video stream associated with the broadcaster user in the virtual live room that the server pulls from the broadcaster client; the size of the video frame in the first video sequence is less than or equal to the size of the video frame in the captured video sequence.
[0186] In step S303, the server can determine the first video stream that matches the terminal screen information from the original captured video stream corresponding to the virtual live room based on the access request, and return the first video stream to the application client so that the application client can output the first live video associated with the target object on the video live streaming interface of the virtual live room based on the first video stream.
[0187] It should be understood that when a target viewer terminal running the application client receives the first video stream corresponding to the first video sequence returned by the server, it can decode the first video stream to obtain the first live video of the target object associated with the broadcaster user.
[0188] In step S304, the target viewer's terminal can output the first live video on the video live streaming interface corresponding to the virtual live streaming room.
[0189] In step S305, the target viewer terminal can respond to the movement operation of the first live video, move the first live video to the first display area of the live display area corresponding to the live video interface, and determine the second display area to be filled on the live video interface.
[0190] The first display area is the overlapping area between the video display area of the first live video after it has been moved and the live display area; the second display area is the area in the live display area excluding the overlapping area.
[0191] Step S306: When the application client moves the first live video to the first display area of the live display area corresponding to the live video interface, the server can receive the video coordinate position information sent by the application client based on the moved first live video, and determine the movement displacement of the first live video based on the video coordinate position information.
[0192] Specifically, the server can receive video coordinate position information sent by the application client based on the first live video after movement, and determine the start and end position information associated with the moving video frame in the first live video based on the video coordinate position information; the moving video frame is the video frame corresponding to the movement operation determined by the application client in response to the movement operation on the first live video; the start position information is the coordinate position information of the first video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the first image area at the start time of the movement operation of the first video frame in the moving video frame; the end position information is the coordinate position information of the second video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the second image area at the end time of the movement operation of the second video frame in the moving video frame; furthermore, the server can determine the coordinate position change information of the video display area to which the first live video belongs, from the first image area to the second image area, based on the start and end position information, and determine the movement displacement of the first live video based on the coordinate position change information.
[0193] The specific implementation method for the server to determine the movement and displacement of the first live video can be found in the above. Figure 3 The specific process of determining the movement displacement of the first live video based on the coordinate position change information in the corresponding embodiment will not be repeated here.
[0194] In step S307, the server can determine the second video stream associated with the target object from the original acquired video stream based on the movement displacement, and return the second video stream to the application client.
[0195] Specifically, the server can determine the location information of the first live video after movement in the original captured video stream based on the displacement. Based on the location information and the terminal screen information, the server can determine the second video sequence associated with the target object from the original captured video stream. Furthermore, the server can encode the second video sequence to obtain the second video stream corresponding to the second video sequence, and return the second video stream to the application client.
[0196] The specific method by which the server determines the location information of the first live video after movement within the original captured video stream based on the displacement, and then determines the second video sequence associated with the target object from the original captured video stream based on the location information and the terminal screen information, can be described as follows:
[0197] The server can obtain the captured video sequence corresponding to the original captured video stream, and based on the vertices position information of the video frames in the captured video sequence in the positioning coordinate system, determine the movement safety area corresponding to the original captured video stream. Then, it can select video frames with the same video frame number as the moving video frames in the captured video sequence as captured video frames. Furthermore, the server can determine the start and end coordinates of the moving video frames in the captured video frames based on the movement displacement in the positioning coordinate system, and can use the end coordinates as the positioning position information of the first live video after movement in the original captured video stream. Then, it can compare the positioning position information with the movement safety area to obtain the comparison result. It can be understood that the image area formed by the start coordinates is the first positioning area of the first video frame in the moving video frame in the positioning coordinate system, and the image area formed by the end coordinates is... The second video frame in the moving video frame is located in the second positioning area in the positioning coordinate system; the overlapping area between the second positioning area and the first positioning area is the first screenshot area, and the remaining area in the first positioning area excluding the overlapping area is the second screenshot area; it should be understood that if the comparison result indicates that each movement end coordinate information in the positioning location information is within the movement safety area, the server can determine that the first live video after movement is located within the movement safety area, and then, based on the terminal screen information and the first screenshot method, the video data belonging to the first screenshot area in the captured video sequence can be used as the first video data to be displayed in the first display area, and the video data belonging to the second screenshot area in the captured video sequence can be used as the second video data to be displayed in the second display, the first video data and the second video data can be synthesized, and the synthesized video data can be used as the second video sequence associated with the target object.
[0198] For further information, please refer to [link / reference]. Figure 12 , Figure 12 This is a schematic diagram illustrating a scenario where missing frames are automatically filled in within a safe area, as provided in an embodiment of this application. Figure 12 The captured video frame 1201a shown can be the captured video frame corresponding to the first video frame mentioned above, and Figure 12 The moving video frame 1203a shown is the first video frame displayed in the aforementioned live video interface, which is captured by the server from the acquired video frame 1201a based on the size of the screenshot area 1202a. For example... Figure 12 The size of the screenshot area 1202a shown can be the same as the screen size in the terminal screen information of the target viewer's terminal. It should be understood that the size of the screenshot area 1202a used for taking screenshots may not be exactly the same for different viewer terminals in the virtual live broadcast room.
[0199] like Figure 12As shown, the server will capture video frames (e.g., ...) in the video capture sequence. Figure 12 The captured video frame 1201a shown is in the positioning coordinate system (e.g., Figure 12 The vertex position information (shown in the X'Y' plane) is used to determine the position of the vertex. Figure 12 The mobile safety zone corresponding to the original captured video stream shown can be... Figure 12 The safe zone shown.
[0200] Among them, such as Figure 12 The captured video frame 1201b shown can be the captured video frame corresponding to the second video frame mentioned above, and Figure 12 The moving video frame 1203b shown is the second video frame that the server captures from the acquired video frame 1201b based on the size of the screenshot area 1202a, and is used to display in the aforementioned live video interface. For example... Figure 12 As shown, if the server determines that the movement displacement of the first live video is +20px, that is, the viewer has moved the second video frame horizontally to the right by 20 pixels in the target viewer's terminal. Since the second video frame is captured by the server from the captured video frame 1201b based on the size of the screenshot area 1202a, the server can determine in this positioning coordinate system that the captured video frame 1201b should also have moved horizontally to the right by 20 pixels synchronously. Therefore, at this time, the server can determine the first video frame (e.g., based on the aforementioned movement start coordinate information of the first video frame in this positioning coordinate system) by using the first video frame's movement start coordinate information. Figure 12 The moving video frame 1203a shown is located in the first positioning area of the positioning coordinate system. Simultaneously, the server can also determine the location of the second video frame (e.g., based on the aforementioned second video frame's end-of-motion coordinates in this positioning coordinate system). Figure 12 The second positioning area of the moving video frame 1203b shown in the positioning coordinate system.
[0201] At this point, it is understandable that the server can use the end-of-motion coordinates of the moving video frame 1203a as the location information of the first live video after the movement within the original captured video stream. For example... Figure 12 As shown, since the coordinates of the end of movement at the four corners of the second video frame are all within the safe area, the server can determine that the first live video after the movement is within the safe movement area. Based on this, the server can further... Figure 12The overlapping area between the first and second positioning areas defines the first screenshot area for capturing the first video data. The remaining area within the first positioning area, excluding the overlapping area, can be used as the second screenshot area for capturing the second video data. It is understood that when a viewer moves the first live video and prepares to release it, the second screenshot area represents the missing video area displayed against the background color. The remaining area within the second positioning area, excluding the overlapping area, represents a partial missing portion of the video footage from the aforementioned second video frame. Therefore, during the movement of the first live video, the video data in this partially missing area is invisible to the viewer.
[0202] Furthermore, such as Figure 12 As shown, the server can, based on the terminal screen information and the first screenshot method, use video data belonging to the first screenshot area as the first video data to be displayed in the first display area, and video data belonging to the second screenshot area as the second video data to be displayed in the second display area. For example, as... Figure 12 As shown, the server can combine the first video data captured in the first screenshot area with the second video data captured in the second screenshot area to obtain... Figure 12 The second video sequence shown is video frame 1204a. It should be understood that, at this point, all other video frames in the second video sequence can be synthesized in the same way as video frame 1204a, which will not be elaborated further here.
[0203] Optionally, if the comparison result indicates that each movement end coordinate information in the positioning location information contains movement end coordinate information outside the movement safety area, the server can determine that a local video area in the first live video after movement is located outside the movement safety area; the local video area is the area in the second positioning area composed of the movement frame position information outside the movement safety area; further, the server can fill the image data in the local video area into the second screenshot area based on the terminal screen information and the second screenshot method, use the video data belonging to the first screenshot area as the first video data to be displayed in the first display area in the captured video sequence, and use the video data in the second screenshot area as the second video data to be displayed in the second display in the captured video sequence, perform composite processing on the first video data and the composited video data as the second video sequence associated with the target object.
[0204] For further information, please refer to [link / reference]. Figure 13 , Figure 13 This is a schematic diagram illustrating a scenario where missing frames are automatically filled in within a non-safe area, as provided in an embodiment of this application. For example... Figure 13 The captured video frame 1301a shown can be the captured video frame corresponding to the first video frame mentioned above, and Figure 13 The moving video frame 1303a shown is the first video frame displayed in the aforementioned live video interface, which is captured by the server from the acquired video frame 1301a based on the size of the screenshot area 1302a. For example... Figure 13 The size of the screenshot area 1302a shown can be the same as the screen size in the terminal screen information of the target viewer's terminal. It should be understood that the size of the screenshot area 1302a used for taking screenshots may not be exactly the same for different viewer terminals in the virtual live broadcast room.
[0205] like Figure 13 As shown, the server will capture video frames (e.g., ...) in the video capture sequence. Figure 13 The captured video frame 1301a shown is in the positioning coordinate system (e.g., Figure 13 The vertex position information (shown in the X'Y' plane) is used to determine the position of the vertex. Figure 13 The mobile safety zone corresponding to the original captured video stream shown can be... Figure 13 The safety zone is shown. (As shown in the image) Figure 13 As shown, when a viewer drags video frame 1303b upwards and moves it out of the safe area, the server can determine that a portion of the first live video after the movement is located outside the safe area. At this time, the server can, based on the terminal screen information and the second screenshot method, fill the second screenshot area with image data from the portion of the video area (i.e., image data from the partially missing area of the second video frame), so that the video sequence containing the data belonging to the first screenshot area (i.e., the missing area of the second video frame) is included. Figure 13 The video data in the overlapping area of the first and second positioning areas shown is used as the first video data to be displayed in the first display area, and the video data in the second screenshot area (i.e., the remaining area of the first positioning area excluding the overlapping area) in the captured video sequence is used as the second video data to be displayed in the second display. Figure 13 As shown, at this point, the server can synthesize the captured first video data and the captured second video data to obtain... Figure 13 Video frame 1304a in the second video sequence shown.
[0206] Therefore, when the application client in the target viewer's terminal records the coordinate positions of the video display area of the first live video before and after the movement, it can collectively refer to these recorded coordinate positions (e.g., the aforementioned start and end position information) as video coordinate position information, and send this video coordinate position information to... Figure 12The server shown allows it to determine the displacement of the first live video based on received video coordinates, and then determine the location of the moved first live video within the original captured video stream in a positioning coordinate system. Furthermore, the server can determine whether each movement end coordinate in the positioning information is within the corresponding movement safety area of the original captured video stream. If so, it can automatically fill in the missing frames. Conversely, if not, it can extract video data from the aforementioned partially missing areas as second video data to fill the second display area.
[0207] In step S308, when the target viewer terminal obtains the second live video based on the second video stream, it can output the first video data in the second live video in the first display area and output the second video data in the second live video in the second display area.
[0208] The second display area is the area outside the overlapping area in the live broadcast display area; the display position of the target object in the second live broadcast video is different from the display position of the target object in the first live broadcast video.
[0209] It should be understood that the target viewer's terminal can determine the second live video, composed of the first video data and the second video data, based on the aforementioned moved first live video. It should also be understood that both the first and second video data are associated with the target object. Therefore, when the target viewer's terminal plays the second live video on the live video interface, it can display the first video data in the first display area and the second video data in the second display area.
[0210] Therefore, in this embodiment, the target viewer terminal can customize the display position of the live video area according to actual needs. For example, when the viewer finds that the product they want to follow is obscured, they can manually drag the video display position. At this time, the target viewer terminal will record the video coordinate position information before and after dragging the first live video, and can send the recorded video coordinate position information to the server, so that the server can determine the movement displacement of the first live video based on the received video coordinate position information. Then, based on the movement displacement, the server can quickly obtain the second video data to fill the second display area. Then, the first video data to be displayed in the first display area and the second video data to be displayed in the second display area can be synthesized to obtain the second video sequence to be played in the virtual live room. It should be understood that the target viewer terminal can automatically fill in the screen in the drag gap corresponding to the second display area in the virtual live room based on the second live video corresponding to the received second video stream. This means that this embodiment can not only ensure the integrity of watching the live broadcast, but also prevent the key information in the live broadcast room (i.e., the aforementioned target object) from being obscured, thereby improving the viewing experience and fun of the live broadcast.
[0211] Further, please see Figure 14 , Figure 14 This is a schematic diagram of the structure of a video data processing device provided in an embodiment of this application. The video data processing device 1 may include: a first video acquisition module 10, a video movement module 20, and a video display module 30. Further, the video data processing device 1 may also include: a video element output module 40, a drag condition acquisition module 50, an indication information output module 60, and a switch request sending module 70.
[0212] The first video acquisition module 10 is used to output the first live video associated with the target object on the video live streaming interface of the virtual live streaming room.
[0213] The first video acquisition module 10 includes: an application interface output unit 101, an access request sending unit 102, a first video stream receiving unit 103, and a first video output unit 104.
[0214] The application interface output unit 101 is used to respond to the launch operation for the application client and output the application display interface corresponding to the application client.
[0215] The access request sending unit 102 is used to send an access request for the virtual live room to the server in response to a trigger operation on the virtual live room in the application display interface; the access request is used to instruct the server to extract a first video sequence from the captured video sequence corresponding to the original captured video stream that matches the terminal screen information of the application client; the original captured video stream is the video stream that the server pulls from the broadcaster client and is associated with the broadcaster user in the virtual live room; the size of the video frame in the first video sequence is less than or equal to the size of the video frame in the captured video sequence;
[0216] The first video stream receiving unit 103 is used to receive the first video stream corresponding to the first video sequence returned by the server, and to decode the first video stream to obtain the first live video of the target object associated with the broadcaster user.
[0217] The first video output unit 104 is used to output the first live video on the video live streaming interface corresponding to the virtual live streaming room.
[0218] The specific implementation methods of the application interface output unit 101, the access request sending unit 102, the first video stream receiving unit 103, and the first video output unit 104 can be found in the above description. Figure 3 The specific process of outputting the first live video in the corresponding embodiment will not be repeated here.
[0219] Optionally, when the first video acquisition module 10 is used to perform the step of outputting the first live video associated with the target object on the video live streaming interface of the virtual live streaming room, the video element output module 40 is used to output video auxiliary elements associated with the first live video on the video live streaming interface.
[0220] The video live streaming interface includes a first business layer and a second business layer; the first business layer is used to display the first live video played in the virtual live streaming room; the second business layer is used to display public screen message elements, interactive elements and operation elements associated with the virtual live streaming room.
[0221] The video element output module 40 includes: a region determination unit 401, a broadcast message output unit 402, an element information output unit 403, and an auxiliary element determination unit 404;
[0222] The area determination unit 401 is used to determine the public screen message area, business interaction area and content operation area associated with the second business level on the video live broadcast interface;
[0223] The broadcast message output unit 402 is used to output public broadcast messages associated with the target object in the virtual live broadcast room in the public screen message area, and to use the public broadcast messages as public screen message elements;
[0224] The element information output unit 403 is used to output interactive message elements associated with the anchor user in the virtual live broadcast room in the business interaction area, and to output operation elements associated with the audience user in the virtual live broadcast room in the content operation area.
[0225] The auxiliary element determination unit 404 is used to identify public screen message elements, interactive elements, and operation elements as video auxiliary elements associated with the first live video.
[0226] The specific implementation methods of the region determination unit 401, the broadcast message output unit 402, the element information output unit 403, and the auxiliary element determination unit 404 can be found in the above description. Figure 8 The description of the video auxiliary elements in the corresponding embodiments will not be repeated here.
[0227] Optionally, before the video movement module 20 performs the step in response to the movement operation for the first live video, the drag condition acquisition module 50 is used to acquire the video drag conditions corresponding to the live video interface, and to detect the video auxiliary elements on the second business level based on the video drag conditions.
[0228] The instruction information output module 60 is used to output movement instruction information for the first live video at the first business level on the video live streaming interface when the number of video auxiliary elements is detected to meet the quantity threshold in the video dragging condition; the movement instruction information is used to instruct the audience users in the virtual live streaming room to perform movement operations on the first live video.
[0229] The video movement module 20 is used to respond to a movement operation on the first live video, move the first live video to the first display area of the live display area corresponding to the video live interface, and determine the second display area to be filled on the video live interface; the first display area is the overlapping area between the video display area of the moved first live video and the live display area; the second display area is the remaining area in the live display area excluding the overlapping area.
[0230] The video motion module 20 includes: a motion operation response unit 201, a display area change unit 202, an overlapping area determination unit 203, and a missing area determination unit 204.
[0231] The motion operation response unit 201 is used to respond to a motion operation on the first live video, take the video frame corresponding to the motion operation as the motion video frame, and determine the motion duration corresponding to the motion operation; the motion duration includes the motion start time and the motion end time; the motion video frame corresponding to the motion start time is the first video frame, and the video frame corresponding to the motion end time is the second video frame.
[0232] The mobile operation response unit 201 includes: an indication information triggering unit 2011 and a video movement unit 2012;
[0233] The instruction information triggering unit 2011 is used to obtain the motion instruction information corresponding to the first live video, respond to the triggering operation for the motion execution information, determine the video display area to which the first live video belongs at the first business level of the video live interface, and adjust the business status of the video display area to which the first live video belongs to the movable state.
[0234] The video motion unit 2012 is configured to, in a video display area having a movable state, in response to a motion operation for a first live video, determine the video frame corresponding to the motion operation as a motion video frame in the first live video, and determine the motion duration corresponding to the motion operation based on the motion start time corresponding to the first video frame in the motion video frame and the motion end time corresponding to the second video frame in the motion video frame.
[0235] The specific implementation methods of the indication information triggering unit 2011 and the video movement unit 2012 can be found in the above description. Figure 3 The specific process for determining the movement duration described in the corresponding embodiments will not be repeated here.
[0236] The display area changing unit 202 is used to change the video display area of the first live video from the first image area to the second image area during the movement duration; the first image area is the video display area of the first video frame at the start of the movement; the second image area is the video display area of the second video frame at the end of the movement.
[0237] The overlapping area determination unit 203 is used to determine the overlapping area between the second image area and the live display area corresponding to the video live interface, and to determine the overlapping area as the first display area corresponding to the first live video after the movement.
[0238] The missing area determination unit 204 is used to identify the remaining video missing areas (excluding overlapping areas) in the live broadcast display area as the second display area to be filled.
[0239] The specific implementation methods of the movement operation response unit 201, the display area change unit 202, the overlapping area determination unit 203, and the missing area determination unit 204 can be found in the above description. Figure 3 The specific process for determining the first display area and the second display area in the corresponding embodiments will not be repeated here.
[0240] The video display module 30 is used to determine a second live video consisting of first video data and second video data based on the first live video after the movement. When playing the second live video on the video live interface, the first video data is displayed in the first display area, and the second video data is displayed in the second display area. The display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0241] The video display module 30 includes: a start position recording unit 301, an end position recording unit 302, a position sending unit 303, and a second video stream receiving unit 304.
[0242] The starting position recording unit 301 is used to obtain the moving video frame corresponding to the moving operation from the first live video, and record the starting position information of the first image area corresponding to the first video frame in the screen coordinate system to which the live display area belongs, based on the first image area of the first video frame at the start time of the moving operation.
[0243] The end position recording unit 302 is used to record the end position information of the second image area corresponding to the second video frame in the screen coordinate system to which the live display area belongs, based on the second image area of the second video frame in the moving video frame at the end of the moving operation.
[0244] The location sending unit 303 is used to send the start position information and end position information to the server, so that when the server determines the movement displacement of the first live video based on the start position information and end position information, it can determine the second video sequence associated with the target object from the original captured video stream based on the movement displacement and the positioning information of the moving video frame in the captured video frame; the captured video frame is a video frame with the same video frame number as the moving video frame in the captured video sequence cached by the server, and the moving video frame is obtained by the server taking a screenshot of the captured video frame based on the terminal screen information;
[0245] The second video stream receiving unit 304 is used to receive the second video stream corresponding to the second video sequence returned by the server, decode the second video stream to obtain the second live video corresponding to the second video sequence, and when playing the second live video on the video live streaming interface, display the first video data in the second live video in the first display area, and display the second video data in the second live video in the second display area.
[0246] The specific implementation methods of the start position recording unit 301, end position recording unit 302, position sending unit 303, and second video stream receiving unit 304 can be found above. Figure 3The specific process of obtaining the second live video as described in the corresponding embodiments will not be repeated here.
[0247] The first video stream is obtained by the server encoding the first video sequence using the first encoder; the frame number of the video frame in the first video sequence is consistent with the frame number of the video frame in the acquired video sequence.
[0248] The switching request sending module 70 is used to generate an encoding switching request for the virtual live broadcast room based on the start position information and the end position information, and send the encoding switching request to the server. The encoding switching request is used to instruct the server to encode the second video sequence through the second encoder when the second encoder is started, so as to obtain the second video stream corresponding to the second video sequence. The video frames in the second video sequence are determined by the server after taking screenshots of the video frames in the acquired video sequence based on the terminal screen information and the positioning position information, and the display position of the target object in the second video sequence is different from the display position of the target object in the first video sequence.
[0249] The specific implementation methods of the first video acquisition module 10, the video movement module 20, and the video display module 30 can be found in the above description. Figure 3 The descriptions of steps S101 and S103 in the corresponding embodiments will not be repeated here. Optionally, the specific implementations of the video element output module 40, the drag condition acquisition module 50, the indication information output module 60, and the switch request sending module 70 can be found above. Figure 8 or Figure 11 The description of the target audience terminal side in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0250] Further, please see Figure 15 , Figure 15 This is a schematic diagram of the structure of a video data processing device provided in an embodiment of this application. The video data processing device 2 may include: a first video return module 100, a movement displacement determination module 200, and a second video determination module 300. Further, the video data processing device 2 may also include: an encoder start module 400, a frame number synchronization module 500, a keyframe determination module 600, and an encoder stop module 700.
[0251] The first video return module 100 is used to receive an access request sent by the application client based on the terminal screen information for accessing the virtual live room, and based on the access request, determine the first video sequence that matches the terminal screen information from the original captured video stream corresponding to the virtual live room, and return the video stream corresponding to the first video sequence to the application client, so that the application client can output the first live video associated with the target object on the video live interface of the virtual live room based on the video stream corresponding to the first video sequence.
[0252] The displacement determination module 200 is used to receive video coordinate position information sent by the application client based on the moved first live video when the application client moves the first live video to the first display area in the live display area corresponding to the video live interface, and to determine the displacement of the first live video based on the video coordinate position information; the first display area is the overlapping area between the video display area of the moved first live video and the live display area.
[0253] The displacement determination module 200 includes: a coordinate position receiving unit 2001 and a displacement determination unit;
[0254] The coordinate position receiving unit 2001 is used to receive video coordinate position information sent by the application client based on the first live video after movement, and to determine the start position information and end position information associated with the moving video frame in the first live video based on the video coordinate position information; the moving video frame is the video frame corresponding to the movement operation determined by the application client in response to the movement operation for the first live video; the start position information is the coordinate position information of the first video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the first image area at the start time of the movement operation of the first video frame in the moving video frame; the end position information is the coordinate position information of the second video frame in the screen coordinate system to which the live display area belongs, recorded by the application client based on the second image area at the end time of the movement operation of the second video frame in the moving video frame.
[0255] The displacement determination unit 2002 is used to determine the coordinate position change information of the video display area to which the first live video belongs, from the first image area to the second image area, based on the start position information and the end position information, and to determine the displacement of the first live video based on the coordinate position change information.
[0256] The specific implementation methods of the coordinate position receiving unit 2001 and the displacement determining unit can be found in the above description. Figure 11 The specific process of determining the displacement in the server as described in the corresponding embodiments will not be repeated here.
[0257] The second video determination module 300 is used to determine the second video stream associated with the target object from the original acquired video stream based on the movement displacement, and return the second video stream to the application client, so that when the application client obtains the second live video based on the second video stream, it outputs the first video data in the second live video in the first display area, and also outputs the second video data in the second live video in the second display area; the second display area is the area in the live display area excluding the overlapping area; the display position of the target object in the second live video is different from the display position of the target object in the first live video.
[0258] The second video determination module 300 includes: a positioning location determination unit 3001 and a video sequence encoding unit 3002;
[0259] The positioning and location determination unit 3001 is used to determine the positioning information of the first live video after movement in the original acquired video stream based on the movement displacement, and to determine the second video sequence associated with the target object from the original acquired video stream based on the positioning and location information and terminal screen information.
[0260] The video sequence encoding unit 3002 is used to encode the second video sequence to obtain the second video stream corresponding to the second video sequence, and return the second video stream to the application client.
[0261] The specific implementation methods of the positioning and location determination unit 3001 and the video sequence encoding unit 3002 can be found in the above description. Figure 11 The specific process of obtaining the second video stream described in the corresponding embodiments will not be repeated here.
[0262] The location determination unit 3001 includes: a sequence acquisition subunit 30011, a location determination subunit 30012, a safe area determination subunit 30013, and a video sequence synthesis subunit 30014; optionally, the location determination unit 3001 further includes: a non-safe area determination subunit 30015 and a local video filling subunit 30016.
[0263] The acquisition sequence acquisition subunit 30011 is used to acquire the acquisition video sequence corresponding to the original acquisition video stream. Based on the vertex position information of the video frame in the acquisition video sequence in the positioning coordinate system, the moving safe area corresponding to the original acquisition video stream is determined. In the acquisition video sequence, the video frame with the same video frame number as the moving video frame is taken as the acquisition video frame.
[0264] The positioning determination subunit 30012 is used to determine the start and end coordinates of the moving video frame in the captured video frame based on the displacement in the positioning coordinate system. The end coordinates are used as the positioning position information of the first live video after movement in the original captured video stream. The positioning position information is compared with the movement safety area to obtain the comparison result. The image area formed by the start coordinates is the first positioning area of the first video frame in the moving video frame in the positioning coordinate system, and the image area formed by the end coordinates is the second positioning area of the second video frame in the moving video frame in the positioning coordinate system. The overlapping area between the second positioning area and the first positioning area is the first screenshot area, and the remaining area in the first positioning area excluding the overlapping area is the second screenshot area.
[0265] The safe zone determination subunit 30013 is used to determine that the first live video after the movement is located within the safe zone if the comparison result indicates that each movement end coordinate information in the positioning information is within the safe zone.
[0266] The video sequence synthesis subunit 30014 is used to, based on terminal screen information and a first screenshot method, take video data belonging to the first screenshot area in the acquired video sequence as first video data to be displayed in the first display area, and take video data belonging to the second screenshot area in the acquired video sequence as second video data to be displayed in the second display area, synthesize the first video data and the second video data, and use the synthesized video data as the second video sequence associated with the target object.
[0267] Optionally, the non-safe area determination subunit 30015 is used to determine that if the comparison result indicates that each movement end coordinate information in the positioning location information has movement end coordinate information outside the movement safe area, then the first live video after movement has a local video area located outside the movement safe area; the local video area is the area in the second positioning area composed of movement frame position information outside the movement safe area.
[0268] The local video filling subunit 30016 is used to fill the image data in the local video area into the second screenshot area based on the terminal screen information and the second screenshot method. In the video acquisition sequence, the video data belonging to the first screenshot area is used as the first video data to be displayed in the first display area, and the video data in the second screenshot area is used as the second video data to be displayed in the second display. The first video data and the second video data are combined and processed, and the combined video data is used as the second video sequence associated with the target object.
[0269] The specific implementation methods of the acquisition sequence subunit 30011, the location determination subunit 30012, the safe area determination subunit 30013, the video sequence synthesis subunit 30014, the non-safe area determination subunit 30015, and the partial video filling subunit 30016 can be found above. Figure 11 The specific process of obtaining the second video sequence described in the corresponding embodiments will not be repeated here.
[0270] Optionally, the first video stream is obtained by encoding the first video sequence using the first encoder; the video frames in the first video sequence are determined by taking screenshots of the video frames in the captured video sequence corresponding to the original captured video stream based on the terminal screen information; the video frame number of the video frames in the first video sequence is consistent with the video frame number of the video frames in the captured video sequence.
[0271] The encoder startup module 400 is used to keep the first encoder running and start the second encoder when it receives an encoding switch request sent by the application client based on the first live video after the move; the second encoder is used to encode the second video stream corresponding to the second video sequence when the second video sequence associated with the target object is obtained.
[0272] Optionally, the frame number synchronization module 500 synchronizes the frame numbers of video frames with the same video frame content in the first video sequence corresponding to the first video stream and the second video sequence corresponding to the second video stream, and uses the video frames after frame number synchronization as buffered video frames in the first video sequence.
[0273] The keyframe determination module 600 is used to determine the video frame corresponding to the motion operation in the first video sequence as the motion video frame in the first live video, and to determine the next video frame of the motion video frame as the key video frame in the second live video in the second video sequence; the frame spacing between the video frame number of the key video frame and the video frame number of the target video frame in the buffered video frames satisfies the video encoding conditions of the second video stream.
[0274] The encoder shutdown module 700 is used to shut down the first encoder based on a key video frame when it is detected that the application client has finished playing the moving video frame based on the target video frame.
[0275] The first video return module 100, the movement displacement determination module 200, and the second video determination module 300. Further, the video data processing device 2 may also include: an encoder start module 400, a frame number synchronization module 500, a keyframe determination module 600, and an encoder stop module 700. Specific implementations of these modules can be found above. Figure 11The description of the server side in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0276] Further, please see Figure 16 , Figure 16 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 16 As shown, the computer device 1000 can be a viewer terminal, and the viewer terminal can be the aforementioned Figure 1 Optionally, in the corresponding embodiment, the computer device 1000 can also be a service server, which can be the aforementioned... Figure 1 The server 2000 in the corresponding embodiment. At this time, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 16 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0277] The network interface 1004 in the computer device 1000 can also provide network communication functions, and the optional user interface 1003 can also include a display screen and a keyboard. Figure 16 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the aforementioned... Figure 3 , Figure 8 or Figure 11 The description of the video data processing method in the corresponding embodiments can also be performed as described above. Figure 14 The description of the video data processing device 1 in the corresponding embodiment can also be performed as described above. Figure 15The description of the video data processing device 2 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0278] Furthermore, it should be noted that this application embodiment also provides a computer storage medium, which stores a computer program executed by the aforementioned video data processing device 1 or video data processing device 2. The computer program includes program instructions, and when the processor executes the program instructions, it can execute the aforementioned... Figure 3 , Figure 8 or Figure 11 The description of the video data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.
[0279] It is understood that embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned... Figure 3 , Figure 8 or Figure 11 The description of the video data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.
[0280] For further details, please see Figure 17 , Figure 17 This is a schematic diagram of the structure of a video data processing system provided in an embodiment of this application. The video data processing system 3 may include a video data processing device 1a and a video data processing device 2a. The video data processing device 1a may be the one described above. Figure 14 As can be understood from the video data processing device 1 in the corresponding embodiment, this video data processing device 1a can be integrated into the above-mentioned... Figure 1 The corresponding viewer terminal 4000a in the embodiment will not be described in detail here. The video data processing device 2a can be the one described above. Figure 15The video data processing device 2 in the corresponding embodiment can be understood to be integrated into the server 2000 in the corresponding embodiment described above; therefore, it will not be described again here. Furthermore, the beneficial effects of using the same method will also not be described again. For technical details not disclosed in the embodiments of the video data processing system involved in this application, please refer to the description of the method embodiments of this application.
[0281] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0282] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A video data processing method, characterized in that, include: Output the first live video associated with the target object at the first business level of the video live streaming interface in the virtual live streaming room. In response to a movement operation on the first live video at the first service level, the first live video is moved to a first display area in the terminal screen window corresponding to the live video interface, and a second display area to be filled is determined on the live video interface. The moving video frame corresponding to the movement operation includes a first video frame at the start of the movement and a second video frame at the end of the movement. Moving the first live video means changing the live video window to which the first live video belongs at the first service level from the first image area of the first video frame at the start of the movement to the second image area of the second video frame at the end of the movement. The first display area is the overlapping area between the second image area and the terminal screen window. The second display area is the area in the terminal screen window other than the overlapping area. The start position information of the first image area in the screen coordinate system to which the terminal screen window belongs and the end position information of the second image area in the screen coordinate system are used to determine the movement displacement of the first live video. When determining a second video stream related to the target object from the original acquired video stream related to the first live video based on the movement displacement of the first live video and the positioning information of the moving video frame in the acquired video frame, the second live video composed of the first video data and the second video data is decoded based on the second video stream. When playing the second live video on the video live interface, the first video data is displayed in the first display area and the second video data is displayed in the second display area. The display position of the target object in the second live video is different from the display position of the target object in the first live video; the positioning position information is determined by the movement end coordinate information when the movement start coordinate information and movement end coordinate information of the moving video frame in the captured video frame are determined in the positioning coordinate system; the image area formed by the movement start coordinate information is the first positioning area, and the image area formed by the movement end coordinate information is the second positioning area; the overlapping area between the first positioning area and the second positioning area is the first screenshot area for capturing the first video data, and the remaining area in the first positioning area other than the overlapping area is the second screenshot area for capturing the second video data.
2. The method according to claim 1, characterized in that, The step of outputting the first live video associated with the target object at the first business level of the video live streaming interface in the virtual live streaming room includes: In response to a launch operation for an application client, the application display interface corresponding to the application client is output; In response to a trigger operation on the virtual live streaming room in the application's display interface, an access request for the virtual live streaming room is sent to the server; the access request instructs the server to extract a first video sequence from the captured video sequence corresponding to the original captured video stream that matches the terminal screen information of the application client; the original captured video stream is a video stream pulled by the server from the broadcaster client that is associated with the broadcaster user in the virtual live streaming room; the size of the video frames in the first video sequence is less than or equal to the size of the video frames in the captured video sequence; Receive the first video stream corresponding to the first video sequence returned by the server, decode the first video stream, and obtain the first live video of the target object associated with the broadcaster user; The first live video is output at the first business level of the video live streaming interface corresponding to the virtual live streaming room.
3. The method according to claim 1, characterized in that, When outputting the first live video associated with the target object at the first business level of the video live streaming interface in the virtual live streaming room, the method further includes: Video auxiliary elements associated with the first live video are output on the live video interface.
4. The method according to claim 3, characterized in that, The video live streaming interface includes a first business layer and a second business layer; the first business layer is used to display the first live video played in the virtual live streaming room; the second business layer is used to display public screen message elements, interactive elements and operation elements associated with the virtual live streaming room; The step of outputting video auxiliary elements associated with the first live video on the live video interface includes: On the video live streaming interface, identify the public screen message area, business interaction area, and content operation area associated with the second business level; Output a public broadcast message associated with the target object in the virtual live broadcast room in the public screen message area, and use the public broadcast message as the public screen message element; The interactive message elements associated with the anchor users in the virtual live broadcast room are output in the business interaction area, and the operation elements associated with the audience users in the virtual live broadcast room are output in the content operation area. The public screen message element, the interactive element, and the operation element are used as video auxiliary elements associated with the first live video.
5. The method according to claim 4, characterized in that, Prior to responding to a movement operation on the first live video at the first service tier, the method further includes: Obtain the video dragging conditions corresponding to the live video interface, and detect the video auxiliary elements on the second business level based on the video dragging conditions; When the number of video auxiliary elements is detected to meet the quantity threshold in the video dragging condition, movement instruction information for the first live video at the first business level is output on the live video interface; the movement instruction information is used to instruct the audience user in the virtual live room to perform movement operation on the first live video at the first business level.
6. The method according to claim 1, characterized in that, The step of responding to a movement operation on the first live video at the first service level, moving the first live video to a first display area in the terminal screen window corresponding to the live video interface, and determining a second display area to be filled on the live video interface, includes: In response to a motion operation on the first live video at the first service level, the video frame corresponding to the motion operation is taken as the motion video frame, and the motion duration corresponding to the motion operation is determined; the motion duration includes a motion start time and a motion end time; the motion video frame corresponding to the motion start time is the first video frame, and the video frame corresponding to the motion end time is the second video frame. During the movement duration, the live video window to which the first live video belongs is changed from a first image region to a second image region; the first image region is the live video window of the first video frame at the start of the movement; the second image region is the live video window of the second video frame at the end of the movement. Determine the overlapping area between the second image area and the terminal screen window corresponding to the video live streaming interface, and define the overlapping area as the first display area for displaying the first video data in the second live streaming video; In the terminal screen window, the remaining video missing area, excluding the overlapping area, is used as the second display area to be filled.
7. The method according to claim 6, characterized in that, The step of responding to a movement operation on the first live video at the first service level, taking the video frame corresponding to the movement operation as the movement video frame, and determining the movement duration corresponding to the movement operation, includes: Obtain the movement indication information corresponding to the first live video, and in response to the trigger operation for the movement execution information, determine the video live window to which the first live video belongs at the first business level of the video live interface, and adjust the business status of the video live window to which the first live video belongs to the movable state. Within the live video window with the movable state, in response to a movement operation on the first live video, the video frame corresponding to the movement operation is determined as a moving video frame in the first live video. Based on the movement start time corresponding to the first video frame in the moving video frame and the movement end time corresponding to the second video frame in the moving video frame, the movement duration corresponding to the movement operation is determined.
8. The method according to claim 2, characterized in that, When determining a second video stream related to the target object from the original acquired video stream related to the first live video based on the movement displacement of the first live video and the positioning information of the moving video frame in the acquired video frame, the second live video composed of first video data and second video data is decoded based on the second video stream. When playing the second live video on the live video interface, the first video data is displayed in the first display area, and the second video data is displayed in the second display area, including: Obtain the mobile video frame corresponding to the mobile operation from the first live video, and record the starting position information of the first image area corresponding to the first video frame in the screen coordinate system to which the terminal screen window belongs, based on the first image area of the first video frame at the start time of the mobile operation. Based on the second image region of the second video frame in the moving video frame at the end time of the moving operation, record the end position information of the second image region corresponding to the second video frame in the screen coordinate system to which the terminal screen window belongs; The start and end position information are sent to the server so that when the server determines the movement displacement of the first live video based on the start and end position information, it can determine a second video sequence associated with the target object from the original captured video stream based on the movement displacement and the positioning information of the moving video frame in the captured video frame; the captured video frame is a video frame with the same video frame number as the moving video frame in the captured video sequence cached by the server, and the moving video frame is obtained by the server taking a screenshot of the captured video frame based on the terminal screen information; The system receives the second video stream corresponding to the second video sequence returned by the server, decodes the second video stream to obtain the second live video corresponding to the second video sequence, and displays the first video data in the second live video in the first display area and the second video data in the second live video in the second display area when playing the second live video on the video live interface.
9. The method according to claim 8, characterized in that, The first video stream is obtained by the server encoding the first video sequence using the first encoder; the frame number of the video frame in the first video sequence is consistent with the frame number of the video frame in the acquired video sequence. The method further includes: Based on the start position information and end position information, an encoding switching request is generated for the virtual live streaming room, and the encoding switching request is sent to the server; The encoding switching request is used to instruct the server to encode the second video sequence through the second encoder when the second encoder is started, so as to obtain the second video stream corresponding to the second video sequence; the video frames in the second video sequence are determined by the server after taking screenshots of the video frames in the acquired video sequence based on the terminal screen information and the positioning information, and the display position of the target object in the second video sequence is different from the display position of the target object in the first video sequence.
10. A video data processing method, characterized in that, include: The application client receives an access request for accessing a virtual live streaming room based on terminal screen information. Based on the access request, the application client determines a first video stream that matches the terminal screen information from the original captured video stream corresponding to the virtual live streaming room. The application client then returns the first video stream to the application client so that the application client outputs a first live video associated with the target object at the first business level of the video live streaming interface of the virtual live streaming room based on the first video stream. When the application client responds to a movement operation on the first live video at the first service level and moves the first live video to the first display area in the terminal screen window corresponding to the live video interface, it receives video coordinate position information sent by the application client based on the movement of the first live video, and determines the movement displacement of the first live video based on the video coordinate position information. The moving video frame corresponding to the movement operation includes a first video frame corresponding to the start time of the movement and a second video frame corresponding to the end time of the movement. Moving the first live video means changing the live video window to which the first live video belongs at the first service level from the first image area of the first video frame at the start time of the movement to the second image area of the second video frame at the end time of the movement. The first display area is the overlapping area between the second image area and the terminal screen window. The start position information of the first image area in the screen coordinate system to which the terminal screen window belongs and the end position information of the second image area in the screen coordinate system are used to determine the movement displacement. Based on the displacement and the positioning information of the moving video frame in the acquired video frame, a second video stream associated with the target object is determined from the original acquired video stream. The second video stream is then returned to the application client, so that when the application client obtains the second live video based on the second video stream, it outputs the first video data in the second live video in the first display area and the second video data in the second live video in the second display area. The second display area is the area in the terminal screen window excluding the overlapping area. The display position of the target object in the second live video is different from the display position of the target object in the first live video; the positioning position information is determined by the movement end coordinate information when the movement start coordinate information and movement end coordinate information of the moving video frame in the captured video frame are determined in the positioning coordinate system; the image area formed by the movement start coordinate information is the first positioning area, and the image area formed by the movement end coordinate information is the second positioning area; the overlapping area between the first positioning area and the second positioning area is the first screenshot area for capturing the first video data, and the remaining area in the first positioning area other than the overlapping area is the second screenshot area for capturing the second video data.
11. The method according to claim 10, characterized in that, Receiving the video coordinate position information sent by the application client based on the movement of the first live video, and determining the movement displacement of the first live video based on the video coordinate position information, includes: The system receives video coordinate position information sent by the application client based on the movement of the first live video, and determines start and end position information associated with the moving video frame in the first live video based on the video coordinate position information; the moving video frame is the video frame corresponding to the movement operation determined by the application client in response to the movement operation on the first live video; the start position information is the coordinate position information of the first video frame in the moving video frame in the screen coordinate system to which the terminal screen window belongs, recorded by the application client based on the first image area at the start time of the movement operation; the end position information is the coordinate position information of the second video frame in the moving video frame in the screen coordinate system to which the terminal screen window belongs, recorded by the application client based on the second image area at the end time of the movement operation. Based on the start position information and end position information, determine the coordinate position change information of the video live streaming window to which the first live streaming video belongs, from the first image area to the second image area, and determine the movement displacement of the first live streaming video based on the coordinate position change information.
12. A video data processing apparatus, characterized in that, include: The first video acquisition module is used to output the first live video associated with the target object at the first business level of the video live streaming interface in the virtual live streaming room. A video movement module is configured to respond to a movement operation on the first live video at the first service level by moving the first live video to a first display area in the terminal screen window corresponding to the live video interface, and determining a second display area to be filled in the live video interface. The movement video frame corresponding to the movement operation includes a first video frame at the start time of the movement and a second video frame at the end time of the movement. Moving the first live video means changing the live video window to which the first live video belongs at the first service level from the first image area of the first video frame at the start time of the movement to the second image area of the second video frame at the end time of the movement. The first display area is the overlapping area between the second image area and the terminal screen window. The second display area is the remaining area in the terminal screen window excluding the overlapping area. The start position information of the first image area in the screen coordinate system to which the terminal screen window belongs and the end position information of the second image area in the screen coordinate system are used to determine the movement displacement of the first live video. The video display module is used to determine a second video stream related to the target object from the original acquired video stream related to the first live video based on the movement displacement of the first live video and the positioning information of the moving video frame in the acquired video frame; decode the second video stream based on the second video stream to obtain a second live video composed of first video data and second video data; when playing the second live video on the video live interface, display the first video data in the first display area and the second video data in the second display area. Both the first video data and the second video data originate from the second live video, and the display position of the target object in the second live video is different from the display position of the target object in the first live video; the positioning information is determined by the movement end coordinate information when the movement start coordinate information and movement end coordinate information of the moving video frame in the captured video frame are determined in the positioning coordinate system; the image area formed by the movement start coordinate information is the first positioning area, and the image area formed by the movement end coordinate information is the second positioning area; the overlapping area between the first positioning area and the second positioning area is the first screenshot area for capturing the first video data, and the remaining area in the first positioning area other than the overlapping area is the second screenshot area for capturing the second video data.
13. A video data processing apparatus, characterized in that, include: The first video return module is used to receive an access request for accessing a virtual live room sent by an application client based on terminal screen information, determine a first video sequence that matches the terminal screen information from the original captured video stream corresponding to the virtual live room based on the access request, and return the video stream corresponding to the first video sequence to the application client, so that the application client outputs a first live video associated with the target object at the first business level of the video live interface of the virtual live room based on the video stream corresponding to the first video sequence. The movement displacement determination module is used to receive video coordinate position information sent by the application client based on the movement of the first live video in response to a movement operation on the first live video at the first service level, and to determine the movement displacement of the first live video based on the video coordinate position information when the application client moves the first live video to the first display area of the terminal screen window corresponding to the video live interface; the movement video frame corresponding to the movement operation includes a first video frame corresponding to the start time of the movement and a second video frame corresponding to the end time of the movement; moving the first live video means changing the video live window to which the first live video belongs at the first service level from the first image area of the first video frame at the start time of the movement to the second image area of the second video frame at the end time of the movement; the first display area is the overlapping area between the second image area and the terminal screen window; the start position information of the first image area in the screen coordinate system to which the terminal screen window belongs and the end position information of the second image area in the screen coordinate system are used to determine the movement displacement. The second video determination module is used to determine a second video stream associated with the target object from the original acquired video stream based on the movement displacement and the positioning information of the moving video frame in the acquired video frame, and return the second video stream to the application client, so that when the application client obtains the second live video based on the second video stream, it outputs the first video data in the second live video in the first display area and the second video data in the second live video in the second display area; the second display area is the area in the terminal screen window other than the overlapping area; The display position of the target object in the second live video is different from the display position of the target object in the first live video; the positioning position information is determined by the movement end coordinate information when the movement start coordinate information and movement end coordinate information of the moving video frame in the captured video frame are determined in the positioning coordinate system; the image area formed by the movement start coordinate information is the first positioning area, and the image area formed by the movement end coordinate information is the second positioning area; the overlapping area between the first positioning area and the second positioning area is the first screenshot area for capturing the first video data, and the remaining area in the first positioning area other than the overlapping area is the second screenshot area for capturing the second video data.
14. A computer device, characterized in that, include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to cause the computer device to perform the method according to any one of claims 1-11.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-11.