Communication method and device
By adding animated video streams to the other end when it does not support animation effects, the problem of the other end being unable to display animation effects is solved, thus improving the fun and user experience of VoLTE/RTC video calls.
Patent Information
- Application Number
- CN202410878388.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-12-30
AI Technical Summary
In traditional VoLTE/RTC video calls, if the other end does not support animation effects, the animation effects cannot be displayed, which affects the fun of the call.
If the other end does not support motion effects, the first device adds a motion effect video stream to it and sends it together with the local video by overlaying the motion effect video, so as to ensure that the other end can display the motion effects.
It enables the display of animation effects even when the other end does not support them, thereby enhancing the fun of calls and the user experience.
Smart Images

Figure CN121239775A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication, and in particular to a communication method and device. BACKGROUND
[0002] In a conventional voice over long-term evolution (VoLTE) / real-time communication (RTC) video call, a video picture is generally a full-screen window to see the video of the opposite end, and the video of the local end is displayed in a small-screen window. By recognizing gestures / facial expressions in the video picture and adding corresponding dynamic effects in the picture, the interestingness of the call can be increased, and the experience effect is better. For example, the local end or the network side can transmit the information of the dynamic effects to the opposite end, or generate the dynamic effects by the opposite end, so that the opposite end can display the dynamic effects in the call.
[0003] However, if the opposite end does not support dynamic effects, the opposite end cannot display dynamic effects in the call process, resulting in that the interestingness of the call is affected. SUMMARY
[0004] Embodiments of the present application provide a communication method and device to enable the opposite end to display dynamic effects even if the opposite end does not support dynamic effects.
[0005] To achieve the above object, the present application adopts the following technical solutions:
[0006] In a first aspect, a communication method is provided, applied to a first device, and the method comprises: the first device receives a video of a second device; in a case where the first device supports dynamic effects and the second device does not support dynamic effects, the first device sends a video stream to the second device. The video stream is a video stream containing dynamic effects determined according to the video of the first device and a first dynamic effect video, the first dynamic effect video being a video in which the video of the second device is added with dynamic effects, or the video stream is a video stream containing dynamic effects determined according to a second dynamic effect video and the video of the second device, the second dynamic effect video being a video in which the video of the first device is added with dynamic effects.
[0007] Therefore, in a case where the second device (i.e. the opposite end) does not support dynamic effects, the first device (i.e. the local end) can add dynamic effects to the video of the second device, such as a first dynamic effect video, and send the video stream together with the video of the first device to the second device; or the first device can also add dynamic effects to its own video, such as a second dynamic effect video, and send the video stream together with the video of the second device to the second device. The second device does not need to perform processing related to dynamic effects, and can display the corresponding dynamic effects by displaying the video stream, so that the opposite end can display dynamic effects even if the opposite end does not support dynamic effects.
[0008] In a possible design, the method in the first aspect can further include: the first device receiving capability information of the second device, the capability information of the second device indicating whether the second device supports the dynamic effect, that is, dynamically determining whether the second device supports the dynamic effect through capability interaction, so as to select a corresponding video processing mode according to the capability, for example, if the dynamic effect is not supported, using the method of the present application, or if the dynamic effect is supported, using the implementation of the prior art, which is relatively more flexible. Alternatively, the first device can also default that the second device does not support the dynamic effect, and default to use the method of the present application, thereby avoiding signaling overhead caused by capability interaction.
[0009] Optionally, the method in the first aspect can further include: the first device sending a session invitation message to the second device, the session invitation message being used for the first device to request to establish a video call with the second device. The first device receiving capability information of the second device includes: the first device receiving a response message returned by the second device according to the session invitation message, the response code of the response message being 200 or 183, and the response message containing the capability information of the second device. That is, reusing signaling of a session negotiation process to deliver the capability information, which is relatively low in implementation complexity and friendly to the current protocol, and of course, the capability information can also be delivered through newly defined signaling / messages, which is not limited in particular.
[0010] Further, the session invitation message contains capability information of the first device, the capability information of the first device indicating whether the first device supports the dynamic effect. That is, the first device can also inform the peer of its own capability, thereby avoiding that the dynamic effect display fails due to misalignment of the capabilities of the two parties, for example, the peer does not cooperate with the method of the present application due to misbelief that the first device does not support the dynamic effect, thereby causing the peer to be unable to display the corresponding dynamic effect.
[0011] Further, in the case that the first device supports the dynamic effect, the capability information of the first device further indicates that the dynamic effect supported by the first device is: a dynamic effect for a video user, or a dynamic effect for a video display window.
[0012] It should be understood that the dynamic effect for a video user can mean that the dynamic effect is applied to the user, or a device (such as the first device or the second device) corresponding to the user, and the dynamic effect changes when the display mode of the video of the user changes, for example, from small window display to large window display, and the dynamic effect also changes accordingly, for example, from display on the small window to display on the large window. Similarly, the dynamic effect for a video display window can mean that the dynamic effect is applied to the display window, and the dynamic effect does not change when the display mode of the video of the user changes, for example, from large window display to small window display, and the dynamic effect is still displayed on the current large window.
[0013] In a possible design, the video stream is obtained by superimposing the picture of the first motion effect video on the picture of the video of the first device, and the picture size of the first motion effect video is smaller than the picture size of the video of the first device; or the video stream is obtained by superimposing the picture of the video of the first device on the picture of the first motion effect video, and the picture size of the video of the first device is smaller than the picture size of the first motion effect video. In this way, the effect of superimposing the motion effect is equivalent to being fused in one picture, so that the second device can only decode the video stream to display the corresponding motion effect.
[0014] Optionally, the method in the first aspect can further include: if the second device is configured to display the video of the first device, and the first device supports the motion effect for the video user, the first device superimposes the picture of the first motion effect video on the picture of the video of the first device and encodes to obtain the video stream; or if the second device is configured to display the video of the second device, and the first device supports the motion effect for the video user, the first device superimposes the picture of the video of the first device on the picture of the first motion effect video and encodes to obtain the video stream.
[0015] It should be understood that the second device can include a first window and a second window. The second device displaying the video of the first device can refer to displaying the video of the first device in the first window, and the video of the second device can be displayed in the second window at this time. The second window is contained in the first window, and it can also be considered that the large window / large screen displays the video of the first device, and the small window / small screen displays the video of the second device. Therefore, the first device can determine / adjust the superimposition mode of the video picture according to the display of the second device, such as superimposing the picture of the first motion effect video on the picture of the video of the first device, to ensure that the large window of the second device still displays the video of the first device when decoding and displaying the video stream, and the video of the second device with the motion effect is displayed in the small window.
[0016] Similarly, the second device displaying the video of the second device can refer to displaying the video of the second device in the second window, and the video of the first device can be displayed in the first window at this time. The first window is contained in the second window, and it can also be considered that the large screen / large window displays the video of the first device, and the small screen / small window displays the video of the second device. Therefore, the first device can also determine / adjust the superimposition mode of the video picture according to the display of the second device, such as superimposing the picture of the video of the first device on the picture of the first motion effect video, to ensure that the large window of the second device still displays the video of the first device when decoding the video stream, and the video of the second device with the motion effect is displayed in the small window.
[0017] Thus, in the case that the motion effect is applied to the video user, the first device determines / adjusts the superimposition manner of the video picture according to the display condition of the second device, so that the motion effect can change with the user when the display window switches, such as switching from the second window being contained in the first window to the first window being contained in the second window, i.e., the video of the second device switches from small window display to large window display, and the motion effect changes with the user, realizing seamless switching of the motion effect and better user experience.
[0018] Optionally, the method of the first aspect can further include that the first device sends information indicating the superimposition region to the second device; if the video stream is obtained by superimposing the picture of the first motion effect video on the picture of the video of the first device, the superimposition region is a region occupied by the picture of the first motion effect video in the picture of the video of the first device; or if the video stream is obtained by superimposing the picture of the video of the first device on the picture of the first motion effect video, the superimposition region is a region occupied by the picture of the video of the first device in the picture of the first motion effect video. Thus, the second device can determine the location of the superimposition region in the large window according to the information indicating the superimposition region, so that the user can operate (such as clicking or sliding, etc.) in the superimposition region, and the second device can trigger switching of the window according to the location of the superimposition region, or in other words, switching of the display window.
[0019] In a possible design, the video stream is obtained by superimposing the picture of the first motion effect video on the picture of the video of the first device, and the picture size of the first motion effect video is smaller than the picture size of the video of the first device; or the video stream is obtained by superimposing the picture of the video of the second device on the picture of the second motion effect video, and the picture size of the video of the second device is smaller than the picture size of the second motion effect video. In this way, it is equivalent to superimposing the effect of fusing the motion effect in one picture, so that the opposite end can only decode the video stream to display the corresponding motion effect.
[0020] Optionally, the method of the first aspect can further include that if the second device is used to display the video of the first device, and the first device supports the motion effect for the video display window, the first device superimposes the picture of the first motion effect video on the picture of the video of the first device and encodes to obtain the video stream; or if the second device is used to display the video of the second device, and the first device supports the motion effect for the video display window, the first device superimposes the picture of the video of the second device on the picture of the second motion effect video and encodes to obtain the video stream. Similar to the above, in the case that the motion effect is applied to the display window, the first device determines / adjusts the superimposition manner of the video picture according to the display condition of the second device, so that the motion effect does not change with the user when the display window switches, such as still being in the current large window display when the first window is contained in the second window, to realize the effect of the motion effect following the window.
[0021] Optionally, the method described in the first aspect may further include: the first device sending information indicating an overlay area to the second device; if the video stream is obtained by overlaying the image of a first motion effect video onto the image of the first device's video, then the overlay area is the area occupied by the image of the first motion effect video within the image of the first device's video; or; if the video stream is obtained by overlaying the image of the second device's video onto the image of a second motion effect video, then the overlay area is the area occupied by the image of the second device's video within the image of the second motion effect video. Thus, the second device can determine the location of the overlay area within the large window based on the information indicating the overlay area, so that the user can operate within the overlay area, and the second device can trigger a display window switch based on its location.
[0022] Optionally, the method described in the first aspect may further include: the first device receiving instruction information from the second device, the instruction information instructing the second device to display the video of the first device, or the instruction information instructing the second device to display the video of the second device, so that the first device can determine the corresponding video overlay method accordingly to avoid errors in the animation display.
[0023] In one possible design, the first device supports motion effects for video users, and the video stream includes a first video stream and a second video stream; the first video stream is the video stream of the first device, and the second video stream is the video stream of the first motion effect.
[0024] Optionally, the first device sends a video stream to the second device, including: the first device sending a first video stream to the second device through a first transmission channel, and sending a second video stream to the second device through a second transmission channel. For the second device, it receives the first video stream from the first device through the first transmission channel, receives the second video stream from the first device through the second transmission channel, displays the video from the first video stream in a first window, and displays the video from the second video stream in a second window. In this case, regardless of how the containment relationship between the first and second windows changes or how they switch, the animation effects can change with the user, achieving seamless switching of animation effects and a better user experience.
[0025] In one possible design, the first device supports animation effects for the video display window, and the video stream includes a first video stream and a second video stream, wherein the first video stream is obtained by encoding the video of the first device, and the second video stream is obtained by encoding the first animation effect video; or, the video stream includes a third video stream and a fourth video stream, wherein the third video stream is obtained by encoding the video of the second device, and the fourth video stream is obtained by encoding the second animation effect video.
[0026] Optionally, if the second device is used to display the video of the second device, then the first device sends a first video stream to the second device through a first transmission channel and sends a second video stream to the second device through a second transmission channel; or, if the second device is used to display the video of the first device, then the first device sends a fourth video stream to the second device through a first transmission channel and sends a third video stream to the second device through a second transmission channel.
[0027] For the second device, it can receive a first video stream from the first device through the first transmission channel, and a second video stream from the first device through the second transmission channel. It can also display the video from the first video stream in a first window and the video from the second video stream in a second window. Alternatively, the second device can receive a fourth video stream from the first device through the first transmission channel, and a third video stream from the first device through the second transmission channel. It can also display the video from the fourth video stream in a first window and the video from the third video stream in a second window. In this case, if the containment relationship between the first and second windows changes, by adjusting the transmission of the video stream containing the animation to different transmission channels (e.g., changing the transmission of the second video stream (including the animation) through the second transmission channel to the transmission of the fourth video stream (including the animation) through the first transmission channel), the animation can remain unchanged with the user, such as still being displayed in the current large window, thus achieving the effect of the animation following the window.
[0028] Optionally, the method described in the first aspect may further include: the first device receiving instruction information from the second device, the instruction information instructing the second device to display the video of the first device, or the instruction information instructing the second device to display the video of the second device, so that the first device can use the corresponding stream to transmit the corresponding video accordingly, thereby avoiding errors in the animation display due to transmission errors.
[0029] Optionally, the first transmission channel is a transmission channel for video calls negotiated and established between the first device and the second device. The second transmission channel is a transmission channel negotiated and established between the first device and the second device when the second device does not support animation effects. That is, it is an additional transmission channel established when the second device does not support animation effects, so as to avoid redundant overhead caused by maintaining the second transmission channel when it is not needed.
[0030] Secondly, a communication method is provided, applied to a second device, the method comprising: the second device sending a video stream of the second device to a first device; the second device receiving a video stream from the first device and displaying the video in the video stream. Wherein, if the first device supports motion effects but the second device does not, the video stream is determined based on the video of the first device and a first motion effect video, wherein the first motion effect video is a video of the second device with motion effects added; or the video stream is determined based on a second motion effect video and the video of the second device, wherein the second motion effect video is a video of the first device with motion effects added.
[0031] In one possible design, the method described in the second aspect may further include: the second device sending capability information of the second device to the first device, the capability information of the second device indicating whether the second device supports motion effects.
[0032] Optionally, the method in the second aspect may further include: the second device receiving a session invitation message from the first device, the session invitation message being used by the first device to request to establish a video call with the second device; the second device sending capability information of the second device to the first device, including: the second device sending a response message to the first device according to the session invitation message, the response message having a response code of 200 or 183, and the response message containing capability information of the second device.
[0033] Furthermore, the session invitation message contains capability information of the first device, which indicates whether the first device supports motion effects.
[0034] Furthermore, if the first device supports motion effects, the capability information of the first device also indicates that the motion effects supported by the first device are: motion effects for video users, or motion effects for video display windows.
[0035] In one possible design, the second device includes a first window and a second window, and the video stream is a single video stream. The second device displays the video in the video stream, including: if the second window is contained within the first window, or in other words, the first window contains the video of the second window, then the second device displays the video in the video stream in the first window and hides the second window; if the first window is contained within the second window, or in other words, the second window contains the video of the first window, then the second device displays the video in the video stream in the second window and hides the first window, so as to avoid the display of the first window affecting the animation experience.
[0036] Optionally, the method in the second aspect may further include: the second device sending instruction information to the first device; if the second window is contained within the first window, the instruction information instructs the second device to display the video of the first device; if the first window is contained within the second window, the instruction information instructs the second device to display the video of the second device.
[0037] Furthermore, the second device sends instruction information to the first device, including: in response to a user's window switching operation triggered within the overlay area, the second device sends instruction information to the first device; if the overlay area is the area occupied by the image of the first motion effect video in the video image of the first device, or the area occupied by the video image of the second device in the image of the second motion effect video, then the window switching operation instruction changes from the second window being contained in the first window to the first window being contained in the second window; if the overlay area is the area occupied by the video image of the first device in the image of the first motion effect video, then the window switching operation instruction changes from the first window being contained in the second window to the second window being contained in the first window.
[0038] Furthermore, the method described in the second aspect may further include: the second device receiving information from the first device for indicating the overlay area.
[0039] In one possible design, the second device includes a first window and a second window, and the video stream includes a first video stream and a second video stream. The first video stream is obtained by encoding the video of the first device, and the second video stream is obtained by encoding the first motion effect video. The second device receives the video stream from the first device, including: the second device receives the first video stream from the first device through a first transmission channel and receives the second video stream from the first device through a second transmission channel; the second device displays the video in the video stream, including: the second device displays the video in the first video stream in the first window and displays the video in the second video stream in the second window.
[0040] In one possible design, the second device includes a first window and a second window. If the video stream includes a first video stream and a second video stream, where the first video stream is obtained by encoding the video of the first device and the second video stream is obtained by encoding a first motion effect video, the second device receiving the video stream from the first device includes: the second device receiving the first video stream from the first device through a first transmission channel and receiving the second video stream from the first device through a second transmission channel; the second device displaying the video in the video stream includes: the second device displaying the video in the first video stream in a first window and displaying the video in the second video stream in a second window, wherein the first window is contained within the second window; if the video stream includes a third video stream and a fourth video stream, where the third video stream is obtained by encoding the video of the second device and the fourth video stream is obtained by encoding a second motion effect video, the second device receiving the video stream from the first device includes: the second device receiving the fourth video stream from the first device through a first transmission channel and receiving the third video stream from the first device through a second transmission channel; the second device displaying the video in the video stream includes: the second device displaying the video in the fourth video stream in a first window and displaying the video in the third video stream in a second window, wherein the second window is contained within the first window.
[0041] Optionally, the method in the second aspect may further include: the second device sending instruction information to the first device; if the second window is contained within the first window, the instruction information instructs the second device to display the video of the first device; if the first window is contained within the second window, the instruction information instructs the second device to display the video of the second device.
[0042] Furthermore, the second device sends instruction information to the first device, including: in response to a user-triggered window switching operation, the second device sends instruction information to the first device; the window switching operation indicates switching from the second window being contained in the first window to the first window being contained in the second window, or switching from the first window being contained in the second window to the second window being contained in the first window.
[0043] It is understandable that the technical effects of the method described in the second aspect can also refer to the relevant introduction of the method described in the first aspect above, and will not be repeated here.
[0044] Thirdly, a communication device is provided. This communication device is used to perform the communication method described in either the first or second aspect.
[0045] In this application, the communication device described in the third aspect can be a terminal device or a network device, or a chip (system) or other component or assembly, or a device containing the terminal device or network device. The aforementioned chip (system) or other component or assembly can all be disposed within the terminal device or network device.
[0046] It should be understood that the communication apparatus described in the third aspect includes modules, units, or means that implement the communication method described in either the first or second aspect. These modules, units, or means can be implemented in hardware, software, or by hardware executing corresponding software. The hardware or software includes one or more modules or units for performing the functions involved in the aforementioned communication method.
[0047] Fourthly, a communication device is provided. The communication device includes a processor configured to execute the communication method described in either the first or second aspect.
[0048] In one possible design, the communication device described in the fourth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the fourth aspect and other communication devices.
[0049] In one possible design, the communication device described in the fourth aspect may further include a memory. This memory may be integrated with the processor or disposed separately. The memory may be used to store computer programs and / or data related to the communication method described in either the first or second aspect.
[0050] In this application, the communication device described in the fourth aspect can be a terminal device or a network device, or a chip (system) or other component or assembly, or a device containing the terminal device or network device. The aforementioned chip (system) or other component or assembly can all be disposed within the terminal device or network device.
[0051] Fifthly, a communication device is provided. The communication device includes a processor coupled to a memory, the processor executing a computer program stored in the memory, such that the communication device performs the communication method described in either the first or second aspect.
[0052] In one possible design, the communication device described in the fifth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the fifth aspect and other communication devices.
[0053] In this application, the communication device described in the fifth aspect can be a terminal device or a network device, or a chip (system) or other component or assembly, or a device containing the terminal device or network device. The aforementioned chip (system) or other component or assembly can all be disposed within the terminal device or network device.
[0054] A sixth aspect provides a communication device, comprising: a processor and a memory; the memory being used to store a computer program, which, when executed by the processor, causes the communication device to perform the communication method described in either the first or second aspect.
[0055] In one possible design, the communication device described in the sixth aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the sixth aspect and other communication devices.
[0056] In this application, the communication device described in the sixth aspect can be a terminal device or a network device, or a chip (system) or other component or assembly, or a device containing the terminal device or network device. The aforementioned chip (system) or other component or assembly can all be disposed within the terminal device or network device.
[0057] A seventh aspect provides a communication device comprising: a processor; the processor being configured to be coupled to a memory, and after reading a computer program from the memory, to execute a communication method as described in any implementation of the first or second aspect according to the computer program.
[0058] In one possible design, the communication device described in the seventh aspect may further include a transceiver. This transceiver may be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the communication device described in the seventh aspect and other communication devices.
[0059] In this application, the communication device described in the seventh aspect can be a terminal device or a network device, or a chip (system) or other component or assembly, or a device containing the terminal device or network device. The aforementioned chip (system) or other component or assembly can all be disposed within the terminal device or network device.
[0060] Eighthly, a processor is provided. The processor is used to execute the communication method described in either the first or second aspect.
[0061] Ninthly, a communication system is provided. The communication system includes one or more terminal devices and one or more network devices.
[0062] A tenth aspect provides a computer-readable storage medium comprising: a computer program or instructions; which, when executed on a computer, causes the computer to perform the communication method described in any possible design of the first or second aspect.
[0063] Eleventhly, a computer program product is provided, comprising a computer program or instructions that, when executed on a computer, cause the computer to perform the communication method described in any possible design of the first or second aspect.
[0064] Furthermore, the technical effects of the communication devices described in the third to eleventh aspects above can be referred to the technical effects of the communication methods described in the first or second aspects above, and will not be repeated here. Attached Figure Description
[0065] Figure 1 Illustration of video call application scenarios Figure 1 ;
[0066] Figure 2 Illustration of video call application scenarios Figure 2 ;
[0067] Figure 3 Schematic diagram of the communication system architecture provided in the embodiments of this application Figure 1 ;
[0068] Figure 4 Schematic diagram of the communication system architecture provided in the embodiments of this application Figure 2 ;
[0069] Figure 5 Flowchart of the communication method provided in the embodiments of this application Figure 1 ;
[0070] Figure 6 This application scenario illustrates the application scenarios of the communication method provided in the embodiments of this application. Figure 1 ;
[0071] Figure 7 This application scenario illustrates the application scenarios of the communication method provided in the embodiments of this application. Figure 2 ;
[0072] Figure 8 Flowchart of the communication method provided in the embodiments of this application Figure 2 ;
[0073] Figure 9 Flowchart of the communication method provided in the embodiments of this application Figure 3 ;
[0074] Figure 10 Flowchart of the communication method provided in the embodiments of this application Figure 4 ;
[0075] Figure 11 Flowchart of the communication method provided in the embodiments of this application Figure 5 ;
[0076] Figure 12 Schematic diagram of the communication device provided in the embodiments of this application Figure 1 ;
[0077] Figure 13 Schematic diagram of the communication device provided in the embodiments of this application Figure 2 . Detailed Implementation
[0078] The technical solutions of this application embodiment can be applied to various communication systems, such as Wi-Fi wireless network systems, vehicle-to-everything (V2X) communication systems, device-to-device (D2D) communication systems, vehicle-to-everything (V2X) communication systems, fourth-generation (4G) mobile communication systems, such as long-term evolution (LTE) systems, worldwide interoperability for microwave access (WiMAX) communication systems, fifth-generation (5G) mobile communication systems, such as new radio (NR) systems, and future communication systems.
[0079] The technical terms and related technical solutions in this application will be described below with reference to the accompanying drawings.
[0080] In traditional Voice Over Long-Term Evolution (VoLTE) / Real-Time Communication (RTC) video calls, the video feed typically shows the other end's video in a full-screen window, while the local video is displayed in a smaller window. By recognizing gestures and facial expressions in the video feed and adding corresponding animations, the call can be made more engaging and the overall experience better.
[0081] In one possible design, such as Figure 1 As shown, both parties in the call can display animations on their respective large screens. For example, if a user makes a heart gesture on their local screen, the local device recognizes the gesture and adds a heart-shaped animation. Simultaneously, the local device overlays the heart-shaped animation onto its own camera feed and transmits it to the other end via the call media channel, where the other end can also see the heart-shaped animation on their large screen. Another possible design, such as... Figure 2 As shown, both parties in a call can display their local animations within their respective windows. For example, if one party makes a gesture, the animation is displayed in a small window on that party, while the other party can display the animation in a larger window. The local party supports previewing various animation modes (including stickers or real-time animations), and users can select a mode. The local party can transmit its animations to the other party as a video stream or as an identifier (ID) for display. If the local party transmits the identifier, the other party needs to retrieve the corresponding animation based on the identifier and then overlay it onto the video stream for display.
[0082] However, neither of the above designs takes into account whether the other end supports animation effects. If the other end does not support animation effects, it usually cannot display animation effects during the call, which affects the fun of the call.
[0083] To address the aforementioned technical problems, this application proposes the following technical solutions. The technical solutions in this application will now be described in conjunction with the accompanying drawings.
[0084] This application will present various aspects, embodiments, or features relating to systems that may include multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.
[0085] Furthermore, in the embodiments of this application, words such as "exemplarily" and "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as an "example" in this application should not be construed as being better or more advantageous than other embodiments or designs. Rather, the use of the word "example" is intended to present the concept in a specific manner.
[0086] First, in this application, "for indicating" can include both direct and indirect indication. When describing "information" for indicating A, it can include whether the information directly indicates A or indirectly indicates A, but does not necessarily mean that the information carries A.
[0087] The information indicated by a given piece of information is called the information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as, but not limited to, directly indicating the information to be indicated, such as the information to be indicated itself or its index. It can also be indirectly indicated by indicating other information, where there is a relationship between the other information and the information to be indicated. It can also indicate only a part of the information to be indicated, while the other parts are known or pre-agreed upon. For example, the indication of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing the indication overhead to some extent. At the same time, common parts of various pieces of information can be identified and indicated uniformly to reduce the indication overhead caused by individually indicating the same information.
[0088] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be repeated here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In the specific implementation process, the required indication method can be selected according to specific needs. This application embodiment does not limit the selected indication method; therefore, the indication methods involved in this application embodiment should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.
[0089] The information to be instructed can be sent as a whole or divided into multiple sub-information messages, and the sending period and / or timing of these sub-information messages can be the same or different. This application does not limit the specific sending method. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the transmitting device by sending configuration information to the receiving device. This configuration information can include, for example, but not limited to, one or a combination of at least two of radio resource control (RRC) signaling, medium access control (MAC) layer signaling, and physical layer signaling. MAC layer signaling includes, for example, a MAC control element (CE); physical (PHY) layer signaling includes, for example, downlink control information (DCI).
[0090] Second, in the embodiments shown below, the first, second, and various numerical designations are merely distinctions for descriptive convenience and are not intended to limit the scope of the embodiments of this application. For example, to distinguish different indication information.
[0091] Third, "pre-defined," "pre-configured," or "pre-specified" can be achieved by pre-saving corresponding codes, tables, or other means of indicating relevant information in the device (e.g., including terminal devices and network devices), or by pre-defining them in a protocol. This application does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into the encoder, decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or communication device. The type of memory can be any form of storage medium, and this application does not limit this.
[0092] Fourth, the “protocol” involved in the embodiments of this application may refer to standard protocols in the field of communication, such as 3GPP’s LTE protocols (such as technical specification (TS) 36, i.e., the TS36 series of technical specifications), NR protocols (such as the TS38 series of technical specifications), and related protocols applied to future communication systems. This application does not limit this.
[0093] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0094] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0095] To facilitate understanding of the embodiments of this application, let's first take... Figure 3 The communication system illustrated herein is used as an example to illustrate a communication system applicable to embodiments of this application. For example, Figure 3 This is a schematic diagram of the architecture of a communication system to which the method provided in the embodiments of this application applies.
[0096] For example, a network device may include a first device and a second device.
[0097] Both the first device and the second device can be terminal equipment.
[0098] Terminal equipment can be a terminal with transceiver capabilities, or it can be a chip or chip system installed in the terminal equipment. This terminal equipment can also be referred to as user equipment (UE), access terminal, subscriber unit, user station, mobile station (MS), mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device. The terminal devices in the embodiments of this application may be mobile phones, cellular phones, smartphones, tablets, wireless data cards, personal digital assistants (PDAs), wireless modems, handsets, laptop computers, machine-type communication (MTC) terminals, computers with wireless transceiver capabilities, virtual reality (VR) terminals, augmented reality (AR) terminals, smart home devices (e.g., refrigerators, televisions, air conditioners, electricity meters, etc.), intelligent robots, robotic arms, workshop equipment, wireless terminals in autonomous driving, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in telemedicine, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, vehicle-mounted terminals, and roadside units with terminal functions. The terminal device in this application can also be an onboard module, onboard unit, onboard component, onboard chip, or onboard unit built into a vehicle as one or more components or units. The terminal device can also be other devices with terminal functions; for example, it can be a device that performs terminal functions in D2D communication. The embodiments of this application do not limit the device form of the terminal device. The device used to implement the terminal function can be a terminal device; it can also be a device that supports the terminal in implementing the function, such as a chip system. This device can be installed in the terminal or used in conjunction with the terminal. In the embodiments of this application, the chip system can be composed of chips or include chips and other discrete devices.
[0099] The first device and the second device can make voice or video calls through the Internet Protocol (IP) Multimedia Subsystem (IMS).
[0100] like Figure 4 As shown, the IMS network may include: a telephony application server (TAS), a proxy-call session control function (P-CSCF) entity, a serving-call session control function (S-CSCF) entity, an IMS access gateway (IMS-AGW), a transition gateway (TrGW), an interconnection border control function (IBCF) entity, a breakout gateway control function (BGCF) entity, a media gateway control function (MGCF) entity, etc.
[0101] TAS provides voice and multimedia calling services for users of fixed and mobile converged networks, supports related basic and supplementary services, integrates fixed and mobile converged services on the same platform, and provides a unified service experience for fixed and mobile network users.
[0102] The S-CSCF entity is the central node of the IMS network, responsible for user registration, authentication, sessions, routing, and service triggering.
[0103] The P-CSCF entity is the entry point node for Session Initiation Protocol (SIP) users to access the IMS network, and is mainly responsible for forwarding SIP signaling between SIP users and the home network.
[0104] IMS-AGW is the IMS access gateway, primarily responsible for media plane interoperability between the user and network interfaces.
[0105] TrGW is the IMS interconnection gateway, responsible for media plane interconnection between network interfaces.
[0106] The IBCF entity is primarily used to enable interoperability between the IMS network and other IMS network control planes. For example, if the calling party is on China Mobile's IMS network and the called party is on China Telecom's IMS network, the BGCF entity is responsible for selecting an MGCF entity for the call to connect to the CS network when the calling party is an IMS user and the called party is a circuit-switched (CS) network user.
[0107] The MGCF entity is primarily used to enable interoperability between the IMS network and the control plane of other non-IP networks (such as the public switched telephone network, PSTN).
[0108] It is understandable that other networks besides IMS (such as CS or IMS) can also be referred to as B party when acting as the caller or the called party in real-time audio and video communication.
[0109] The first and second devices can also make voice or video calls through over-the-top (OTT) services, i.e., video services based on the open internet. Alternatively, in future scenarios, the first and second devices can make voice or video calls through any possible means, without specific restrictions.
[0110] In this communication system, the first device (i.e., the local end) can acquire the capabilities of the second device (i.e., the peer end), such as if the second device does not support animation effects. In this case, the first device can add animation effects to the second device's video, such as a first animation effect video, and send it to the second device along with the first device's video stream; alternatively, the first device can also add animation effects to its own video, such as a second animation effect video, and send it to the second device along with the second device's video stream. The second device does not need to perform any animation effect-related processing; it can display the corresponding animation effects simply by displaying the video stream, thus enabling the peer end to display animation effects even when the peer end does not support them.
[0111] It should be understood that the communication method provided in the embodiments of this application can be applied to... Figure 3-4 The devices shown, such as the first device and the second device, can be specifically implemented as described in the following method embodiments, which will not be repeated here. The solutions in the embodiments of this application can also be applied to other communication systems, and the corresponding names can be replaced by the names of the corresponding functions in other communication systems.
[0112] It should also be understood that Figure 3-4 This is a simplified diagram for ease of understanding only. The communication system may also include other network devices and / or other terminal devices. Figure 3-4 It was not drawn in the middle.
[0113] The following will combine Figure 5 This application provides a detailed description of the interaction process between devices in the aforementioned communication system through method embodiments. The communication method provided in this application can be applied to the aforementioned communication system, such as the interaction between a first device and a second device, which will be described in detail below.
[0114] like Figure 5 As shown, the flow of this communication method is as follows:
[0115] S501, the second device sends the video of the second device to the first device, and the first device receives the video of the second device.
[0116] The first device (local end) and the second device (peer end) can conduct video calls. The video from the second device can be captured during the video call, such as video of the peer user participating in the video call captured by the second device's camera. The second device can encode the video and send it to the first device.
[0117] S502, if the first device supports motion effects but the second device does not, the first device sends a video stream to the second device. The second device receives the video stream from the first device.
[0118] Animation effects refer to animated effects, such as displaying hearts, cakes, small animals, small icons, fireworks, etc. Animation effects can be applied to videos, such as adding and displaying animation effects during video calls. In this case, the types of animation effects can include: animation effects for the video user and animation effects for the video display window, which will be introduced separately below.
[0119] Animation effects for video users refer to effects that apply to the user, or the user's device (such as a first device or a second device). When the display method of the user's video changes, such as from a small window to a large window, the animation effect should also change accordingly; for example, the animation effect displayed in a small window changes to being displayed in a large window. Figure 6As shown, the first device is the local end, denoted as B, and the second device is the remote end, denoted as A. When the user at the remote end A performs animated actions, such as making a heart shape or smiling, the video call can display the corresponding animated effect. For example, the local end B displays A (animated effect) in a large window, meaning the remote end A contains the animated video; the local end B displays B in a small window, meaning the local end B does not contain the animated video. Before the window at the remote end A switches, the remote end A displays B in a large window and A (animated effect) in a small window. When the window size at the remote end A switches, if the remote end's video changes to a large window display at the remote end A, correspondingly, the local end B's video changes to a small window display at the remote end A. Since the animated effect is for the remote end A, or rather, for the user at the remote end A, the display of the animated effect must change accordingly; that is, the remote end A displays A (animated effect) in a large window and B in a small window.
[0120] Animation effects for video display windows refer to effects that apply to the display window itself. When the user's video display method changes, such as from a large window to a small window, the animation effect remains unchanged, continuing to display on the current large window. For example, ... Figure 6 As shown, when a user on the other end (A) performs a motion effect, the video call can display the corresponding motion effect. For example, the local end (B) displays A (motion effect) in a large window, and B in a small window, meaning the local end (B) does not contain the video with the motion effect. If the motion effect applies to the large window, then before the window on the other end (A) switches, the other end (A) also displays B (motion effect) in a large window, meaning the local end (B) contains the video with the motion effect, while the small window on the other end (A) displays A, meaning the other end (A) does not contain the video with the motion effect. When the window size on the other end (A) switches, if the video on the other end (A) changes to a large window display, correspondingly, the video on the local end (B) changes to a small window display. Since the motion effect applies to the large window, whichever end's video is displayed in a large window on the other end (A) will display the motion effect in that large window; that is, the large window on the other end (A) displays A (motion effect), and the small window on the other end (A) displays B.
[0121] The first device supporting motion effects means that the first device has the ability to add motion effects to the video during a video call, specifically supporting motion effects for the video user and / or motion effects for the video display window. The second device not supporting motion effects means that the first device does not have the ability to add motion effects to the video during a video call.
[0122] The first and second devices can learn about each other's capabilities through interaction.
[0123] The second device can send its capability information to the first device, and correspondingly, the first device can receive the capability information of the second device. The capability information of the second device can indicate whether the second device supports motion effects, specifically indicating that the second device does not support motion effects. That is, the first device dynamically determines whether the second device supports motion effects through capability interaction, so as to select the appropriate video processing method according to its capabilities. If it does not support motion effects, the method of this application is used; if it does support motion effects, the implementation of existing technology is used, which is relatively more flexible.
[0124] For example, in a session negotiation process, the first device can send a session invitation message to the second device, and the second device can receive the session invitation message from the first device. The session invitation message can be used by the first device to request to establish a video call with the second device. The second device can send a response message to the first device based on the session invitation message, and the first device can receive the response message returned by the second device based on the session invitation message. The response code of the response message is 200 or 183, also known as a SIP message, such as SIP 200 OK or SIP 183, indicating that the second device is responding to the first device's request to establish a video call. Alternatively, it can be other response codes; there are no specific restrictions. Furthermore, the response message can contain capability information of the second device, such as in the information elements of the response message or by adding fields directly to the response message. Taking the information element implementation as an example, the information element can be a Session Description Protocol (SDP). The SDP carries the capability information of the second device. Specifically, new fields can be added to the SDP, such as `a = Animation`, whose values can be 0, 1, or 2. For example, `Animation: 0` indicates that animation effects are not supported, `Animation: 1` indicates that animation effects for video users are supported, and `Animation: 2` indicates that animation effects for the video display window are supported. That is, a single field jointly indicates supported animation effects and the specific type of animation effect supported. Alternatively, they can be indicated separately. For example, one field is `a = Animation`, where `Animation: 0` indicates that animation effects are not supported, and `Animation: 1` indicates that animation effect 1 is supported. Another field is `b = Animation type`, whose values can be 0 or 1. `Animation type: 0` indicates that animation effects for video users are supported, and `Animation type: 1` indicates that animation effects for the video display window are supported. Or, any other possible implementation methods are possible, without specific restrictions. For the second device, `Animation: 0` in the SDP indicates that the second device does not support animation effects.
[0125] It can be seen that the second device can transmit capability information by reusing the signaling of the session negotiation process, which has relatively low implementation complexity and is more friendly to the current protocol. Of course, it can also be implemented by newly defined signaling / messages, and there are no specific restrictions.
[0126] Optionally, the aforementioned session invitation message may also include capability information of the first device. This capability information indicates whether the first device supports animation effects. If the first device supports animation effects, it may specifically indicate whether the supported animation effects are: animation effects for video users or animation effects for the video display window. That is, the first device can also inform the other end of its capabilities to avoid animation effect display failure due to mismatched capabilities. For example, if the other end mistakenly believes that the first device does not support animation effects and refuses to cooperate in implementing the method of this application, it may be unable to display the corresponding animation effects.
[0127] For example, a field indication can be added to the information element of the session invitation message or directly to the session invitation message. Taking the implementation of the information element as an example, the information element can also be an SDP, which carries the capability information of the first device. Specifically, a new field joint indication can be added to the SDP, such as a = Animation:1 / 2, indicating that the first device supports animation effects for video users or animation effects for video display windows. Alternatively, it can be indicated separately by different fields, such as Animation:1 + Animation type:0, indicating that animation effects are supported, specifically animation effects for video users, or Animation:1 + Animation type:1, indicating that animation effects are supported, specifically animation effects for video display windows.
[0128] It is understandable that the session invitation message and response message can also be replaced with any other possible message. For example, in the case of voice or video calls via OTT, the session invitation message and response message can also be replaced with the corresponding messages in OTT, such as APP request message and APP response message, or other named / type messages. The specific implementation is similar to the above, and can be referred to for understanding, so it will not be elaborated here.
[0129] Optionally, the second device can determine whether it needs to execute the method of this application embodiment based on the capability information received from the first device, such as S503 below; otherwise, it can execute the existing process to achieve backward compatibility.
[0130] It is understood that the above-mentioned capability interaction is only an example. For example, the first device may assume that the second device does not support animation, and the second device may assume that the first device supports animation, and the method of this application will be used by default to avoid the signaling overhead caused by capability interaction.
[0131] The video stream can be a video stream containing motion effects determined based on the video from the first device and the first motion effect video, where the first motion effect video is a video with motion effects added to the video from the second device. Alternatively, the video stream can also be a video stream containing motion effects determined based on the second motion effect video and the video from the second device, where the second motion effect video is a video with motion effects added to the video from the first device. These scenarios will be described below.
[0132] Scenario 1:
[0133] For video user animations, such as animations for a second device or a user of a second device, regardless of how the display window of the second device switches, the video stream can be a video stream containing animations determined based on the video of the first device and the first animation video. That is, animations are always added to the video of the second device to achieve animations following the user.
[0134] The first possible design is that the video stream can be a single stream.
[0135] When the second device displays the video from the first device, the single-stream video can be obtained by overlaying the image of the first motion effect video onto the image of the first device's video, and the image size of the first motion effect video is smaller than the image size of the first device's video. Alternatively, when the second device displays the video from the second device, the single-stream video can be obtained by overlaying the image of the first device's video onto the image of the first motion effect video, and the image size of the first device's video is smaller than the image size of the first motion effect video.
[0136] In one example, the second device may include a first window and a second window. The second device displaying the video of the first device can mean that the video of the first device is displayed in the first window; that is, the video of the second device can be displayed in the second window, and the second window is contained within the first window. In other words, the first window is a large window, and the second window is a small window. Alternatively, it can be considered that the large window / large screen displays the video of the first device, and the small window / small screen displays the video of the second device itself. Similarly, the second device displaying the video of the second device can mean that the video of the second device is displayed in the second window. In this case, the video of the first device can be displayed in the first window, and the first window is contained within the first window. In other words, the second window is a large window, and the first window is a small window. Alternatively, it can be considered that the large window / large screen displays the video of the second device itself, and the small window / small screen displays the video of the first device.
[0137] It should be understood that the second device can switch between displaying the video of the first device and the video of the second device. For example, when the second device establishes a video call with the first device, it can display the video of the first device, and then switch to display the video of the second device by changing the size of the window. The second device can inform the first device which device's video is currently displayed in the large window when establishing a video call with the first device or when changing the size of its own display window. For example, the second device can send an instruction message to the first device, and the first device can receive the instruction message from the second device. This instruction message can contain 1 bit of information, with its 0 / 1 value indicating whether the second device is displaying the video of the first device or the video of the second device. For example, if the second window is contained within the first window (i.e., the large window is used to display the video of the first device), the instruction message can instruct the second device to display the video of the first device; if the first window is contained within the second window (i.e., the large window is used to display the video of the second device), the instruction message can instruct the second device to display the video of the second device. Furthermore, the instruction message can be carried in any possible signaling between the second device and the first device, and the specific implementation is not limited.
[0138] The first device can determine the display status of the second device based on the above-mentioned instruction information, thereby determining / adjusting the superposition method of the video images to avoid errors in the animation display.
[0139] For example, if the second device is used to display the video from the first device, and the first device supports motion effects for video users, then the first device can overlay the image of the first motion effect video onto the image of its own video (i.e., peer A (motion effect) + local B, where "+" indicates overlay) and encode it to obtain a video stream. The first device can also overlay the image of its own video onto the image of the first motion effect video (i.e., local B + peer A (motion effect)) and display it in a large window on the first device. Alternatively, the first device can use other display methods, such as displaying only the first motion effect video or only the video from the first device; no specific limitation is imposed. The size at which the image of the first motion effect video is overlaid onto the image of the first device's video, or the size at which the image of the first device's video is overlaid onto the image of the first motion effect video, can be determined by the first device itself, or it can be pre-configured or predefined by the protocol; this embodiment does not impose any limitations.
[0140] For example, if the second device is used to display the video of the second device, and the first device supports motion effects for video users, then the first device can overlay the video of the first device onto the video of the first motion effect (i.e., local device B + peer device A (motion effect)) and encode it to obtain a video stream. The first device can also overlay the video of the first device onto the video of the first motion effect (i.e., local device B + peer device A (motion effect)) and display it in a large window of the first device. Furthermore, the size at which the video of the first device is overlaid onto the video of the first motion effect can be determined by the first device itself, or it can be pre-configured or predefined by the protocol; this embodiment of the application does not impose any limitations.
[0141] It is understood that during a video call, the first device may execute the method of this application embodiment to add motion effects to the video of the first device / second device when it determines that the video of the second device contains motion effect actions that trigger motion effects. Alternatively, the method of this application embodiment may be executed by default during the video call. The same applies below, and will not be repeated.
[0142] Optionally, the first device may also send information to the second device to indicate the overlay area, and the second device may receive the information from the first device to indicate the overlay area.
[0143] The overlay area can be the region occupied by the first motion effect video frame within the video frame of the first device. The size of the overlay area varies depending on the size of the first motion effect video frame. Alternatively, the overlay area can also be the region occupied by the first device's video frame within the first motion effect video frame; the size of the overlay area also varies depending on the size of the first device's video frame. Information indicating the overlay area (or overlay area information, or other naming conventions) can be carried in Real-Time Transport Protocol (RTP) packets, or in any possible type of packet, without restriction. For example, in an RTP packet, a one-byte header or any other possible naming convention can be added to the RTP packet's header extension field as overlay area information.
[0144] For example, such as Figure 7As shown, setting X=1 in the RTP packet indicates that there are some additional RTP extension headers or extension fields after CSRC. Specifically, the first 16 bits after the RTP header are fixed as 0xBE0xDE, meaning that the extension header is a one-byte extension. Length=3 indicates that there are 4 extension headers. Each extension header starts with a byte. The first 4 bits are the identifier (ID) of this extension header, and the last 4 bits are the length of the data - 1. For example, L-1=0 means that there is 1 byte of data. Setting L=3 means that there are 4 one-byte data. These 4 data are the coordinate parameters of the superimposed area in the video screen of the first device, such as right / bottom / width / height. For example, if the screen resolution is 1920*1080, these 4 parameters can be 0 / 0 / 640 / 360 respectively. Of course, if some of these four parameters are known to the second device, such as through pre-configuration / protocol pre-definition, the overlay area information can also include only the other part of the parameters that are unknown to the second device, in order to reduce overhead and avoid redundancy.
[0145] The first device can send overlay area information to the second device when the size of the overlay area changes, or it can send it periodically. During this process, the overlay area information can be carried along with the video stream, such as being included in the aforementioned video stream, or it can be transmitted separately to the second device, decoupled from the transmission of the video stream. The second device can save the received overlay area information and reuse the saved overlay area information when no new overlay area information is received. For example, the second device can determine the location of the overlay area in the large window based on the information used to indicate the overlay area, so that the user can operate in the overlay area (such as clicking or swiping). The second device can trigger a window switching operation based on its location, determining to switch the video of the first device / second device to be displayed in the large window. For example, if the overlay area is the area occupied by the first motion effect video frame within the video frame of the first device, then the window switching operation can indicate a switch from the second window being contained within the first window to the first window being contained within the second window, or a switch from a video where the first window contains the second window to a video where the second window contains the first window. Similarly, if the overlay area is the area occupied by the first device's video frame within the first motion effect video frame, then the window switching operation can indicate a switch from the first window being contained within the second window to the second window being contained within the first window, or a switch from a video where the second window contains the first window to a video where the first window contains the second window. In response to a window switching operation triggered by the user within the overlay area, the second device can send instruction information to the first device.
[0146] It should be understood that in the case of decoupled transmission, the overlay area information can be transmitted to the second device before the video stream. The second device can also determine whether it needs to execute the method of this application embodiment based on the received overlay area information, as shown in S503 below; otherwise, it executes the existing process. In addition, the overlay area information is optional information. If the second device knows the location of the overlay area in advance through pre-configuration or protocol pre-definition, the first device and the second device may not need to exchange this information.
[0147] The second possible design is that the video stream can be multiple streams.
[0148] For example, a video stream can contain a first video stream and a second video stream.
[0149] The first video stream is a video stream from a first device. The first device can encode its own video to obtain the first video stream, and then send the first video stream to a second device through a first transmission channel. The second device can receive the first video stream from the first device through the first transmission channel. The first transmission channel can be a transmission channel established between the first device and the second device for a video call, such as one negotiated during the video call setup process, used by the first device to transmit its own video stream, such as the first video stream, to the second device. The first transmission channel can be any existing type of transmission channel; the embodiments of this application do not limit the negotiation and establishment methods.
[0150] The second video stream is a stream of the first video with motion effects. The first device can add motion effects and encoding to the video from the second device to obtain the second video stream. The first device can send the second video stream to the second device through a second transmission channel. The second device can receive the second video stream from the first device through the second transmission channel. The second transmission channel can be an additionally negotiated transmission channel between the first and second devices, such as being negotiated during the establishment of a video call, or being negotiated when the motion effect is triggered; the specific timing is not limited. The function of the second transmission channel can differ from that of the first transmission channel; for example, the second transmission channel can be used by the first device to transmit the video stream from the second device to the second device, such as the second video stream. The first transmission channel can also be any existing transmission channel of any possible type, and the negotiation and establishment methods of this application embodiment are not limited. Thus, the video with added motion effects and the video without added motion effects are transmitted separately.
[0151] It is understandable that in the case of split transmission, the second device can decide on its own how to display, overlay, and switch display windows of different video streams. Therefore, unlike the first possible design scheme mentioned above, the second device does not need to indicate to the first device which device's video is currently displayed in the large window, i.e., it does not need to send instruction information. The first device also does not need to overlay the video or send overlay area information to the second device. The same applies below, and will not be elaborated here.
[0152] Optionally, the second device may also establish a second transmission channel according to the negotiation, or establish multiple transmission channels for a video call, and determine whether it needs to execute the method of the embodiment of this application, as shown in S503 below; otherwise, it executes the existing process.
[0153] It should also be understood that when the motion effects are applied to video users, the first device can determine / adjust the overlay method of the video frame according to the display situation of the second device. This allows the motion effects to switch seamlessly with the user when the display window is switched, such as when the second window is contained in the first window and the first window is contained in the second window, that is, when the video of the second device switches from a small window display to a large window display. This results in a better user experience.
[0154] Scenario 2:
[0155] Regarding the animation effects for the video display window, such as the animation effect for a large window, the switching of the display window of the second device will cause the videos of different users to be displayed in large windows respectively. Therefore, the video stream can be a video stream containing animation effects determined based on the video of the first device and the first animation effect video, or a video stream containing animation effects determined based on the second animation effect video and the video of the second device. It changes with the switching of the display window, so that the animation effect follows the display window.
[0156] The third possible design is that the video stream can be a single stream.
[0157] For example, when the second device displays the video from the first device, the single-stream video stream can be obtained by overlaying the video from the second device onto the video from the second motion effect, and the screen size of the video from the second device is smaller than the screen size of the video from the second motion effect. Alternatively, when the second device displays the video from the second device, the single-stream video stream can be obtained by overlaying the video from the first device onto the video from the first motion effect, and the screen size of the video from the first device is smaller than the screen size of the video from the first motion effect. In other words, if the motion effect is for a large window / screen, then the first device adds the motion effect to the video from the device that the second device displays in a large window / screen, so that the motion effect follows the large window / screen. The same principle applies to small windows / screens, and will not be elaborated further.
[0158] It should be understood that the second device can switch between displaying the video from the first device and the second device. For example, when the second device establishes a video call with the first device, it can display the video from the first device, and then switch to display the video from the second device by changing the size of the window. The second device can inform the first device which device's video is currently displayed in the large window when establishing a video call with the first device, or when changing the size of its own display window, sending an instruction message. For specific implementation details, please refer to the relevant description in the "First Possible Design Scheme" above, which will not be repeated here. In this way, the first device can also determine the display status of the second device based on the instruction message, thereby determining / adjusting the video overlay method and avoiding errors in animation display.
[0159] For example, if the second device is used to display the video from the first device, and the first device supports animation effects for the video display window, then the first device can overlay the video from the second device onto the video of the second animation effect (i.e., local B (animation effect) + peer A) and encode it to obtain a video stream. The first device can also overlay the video from the first device onto the video of the first animation effect (i.e., local B + peer A (animation effect)) and display it in a large window on the first device. The size at which the video of the second animation effect is overlaid onto the video of the second device, or the size at which the video of the first device is overlaid onto the video of the first animation effect, can be determined by the first device itself, or it can be pre-configured or predefined by the protocol; this embodiment does not impose any limitations.
[0160] For example, if the second device is used to display the video of the second device, and the first device supports animation effects for the video display window, then the first device can overlay the video frame of the first device onto the frame of the first animation effect video (i.e., local B + peer A (animation effect)) and encode it to obtain a video stream. Similarly, the first device can also overlay the video frame of the first device onto the frame of the first animation effect video (i.e., local B + peer A (animation effect)) and display it in the large window of the first device. Furthermore, the size at which the video frame of the first device is overlaid onto the frame of the first animation effect video can be determined by the first device itself, or it can be pre-configured or predefined by the protocol; this application embodiment does not impose any restrictions.
[0161] Optionally, the first device can also send overlay area information to the second device. The second device can receive the overlay area information from the first device to respond to a user's window switching operation triggered within the overlay area. The second device then sends instruction information to the first device. If the video frame of the second device occupies the area within the frame of the second motion effect video, the window switching operation instruction changes from "the second window is contained within the first window" to "the first window is contained within the second window." If the overlay area is the area occupied by the video frame of the first device within the frame of the first motion effect video, the window switching operation instruction changes from "the first window is contained within the second window" to "the second window is contained within the first window." Specific implementation details can be found in the above description and will not be repeated here.
[0162] It should also be understood that the overlay area information is optional. If the second device knows the location of the overlay area in advance through pre-configuration or protocol pre-definition, the first device and the second device may not need to exchange this information.
[0163] The fourth possible design is that the video stream can be multiple streams.
[0164] For example, when the second device displays the video of the second device, the video stream may include a first video stream and a second video stream. The first video stream is obtained by encoding the video of the first device, and the second video stream is obtained by encoding the first motion effect video. Alternatively, when the second device displays the video of the first device, the video stream may include a third video stream and a fourth video stream. The third video stream is obtained by encoding the video of the second device, and the fourth video stream is obtained by encoding the second motion effect video. That is, if the motion effect is for a large window / screen, then the first device adds the motion effect to the video of the device that the second device displays in a large window / screen and transmits it through the corresponding stream. Similarly, if the motion effect is for a small window / screen, then the first device adds the motion effect to the video of the device that the second device displays in a small window / screen and transmits it through the corresponding stream, so that the motion effect follows the large window / screen.
[0165] It should be understood that the second device can switch between displaying the video from the first device and the video from the second device. For example, when the second device establishes a video call with the first device, it can display the video from the first device, and then switch to display the video from the second device by changing the size of the window. The second device can inform the first device which device's video is currently displayed in its large window when establishing a video call with the first device or when changing the size of its own display window, such as by sending an instruction. For example, in response to a user-triggered window switching operation, the second device sends an instruction to the first device; this instruction indicates a change from the second window being contained within the first window to the first window being contained within the second window, or vice versa. Therefore, the first device can also determine the display status of the second device based on the instruction, thereby determining / adjusting the video for which animation effects need to be added, and avoiding errors in animation effect display.
[0166] For example, if the second device is used to display video from the second device, and the first device supports animation effects for the video display window, then the first device can encode the video from the first device to obtain a second animation effect video, and send the first video stream to the second device through the first transmission channel. The first device can also add animation effects to the video from the second device and encode it to obtain a second animation effect video, and send the second video stream to the second device through the second transmission channel. Correspondingly, the second device can receive the first video stream from the first device through the first transmission channel, and can also receive the second video stream from the first device through the second transmission channel.
[0167] For example, if the second device is used to display the video from the first device, and the first device supports animation effects for the video display window, then the first device can add animation effects to the video from the first device and encode it to obtain a fourth video stream, and send the fourth video stream to the second device through the first transmission channel. The first device can also encode the video from the second device to obtain a third video stream, and send the third video stream to the second device through the second transmission channel. Correspondingly, the second device can receive the fourth video stream from the first device through the first transmission channel, and can also receive the third video stream from the first device through the second transmission channel.
[0168] It should be understood that the first and second transmission channels can be referred to in the relevant introduction above, and will not be repeated here.
[0169] As can be seen, if the containment relationship between the first window and the second window changes, by adjusting the stream containing the animation video to be transmitted in different transmission channels, such as changing the transmission of the second video stream (containing the animation) through the second transmission channel to the transmission of the fourth video stream (containing the animation) through the first transmission channel, the animation can also remain unchanged with the user, such as still being displayed in the current large / small window, thus achieving the effect of the animation following the window.
[0170] S503, the second device displays the video in the video stream.
[0171] Continuing with the first possible design scheme mentioned above:
[0172] The second device can, by default, display the video stream in a large window / large screen. For example, if the second window is contained within the first window, or if the first window contains the video from the second window (the same applies below), the first window can be considered a large window. In this case, the second device displays the video from the video stream in the first window, such as decoding the video stream and displaying it in the first window. Alternatively, the second device can hide the second window to prevent the video displayed in the second window from affecting the animation experience. Similarly, if the first window is contained within the second window, or if the second window contains the video from the first window (the same applies below), the second window can be considered a large window. In this case, the second device displays the video from the video stream in the second window. Alternatively, the second device can hide the first window to prevent the video displayed in the first window from affecting the animation experience.
[0173] Continuing with the second possible design scheme mentioned above:
[0174] In one possible implementation, the second device can employ a principle similar to that of the first device in the first possible design, superimposing the video from the first video stream onto the video from the second video stream. For example, given the first video stream received through the first transmission channel and the second video stream received through the second transmission channel, the second device can decode the first and second video streams to obtain the video from the first device and the first motion effect video, respectively. If the second device is used to display the video from the first device in a large window, it can superimpose the image of the first motion effect video onto the image of the first device's video, display it in the large window, and hide the small window. For details, please refer to the relevant descriptions above, which will not be repeated here. Similarly, if the second device is used to display the video from the second device in a large window, it can superimpose the image of the first device's video onto the image of the first motion effect video, display it in the large window, and hide the small window. For details, please refer to the relevant descriptions above, which will not be repeated here.
[0175] It should be understood that the size at which the image of the first motion effect video is superimposed on the image of the video of the first device, or the size at which the image of the video of the first device is superimposed on the image of the first motion effect video, can be determined by the second device itself, or can be pre-configured or predefined by the protocol. This application embodiment does not impose any restrictions.
[0176] In another possible implementation, the second device can configure the correspondence between the transmission channel and the display window when establishing the second transmission channel, such as the first transmission channel corresponding to the first window and the second transmission channel corresponding to the second window. Thus, for a first video stream received through the first transmission channel, the second device can display the video in the first video stream in the first window, such as decoding the first video stream and displaying it in the first window. For a second video stream received through the second transmission channel, the second device can display the video in the second video stream in the second window, such as decoding the second video stream and displaying it in the second window. The second device can independently determine / adjust the containment relationship between the first and second windows, such as the first window being contained within the second window (i.e., the larger window displays the video with animation effects generated by the second device itself, and the smaller window displays the video with animation effects generated by the first device), or the second window being contained within the first window (i.e., the larger window displays the video with animation effects generated by the second device itself, and the smaller window displays the video with animation effects generated by the second device itself).
[0177] It should be understood that the second device can also be configured such that the first transmission channel corresponds to the second window and the second transmission channel corresponds to the first window. The specific implementation principle is similar to that described above, and will not be repeated here.
[0178] Continuing with the third possible design scheme mentioned above:
[0179] The second device can also display the video stream in a large window / screen by default. For example, if the first window is a large window, the video stream will be displayed in the first window; if the second window is a large window, the video stream will be displayed in the second window. The specific implementation is similar to the first possible design scheme mentioned above, which can be referred to for understanding, and will not be elaborated here.
[0180] Continuing with the fourth possible design scheme mentioned above:
[0181] In one possible implementation, the second device can employ a principle similar to that of the first device in the first possible design, superimposing the video from the first video stream with the video from the second video stream, or superimposing the video from the third video stream with the video from the fourth video stream. For example, when the second device is used to display the video of the second device in a large window, the second device can receive the first video stream through the first transmission channel and the second video stream through the second transmission channel. The second device can decode the first and second video streams to obtain the video of the first device and the first motion effect video, respectively. The second device can superimpose the image of the video of the first device onto the image of the first motion effect video and display it in the large window, and can also hide the small window. Specific implementation details can be found in the above description and will not be repeated here. As another example, when the second device is used to display the video of the first device in a large window, the second device can receive the fourth video stream through the first transmission channel and the third video stream through the second transmission channel. The second device can decode the third and fourth video streams to obtain the video of the second device and the second motion effect video, respectively. The second device can overlay the video from the second device onto the video from the second motion effect and display it in a large window, and it can also hide the small window. For details on the implementation, please refer to the above-mentioned introduction, which will not be repeated here.
[0182] In another possible implementation, the second device can configure the correspondence between the transmission channels and display windows when establishing the second transmission channel, such as the first transmission channel corresponding to the first window and the second transmission channel corresponding to the second window. If the first window is contained within the second window, the second device can receive the first video stream through the first transmission channel and the second video stream through the second transmission channel. Therefore, the second device can display the video from the first video stream in the first window (e.g., decode the first video stream and display it in the first window) and display the video from the second video stream in the second window (e.g., decode the second video stream and display it in the second window). Alternatively, if the second window is contained within the first window, the second device can receive the fourth video stream through the first transmission channel and the third video stream through the second transmission channel. Therefore, the second device can display the video from the fourth video stream in the first window (e.g., decode the fourth video stream and display it in the first window) and display the video from the third video stream in the second window (e.g., decode the third video stream and display it in the second window).
[0183] As can be seen, if the containment relationship between the first window and the second window changes or how they switch, the animation video can be transmitted through different transmission channels. For example, if the second video stream (containing the animation) is transmitted through the second transmission channel, it can be changed to the fourth video stream (containing the animation) being transmitted through the first transmission channel. The animation can remain unchanged with the user and will still be displayed in the current large window, thus achieving the effect of the animation following the window.
[0184] Furthermore, the above example uses animation effects targeting video users or video display windows, but there are no limitations. For example, in the case of a single stream, the animation effect can also be full-screen and not targeting video users or video display windows.
[0185] In summary, if the second device (i.e., the peer) does not support animation effects, the first device (this device) can add animation effects to the second device's video, such as a first animation effect video, and send it to the second device along with the first device's video stream; alternatively, the first device can also add animation effects to its own video, such as a second animation effect video, and send it to the second device along with the second device's video stream. The second device does not need to perform any animation effect-related processing; it can display the corresponding animation effects simply by displaying the video stream, thus enabling the peer to display animation effects even when the peer does not support them.
[0186] The above combination Figure 5 The overall flow of the communication method provided in the embodiments of this application is described in detail. The following, in conjunction with... Figure 8-11 This paper describes the specific process of the communication method provided in the embodiments of this application in various scenarios.
[0187] Scene 1:
[0188] Figure 8 Flowchart of the communication method provided in the embodiments of this application Figure 2 This communication method is applicable to the aforementioned communication system and mainly involves the interaction between end B (such as the first device) and end A (such as the second device).
[0189] For example, if client B supports motion effects while client A does not, and the motion effects are specific to video users, client B can add motion effects to the video received from client A, resulting in motion effect video #1. Client B can then overlay the footage of motion effect video #1 onto its own video, or vice versa, and then encode it to obtain a video stream. Client B can send this video stream to client A, which in turn receives and decodes it to display the corresponding motion effect video.
[0190] Specifically, such as Figure 8 As shown, the flow of this communication method is as follows:
[0191] S801, B sends a session invitation message to A.
[0192] The session invitation message is used by the B-end to request to establish a video call with the A-end. The session invitation message can contain the B-end's capability information, such as the B-end supporting animation effects, specifically supporting animation effects for video users, or supporting animation effects for the video display window.
[0193] S802, terminal A sends a response message to terminal B.
[0194] The response message is used by client A to respond to client B's request to establish a video call. If the request is granted, then client B and client A can establish a video call. The response message can also contain information about client A's capabilities, such as client A not supporting animations.
[0195] In addition, S801-S802 can also refer to the relevant introduction in S502 above, and will not be repeated here.
[0196] S803, A sends video from A to B.
[0197] During a video call between B and A, A can send its own video (such as the video from the second device mentioned above) to B. The video from A contains video frames with motion effects. For details, please refer to the relevant descriptions in S501-S502 above, which will not be repeated here.
[0198] S804, Terminal A sends instruction information #1 to Terminal B.
[0199] Instruction message #1 indicates that the large window on end A is used to display B, that is, the video of A is displayed in a small window.
[0200] It is understood that the execution order between S803 and S804 is not limited in the embodiments of this application.
[0201] S805: The B-end recognizes that the video from the A-end contains motion effects, but the A-end does not support motion effects.
[0202] S806, the B-end displays the video overlaid in a large window using the method of A (animation) + B.
[0203] The B-side supports animation effects for video users, so the animation effects are added for the A-side or for A-side users, and can be represented by A (animation effect), as in the first animation effect video mentioned above. In "A (animation effect) + B", "+" indicates overlay, the part before "+" represents the content displayed in the large window, the part after "+" represents the content displayed in the small window, and B represents the video on the B-side (such as the video of the first device mentioned above).
[0204] S807, the B end overlays the B+A (motion effect) and encodes it to obtain video stream #1.
[0205] S808, B sends video stream #1 and overlay area information #1 to A.
[0206] Overlay area information #1 is used to indicate the area occupied by the image of A (motion effect) in the video image of B.
[0207] S809, the A end displays the decoded video stream #1 in a large window.
[0208] Additionally, the A-side can hide the small window.
[0209] For S804-S809, please refer to the relevant introduction of the "first possible design scheme" above, and it will not be repeated here.
[0210] S810, in response to the switching of window size, A sends instruction information #2 to B.
[0211] The window size switching can be achieved by switching the display of video B from a large window to a large window for video A. This can happen if the user triggers the overlay area indicated by overlay information #1, thus triggering the window size switching. Correspondingly, indication information #2 indicates that the large window for video A is used to display video A; that is, video B is displayed in a small window.
[0212] S811, B end is superimposed in the form of A (motion effect) + B, and encoded to obtain video stream #2.
[0213] S812, B sends video stream #2 and overlay area information #2 to A.
[0214] Overlay area information #2 is used to indicate the area occupied by the video image from end B in the image of A (motion effect).
[0215] S813, A-end displays the decoded video stream #2 in a large window.
[0216] Additionally, the A-side can continue to hide the small window.
[0217] For S810-S813, please refer to the relevant introduction of the "first possible design scheme" above, and it will not be repeated here.
[0218] Scene 2:
[0219] Figure 9 Flowchart of the communication method provided in the embodiments of this application Figure 3 This communication method is applicable to the aforementioned communication system and mainly involves the interaction between end B (such as the first device) and end A (such as the second device).
[0220] For example, if terminal B supports motion effects while terminal A does not, and the motion effects are specific to video users, terminal B can add motion effects to the video received from terminal A, resulting in motion effect video #1. Terminal B can encode its own video to obtain video stream #1, which is then transmitted to terminal A via transmission channel #1. Terminal B can also encode motion effect video #1 to obtain video stream #2, which is also transmitted to terminal A via transmission channel #2. Correspondingly, terminal A can receive video stream #1 and video stream #2 through different transmission channels, and terminal A can overlay video stream #1 and video stream #2 to display the corresponding motion effect video.
[0221] Specifically, such asFigure 9 As shown, the flow of this communication method is as follows:
[0222] S901, B sends a session invitation message to A.
[0223] The session invitation message is used by the B-end to request to establish a video call with the A-end. The session invitation message can contain the B-end's capability information, such as the B-end supporting animation effects, specifically supporting animation effects for video users, or supporting animation effects for the video display window.
[0224] S902, Terminal A sends a response message to Terminal B.
[0225] The response message is used by end A to respond to end B's request to establish a video call. If the video call is allowed, end B and end A can then negotiate to establish a video call, such as establishing transmission channel #1 (as mentioned above, the first transmission channel). The response message can also contain end A's capability information, such as end A not supporting animations.
[0226] In addition, S901-S902 can also refer to the relevant introduction in S502 above, and will not be repeated here.
[0227] S903, A sends video from A to B.
[0228] During a video call between B and A, A can send its own video (such as the video from the first device mentioned above) to B. The video from A contains video frames with motion effects. For details, please refer to the relevant descriptions in S501-S502 above, which will not be repeated here.
[0229] S904, the B-end recognizes that the video from the A-end contains motion effects, but the A-end does not support motion effects.
[0230] S905, B and A negotiate to establish transmission channel #2 (as described above as the second transmission channel).
[0231] It is understandable that S905 can be triggered by S904, or it can be established by default during the process of B and A negotiating to establish a video call. There are no restrictions on the specific implementation.
[0232] S906, the B-end displays the video overlaid in a large window using the method of A (animation) + B.
[0233] For S906, the relevant introduction of S806 mentioned above can be referred to, and will not be repeated here.
[0234] S907, the video encoded by B-end is obtained as video stream #1, and the video encoded by A (motion effects) is obtained as video stream #2.
[0235] Optionally, the video from the B end can be displayed in a large window. For overhead considerations, the bitrate of encoding the video from the B end can be higher than the bitrate of encoding the A (motion effects).
[0236] S908, B sends video stream #1 to A through transmission channel #1.
[0237] S909, B sends video stream #2 to A through transmission channel #2.
[0238] The S910 decodes video stream #1 and video stream #2 on the A end and displays the video overlaid in a large window using a B+A (motion effect) method.
[0239] Additionally, the A-side can hide the small window.
[0240] S911, in response to the switching of large and small windows, decodes video stream #1 and video stream #2 on end A, and displays the video overlaid in the form of A (animation) + B in a large window.
[0241] The window size switching function can be used to switch the display of video B from a large window to a large window display of video A, which can be triggered by the user. Additionally, the small window can also be hidden on the A side.
[0242] For S903-S911, please refer to the relevant introduction of the "second possible design scheme" mentioned above, which will not be repeated here.
[0243] Scene 3:
[0244] Figure 10 Flowchart of the communication method provided in the embodiments of this application Figure 4 This communication method is applicable to the aforementioned communication system and mainly involves the interaction between end B (such as the first device) and end A (such as the second device).
[0245] For example, if client B supports animation effects while client A does not, and regarding animation effects for the video display window, client B can add animation effects to the video received from client A, resulting in animation video #1. Client B can then overlay the footage of animation video #1 onto its own video and encode it to obtain a video stream. Alternatively, client B can add animation effects to its own video, resulting in animation video #2. Client B can then overlay the footage of animation video #2 onto its own video and encode it to obtain a video stream. Client B can send this video stream to client A, and client A will receive and decode it to display the corresponding animation video.
[0246] Specifically, such as Figure 10 As shown, the flow of this communication method is as follows:
[0247] S1001, Client B sends a session invitation message to Client A.
[0248] The session invitation message is used by the B-end to request to establish a video call with the A-end. The session invitation message can contain the B-end's capability information, such as the B-end supporting animation effects, specifically supporting animation effects for video users, or supporting animation effects for the video display window.
[0249] S1002, Terminal A sends a response message to Terminal B.
[0250] The response message is used by client A to respond to client B's request to establish a video call. If the request is granted, then client B and client A can establish a video call. The response message can also contain information about client A's capabilities, such as client A not supporting animations.
[0251] In addition, S1001-S1002 can also refer to the relevant introduction in S502 above, and will not be repeated here.
[0252] S1003, Terminal A sends the video from Terminal A to Terminal B.
[0253] During a video call between B and A, A can send its own video (such as the video from the first device mentioned above) to B. The video from A contains video frames with motion effects. For details, please refer to the relevant descriptions in S501-S502 above, which will not be repeated here.
[0254] S1004, Terminal A sends instruction information #1 to Terminal B.
[0255] Instruction message #1 indicates that the large window on end A is used to display B, meaning that the video from A is displayed in a small window.
[0256] It is understood that the execution order between S1003 and S1004 is not limited in the embodiments of this application.
[0257] S1005, the B-end recognizes that the video from the A-end contains motion effects, but the A-end does not support motion effects.
[0258] S1006, the B end displays the video overlaid in a large window using the method of A (animation) + B.
[0259] A (motion effect) is motion effect video #1 (such as the first motion effect video mentioned above). S1006 can also refer to the relevant introduction of S806 mentioned above, and will not be repeated here.
[0260] S1007, the B end is superimposed in the form of B (motion effect) + A, and encoded to obtain video stream #1.
[0261] B (motion effect) is the second motion effect video (such as the second motion effect video mentioned above).
[0262] S1008, B sends video stream #1 and overlay area information #1 to A.
[0263] Overlay area information #1 is used to indicate the area occupied by the B (motion effect) image in the video image of A.
[0264] S1009, Terminal A displays the decoded video stream #1 in a large window.
[0265] Additionally, the A-side can hide the small window.
[0266] For S1004-S1009, please refer to the relevant introduction of the "third possible design scheme" above, and it will not be repeated here.
[0267] S1010, in response to the window size switching, terminal A sends instruction information #2 to terminal B.
[0268] The window size switching can be achieved by switching the display of video B from a large window to a large window for video A. This can happen if the user triggers the overlay area indicated by overlay information #1, thus triggering the window size switching. Correspondingly, indication information #2 indicates that the large window for video A is used to display video A; that is, video B is displayed in a small window.
[0269] S1011, B end is superimposed in the form of A (motion effect) + B, and encoded to obtain video stream #2.
[0270] S1012, B sends video stream #2 and overlay area information #2 to A.
[0271] Overlay area information #2 is used to indicate the area occupied by B's video frame within A's (motion effect) frame.
[0272] S1013, Terminal A displays the decoded video stream #2 in a large window.
[0273] Additionally, the A-side can continue to hide the small window.
[0274] For S1010-S1013, please refer to the relevant introduction of the "third possible design scheme" above, and it will not be repeated here.
[0275] Scene 4:
[0276] Figure 11 Flowchart of the communication method provided in the embodiments of this application Figure 5 This communication method is applicable to the aforementioned communication system and mainly involves the interaction between end B (such as the first device) and end A (such as the second device).
[0277] For example, if terminal B supports animation effects while terminal A does not, and regarding the animation effects for the video display window, terminal B can add animation effects to the video received from terminal A, resulting in animation video #1. Terminal B can then encode its own video to obtain video stream #1, which is transmitted to terminal A via transmission channel #1. Terminal B can also encode animation video #1 to obtain video stream #2, which is then transmitted to terminal A via transmission channel #2. Alternatively, terminal B can add animation effects to its own video, resulting in animation video #2. Terminal B can then encode animation video #2 to obtain video stream #4, which is transmitted to terminal A via transmission channel #1. Terminal B can also encode terminal A's video to obtain video stream #3, which is transmitted to terminal A via transmission channel #2. Correspondingly, terminal A receives and decodes the video stream to display the corresponding animation video.
[0278] Specifically, such as Figure 11 As shown, the flow of this communication method is as follows:
[0279] S1101, B sends a session invitation message to A.
[0280] The session invitation message is used by the B-end to request to establish a video call with the A-end. The session invitation message can contain the B-end's capability information, such as the B-end supporting animation effects, specifically supporting animation effects for video users, or supporting animation effects for the video display window.
[0281] S1102, Terminal A sends a response message to Terminal B.
[0282] The response message is used by client A to respond to client B's request to establish a video call. If the request is granted, then client B and client A can establish a video call. The response message can also contain information about client A's capabilities, such as client A not supporting animations.
[0283] In addition, S1101-S1102 can also refer to the relevant introduction in S502 above, and will not be repeated here.
[0284] S1103, Terminal A sends the video from Terminal A to Terminal B.
[0285] During a video call between B and A, A can send its own video (such as the video from the first device mentioned above) to B. The video from A contains video frames with motion effects. For details, please refer to the relevant descriptions in S501-S502 above, which will not be repeated here.
[0286] S1104, End A sends instruction information #1 to End B.
[0287] Instruction message #1 indicates that the large window on end A is used to display B, meaning that the video from A is displayed in a small window.
[0288] It is understood that the execution order between S1103 and S1104 is not limited in the embodiments of this application.
[0289] S1105, the B-end recognizes that the video from the A-end contains motion effects, but the A-end does not support motion effects.
[0290] S1106, End B and End A negotiate to establish transmission channel #2 (as described above as the second transmission channel).
[0291] S1107, the B-end displays a video overlaid in a large window using the method of A (animation) + B.
[0292] S1106 can also refer to the relevant introduction of S806 above, and will not be repeated here.
[0293] S1108, the video encoded by B (motion effect) at end B is obtained as video stream #4, and the video encoded by end A is obtained as video stream #3.
[0294] B (motion effect) represents motion effect video #2 (such as the second motion effect video mentioned above). Optionally, the B (motion effect) video can be displayed in a large window. For overhead considerations, the bitrate of the B (motion effect) video can be higher than that of the A video.
[0295] S1109, End B sends video stream #4 to End A through transmission channel #1 (as described above as the first transmission channel).
[0296] S1110, End B sends video stream #3 to End A through transmission channel #2.
[0297] S1111, A decodes video stream #3 and video stream #4, and displays the video overlaid in a large window using the B (animation) + A method.
[0298] Additionally, the A-side can hide the small window.
[0299] For S1104-S1111, please refer to the relevant introduction of the "fourth possible design scheme" above, and it will not be repeated here.
[0300] S1112, in response to the window size switching, terminal A sends instruction information #2 to terminal B.
[0301] The window size switching can be achieved by switching the display of video B from a large window to a large window for video A. This can happen if the user triggers the overlay area indicated by overlay information #1, thus triggering the window size switching. Correspondingly, indication information #2 indicates that the large window for video A is used to display video A; that is, video B is displayed in a small window.
[0302] S1113, the video encoded by B-end is obtained as video stream #1, and the video encoded by A (motion effect) is obtained as video stream #2.
[0303] A (motion effect) represents motion effect video #1 (such as the first motion effect video mentioned above). Optionally, the video of A (motion effect) can be displayed in a large window. For overhead considerations, the bitrate of the video of A (motion effect) can be higher than the bitrate of the video of B.
[0304] S1114, End B sends video stream #1 to End A through transmission channel #1.
[0305] S1115, End B sends video stream #2 to End A through transmission channel #2.
[0306] S1116, A decodes video stream #1 and video stream #2, and displays the video overlaid in a large window using the A (motion effect) + B method.
[0307] Additionally, the A-side can hide the small window.
[0308] S1112-S1116 can also refer to the relevant introduction of the "fourth possible design scheme" above, and will not be repeated here.
[0309] The above combination Figure 5-11 The communication method provided in the embodiments of this application is described in detail below. Figure 12-13 This document describes in detail the communication apparatus used to perform the communication method provided in the embodiments of this application.
[0310] Figure 12 This is a schematic diagram of the structure of the communication device provided in the embodiments of this application. Figure 1 For example, such as Figure 12 As shown, the communication device 1200 includes a transceiver module 1201 and a processing module 1202. For ease of explanation, Figure 12 Only the main components of the communication device are shown.
[0311] The communication device 1200 can be applied to the above-described communication method to achieve the corresponding functions. For example, the transceiver module 1201 can be used to implement the transceiver function in the above-described communication method, and the processing module 1202 can be used to implement other functions in the above-described communication method besides the transceiver function.
[0312] Optionally, the transceiver module 1201 may include a transmitting module. Figure 12 (not shown in the image) and receiving module ( Figure 12 (Not shown in the diagram). The transmitting module is used to implement the transmitting function of the communication device 1200, and the receiving module is used to implement the receiving function of the communication device 1200.
[0313] Optionally, the communication device 1200 may also include a storage module. Figure 12(Not shown in the image), the storage module stores programs or instructions. When the processing module 1202 executes the program or instructions, the communication device 1200 can perform the aforementioned... Figure 5-11 The functions in the method shown.
[0314] It is understood that the communication device 1200 may be a network device, or a chip (system) or other component or assembly that can be set in the network device, or a device that includes the network device. This application does not limit this.
[0315] Furthermore, the technical effects of the communication device 1200 can be referenced from the technical effects of the communication method described above, and will not be repeated here.
[0316] Figure 13 Schematic diagram of the communication device provided in the embodiments of this application Figure 2 For example, the communication device can be a terminal, or a chip (system) or other component or assembly that can be set in the terminal. Figure 13 As shown, the communication device 1300 may include a processor 1301. Optionally, the communication device 1300 may also include a memory 1302 and / or a transceiver 1303. The processor 1301 is coupled to the memory 1302 and the transceiver 1303, for example, they may be connected via a communication bus.
[0317] The following is combined Figure 13 A detailed description of each component of the communication device 1300 is provided below:
[0318] The processor 1301 is the control center of the communication device 1300. It can be a single processor or a collective term for multiple processing elements. For example, the processor 1301 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs) or one or more field-programmable gate arrays (FPGAs).
[0319] Optionally, the processor 1301 can perform various functions of the communication device 1300 by running or executing software programs stored in the memory 1302 and calling data stored in the memory 1302, such as performing the above-mentioned functions. Figure 5-11 The communication method shown.
[0320] In a specific implementation, as one example, the processor 1301 may include one or more CPUs, for example... Figure 13 CPU0 and CPU1 are shown in the diagram.
[0321] In a specific implementation, as one example, the communication device 1300 may also include multiple processors, for example... Figure 13 The processors 1301 and 1304 are shown. Each of these processors can be a single-core processor or a multi-core processor. Here, "processor" can refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0322] The memory 1302 is used to store the software program that executes the solution of this application, and is controlled by the processor 1301 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0323] Optionally, the memory 1302 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 1302 may be integrated with the processor 1301 or may exist independently, and may be connected via the interface circuit of the communication device 1300. Figure 13 (Not shown in the image) is coupled to processor 1301, and this embodiment of the application does not specifically limit this.
[0324] Transceiver 1303 is used for communication with other communication devices. For example, if communication device 1300 is a terminal, transceiver 1303 can be used to communicate with a network device or with another terminal device. As another example, if communication device 1300 is a network device, transceiver 1303 can be used to communicate with a terminal or with another network device.
[0325] Optionally, transceiver 1303 may include a receiver and a transmitter. Figure 13 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0326] Optionally, the transceiver 1303 can be integrated with the processor 1301, or it can exist independently and be connected via the interface circuit of the communication device 1300. Figure 13 (Not shown in the image) is coupled to processor 1301, and this embodiment of the application does not specifically limit this.
[0327] Understandable, Figure 13 The structure of the communication device 1300 shown does not constitute a limitation on the communication device. Actual communication devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0328] Furthermore, the technical effects of the communication device 1300 can be referred to the technical effects of the method described in the above method embodiments, and will not be repeated here.
[0329] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0330] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0331] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0332] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0333] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0334] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0335] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0336] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0337] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0338] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0339] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0340] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0341] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A communication method characterized by comprising: The method is applied to a first device, and the method comprises: The first device receives a video of a second device; In a case where the first device supports animation effects and the second device does not support animation effects, the first device sends a video stream to the second device; wherein the video stream is a video stream containing animation effects determined according to the video of the first device and a first animation effect video, the first animation effect video being a video in which the video of the second device is added with animation effects, or the video stream is a video stream containing animation effects determined according to a second animation effect video and the video of the second device, the second animation effect video being a video in which the video of the first device is added with animation effects.
2. The method of claim 1, wherein, The method further comprises: The first device receives capability information of the second device, the capability information of the second device indicating whether the second device supports animation effects.
3. The method of claim 2, wherein, The method further comprises: The first device sends a session invitation message to the second device, the session invitation message being used for the first device to request to establish a video call with the second device; The first device receives capability information of the second device, comprising: The first device receives a response message returned by the second device according to the session invitation message, the response message containing the capability information of the second device.
4. The method of claim 3, wherein, The session invitation message contains capability information of the first device, the capability information of the first device indicating whether the first device supports animation effects.
5. The method of claim 4, wherein, In a case where the first device supports animation effects, the capability information of the first device further indicates that the animation effects supported by the first device are animation effects for a video user or animation effects for a video display window.
6. The method of any of claims 1-5, wherein: The video stream is obtained by superimposing a picture of the first animation effect video on a picture of the video of the first device, the picture of the first animation effect video having a size smaller than a size of the picture of the video of the first device; or The video stream is obtained by superimposing a picture of the video of the first device on a picture of the first animation effect video, the picture of the video of the first device having a size smaller than a size of the picture of the first animation effect video. The method further comprises:
7. The method of claim 6, wherein, If the second device is used to display the video of the first device and the first device supports animation effects for a video user, the first device superimposes a picture of the first animation effect video on a picture of the video of the first device and encodes to obtain the video stream; or If the second device is used to display the video of the second device and the first device supports animation effects for a video user, the first device superimposes a picture of the video of the first device on a picture of the first animation effect video and encodes to obtain the video stream. The method further comprises: The first device sends information indicating a superimposition area to the second device; 8. The method according to claim 6 or 7, characterized in that, If the video stream is obtained by superimposing the picture of the first animation effect video on the picture of the video of the first device, the superimposition area is an area occupied by the picture of the first animation effect video in the picture of the video of the first device; or If the video stream is obtained by superimposing the picture of the video of the first device on the picture of the first animation effect video, the superimposition area is an area occupied by the picture of the video of the first device in the picture of the first animation effect video. If the video stream is obtained by superimposing the picture of the first device video onto the picture of the first dynamic effect video, the superimposition region is a region occupied by the picture of the first dynamic effect video in the picture of the first device video.
9. The method of any of claims 1-5, wherein: the video stream is obtained by superimposing the picture of the first dynamic effect video onto the picture of the first device video, and the picture of the first dynamic effect video has a size smaller than that of the picture of the first device video; or; the video stream is obtained by superimposing the picture of the second device video onto the picture of the second dynamic effect video, and the picture of the second device video has a size smaller than that of the picture of the second dynamic effect video.
10. The method of claim 9, wherein, The method further comprises: if the second device is configured to display the first device video and the first device supports dynamic effects for a video display window, the first device superimposes the picture of the first dynamic effect video onto the picture of the first device video and encodes to obtain the video stream; or; if the second device is configured to display the second device video and the first device supports dynamic effects for a video display window, the first device superimposes the picture of the second device video onto the picture of the second dynamic effect video and encodes to obtain the video stream.
11. The method according to claim 9 or 10, characterized in that, The method further comprises: the first device sends information indicating a superimposition region to the second device; if the video stream is obtained by superimposing the picture of the first dynamic effect video onto the picture of the first device video, the superimposition region is a region occupied by the picture of the first dynamic effect video in the picture of the first device video; or; if the video stream is obtained by superimposing the picture of the second dynamic effect video onto the picture of the second device video, the superimposition region is a region occupied by the picture of the second device video in the picture of the second dynamic effect video.
12. The method according to any one of claims 6-11, characterized in that, The method further comprises: the first device receives indication information from the second device, the indication information indicating that the second device is configured to display the first device video, or the indication information indicating that the second device is configured to display the second device video.
13. The method of any of claims 1-5, wherein: the first device supports dynamic effects for a video user, and the video stream comprises a first video stream and a second video stream; the first video stream is a stream of the first device video, and the second video stream is a stream of the first dynamic effect video.
14. The method of claim 13, wherein, the first device sends a video stream to the second device, comprising: the first device sends the first video stream to the second device through a first transmission channel, and sends the second video stream to the second device through a second transmission channel.
15. The method of any of claims 1-5, wherein: The first device supports a dynamic effect for a video display window, the video stream includes a first video stream and a second video stream, the first video stream is obtained by encoding a video of the first device, and the second video stream is obtained by encoding a first dynamic effect video; or the video stream includes a third video stream and a fourth video stream, the third video stream is obtained by encoding a video of the second device, and the fourth video stream is obtained by encoding a second dynamic effect video.
16. The method of claim 15, wherein, The method further includes: If the second device is used to display the video of the second device, the first device sends the first video stream to the second device through a first transmission channel and sends the second video stream to the second device through a second transmission channel; Or; If the second device is used to display the video of the first device, the first device sends the fourth video stream to the second device through the first transmission channel and sends the third video stream to the second device through the second transmission channel.
17. The method of claim 16, wherein, The method further includes: The first device receives indication information from the second device, the indication information indicating that the second device is used to display the video of the first device, or the indication information indicating that the second device is used to display the video of the second device.
18. The method of claim 14 or 16, characterized in that: The first transmission channel is a transmission channel for a video call established by negotiation between the first device and the second device; The second transmission channel is a transmission channel established by negotiation between the first device and the second device in the case that the second device does not support a dynamic effect.
19. A method of communication, comprising: Applied to a second device, the method includes: The second device sends a video of the second device to a first device; The second device receives a video stream of the first device; wherein, in the case that the first device supports a dynamic effect and the second device does not support a dynamic effect, the video stream is a video stream containing a dynamic effect determined according to a video of the first device and a first dynamic effect video, the first dynamic effect video being a video in which a dynamic effect is added to the video of the second device, or the video stream is a video stream containing a dynamic effect determined according to a second dynamic effect video and the video of the second device, the second dynamic effect video being a video in which a dynamic effect is added to the video of the first device; The second device displays a video in the video stream.
20. The method of claim 19, wherein, The method further includes: The second device sends capability information of the second device to the first device, the capability information of the second device indicating whether the second device supports a dynamic effect.
21. The method of claim 20, wherein, The method further includes: The second device receives a session invitation message from the first device, the session invitation message being used for the first device to request to establish a video call with the second device; The second device sends capability information of the second device to the first device, including: The second device sends a response message to the first device according to the session invitation message, the response message containing the capability information of the second device.
22. The method of claim 21, wherein, The session invitation message comprises capability information of the first device, and the capability information of the first device indicates whether the first device supports animation effect.
23. The method of claim 22, wherein, In the case that the first device supports animation effect, the capability information of the first device further indicates that the animation effect supported by the first device is animation effect for a video user or animation effect for a video display window.
24. The method of any one of claims 19-23, wherein, The second device comprises a first window and a second window, the video stream is one video stream, and the second device displays video in the video stream, comprising: If the second window is contained in the first window, the second device displays video in the video stream in the first window and hides the second window; If the first window is contained in the second window, the second device displays video in the video stream in the second window and hides the first window.
25. The method of claim 24, wherein, The method further comprises: The second device sends indication information to the first device; If the second window is contained in the first window, the indication information indicates that the second device is used to display video of the first device; if the first window is contained in the second window, the indication information indicates that the second device is used to display video of the second device.
26. The method of claim 25, wherein, The second device sends indication information to the first device, comprising: In response to a switching window operation triggered by a user in an overlay area, the second device sends the indication information to the first device; if the overlay area is an area occupied by a picture of the first animation effect video in a video picture of the first device or an area occupied by a video picture of the second device in a picture of the second animation effect video, the switching window operation indicates switching from the second window being contained in the first window to the first window being contained in the second window; if the overlay area is an area occupied by a video picture of the first device in a picture of the first animation effect video, the switching window operation indicates switching from the first window being contained in the second window to the second window being contained in the first window.
27. The method of claim 26, wherein, The method further comprises: The second device receives information from the first device for indicating the overlay area.
28. The method of any one of claims 19-23, wherein, The second device comprises a first window and a second window, the video stream comprises a first video stream and a second video stream, the first video stream is obtained by encoding video of the first device, the second video stream is obtained by encoding the first animation effect video, and the second device receives the video stream of the first device, comprising: The second device receives the first video stream from the first device through a first transmission channel and receives the second video stream from the first device through a second transmission channel; The second device displays video in the video stream, comprising: The second device displays video in the first video stream in the first window and displays video in the second video stream in the second window.
29. The method of any one of claims 19-23, wherein, The second device comprises a first window and a second window; if the video stream comprises a first video stream and a second video stream, the first video stream is encoded from the video of the first device, and the second video stream is encoded from the video of the first device, the second device receives the video stream of the first device, comprising: The second device receives the first video stream from the first device through a first transmission channel, and receives the second video stream from the first device through a second transmission channel; The second device displays the video in the video stream, comprising: The second device displays the video in the first video stream in the first window, and displays the video in the second video stream in the second window, the first window is contained in the second window; If the video stream comprises a third video stream and a fourth video stream, the third video stream is encoded from the video of the second device, and the fourth video stream is encoded from the video of the second device, the second device receives the video stream of the first device, comprising: The second device receives the fourth video stream from the first device through a first transmission channel, and receives the third video stream from the first device through a second transmission channel; The second device displays the video in the video stream, comprising: The second device displays the video in the fourth video stream in the first window, and displays the video in the third video stream in the second window, the second window is contained in the first window.
30. The method of claim 29, wherein, The method further comprises: The second device sends indication information to the first device; If the second window is contained in the first window, the indication information indicates that the second device is used to display the video of the first device; if the first window is contained in the second window, the indication information indicates that the second device is used to display the video of the second device.
31. The method of claim 30, wherein, The second device sends indication information to the first device, comprising: In response to a user triggered window switching operation, the second device sends indication information to the first device; the window switching operation indicates switching from the second window being contained in the first window to the first window being contained in the second window, or switching from the first window being contained in the second window to the second window being contained in the first window.
32. A communications device, characterized by The communication device is configured to perform the method of any one of claims 1-31.
33. A communications device, characterized by Comprising: A processor and a memory; The memory is configured to store computer instructions, when the processor executes the instructions, to make the communication device perform the method of any one of claims 1-31.
34. A communication system, characterized by Comprising: A first device configured to perform the method of any one of claims 1-18, and a second device configured to perform the method of any one of claims 19-31.
35. A communications device, characterized by Comprising: A processor and an interface circuit; wherein, The interface circuit is configured to receive code instructions and transmit them to the processor; The processor is configured to run the code instructions to perform the method of any one of claims 1-31.
36. A communications device, characterized by The communication device comprises a processor and a transceiver for information exchange between the communication device and other communication devices, and the processor executes program instructions to perform the method of any one of claims 1-31.
37. The communication apparatus of claim 36, wherein The communication device is a chip.
38. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a computer program or instructions, which, when executed on a computer, cause the computer to perform the method of any one of claims 1-31.
39. A computer program product, characterised in that, The computer program product comprises a computer program or instructions, which, when executed on a computer, cause the computer to perform the method of any one of claims 1-31.