Display device and video transmission method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JUHAOKAN TECH CO LTD
- Filing Date
- 2022-11-11
- Publication Date
- 2026-07-24
Smart Images

Figure CN117336552B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and more particularly to a display device and a method for video transmission. Background Technology
[0002] With the rapid development of display devices, the functions they offer users are becoming increasingly diverse. Currently, display devices include televisions, set-top boxes, and other products with integrated display screens. Taking televisions as an example, they are expanding into more and more scenarios, no longer limited to watching television programs at home, but also used for video conferencing and other purposes.
[0003] During video conferencing, a common problem is that images containing human figures displayed on the screen are often blurry. Therefore, how to ensure users see clearer images of people during video conferencing has become a pressing issue for those skilled in the art. Summary of the Invention
[0004] This application provides a display device and a video transmission method. In this method, based on the resolution of a first video captured by the display device and the available uplink bandwidth, it is determined whether the first video needs to be processed to obtain a second video including a clear human image, thereby improving the user experience.
[0005] In a first aspect, a display device is provided, comprising:
[0006] A monitor is used to display the user interface.
[0007] User interface, used to receive input signals;
[0008] The controllers, which are connected to the display and the user interface respectively, are configured as follows:
[0009] In response to receiving a command to join a video call, the system receives the first video captured by the camera.
[0010] If the resolution of the first video is greater than the preset resolution, then the available uplink bandwidth is obtained, wherein the preset resolution represents the maximum resolution requirement value of the video in the video call link.
[0011] If the available uplink bandwidth is detected to be no greater than a preset bandwidth, the image in the first video is cropped to convert the first video into a second video, and the second video is sent to the other end, wherein the resolution of the second video is equal to a preset resolution, the preset bandwidth is the bandwidth required to send the first video, and the other end refers to other terminal devices in the video call.
[0012] In some embodiments, the controller is further configured to receive a third video sent by the peer and control the display to play the third video.
[0013] In some embodiments, the first video includes a first video frame; the controller, which performs the step of cropping around a person in the first video to convert the first video into a second video, is further configured to:
[0014] Identify the human image in the first video frame;
[0015] Based on the boundary points of the human figure, a first preset area is determined;
[0016] If the resolution of the first preset region is not greater than the preset resolution, then the image in the first video frame is cropped to obtain a fourth video frame, so that the resolution of the fourth video frame is equal to the preset resolution.
[0017] A second video is generated based on the fourth video frame.
[0018] In some embodiments, the controller, which performs the step of cropping around the human image in the first video to convert the first video into a second video, is further configured to:
[0019] If the resolution of the first preset region is greater than the preset resolution, then the non-human image region of the first video frame is blurred to obtain the fifth video frame, so that the storage space occupied by the fifth video frame is less than the storage space occupied by the first video frame.
[0020] A second video is generated based on the fifth video frame, wherein the bandwidth required to transmit the second video is no greater than a preset bandwidth.
[0021] In some embodiments, the controller is further configured to send the first video to the other end if the resolution of the first video is not greater than a preset resolution.
[0022] In some embodiments, the controller is further configured to:
[0023] When the detected uplink available bandwidth is greater than the preset bandwidth, a first encoding state is determined;
[0024] If the first encoding state is an encodeable state, indicating that a video with a resolution greater than a preset resolution is supported by the peer, a first request is sent to the server. The first request is used to request feedback information indicating whether a first video with a resolution greater than the preset resolution can be sent. This allows the server to determine and send feedback information to the display device based on the second encoding state, where the second encoding state refers to the encoding state stored in the server. If the feedback information indicates that a first video with a resolution greater than the preset resolution can be sent, the server changes the second encoding state from an encodeable state to an unencodeable state, indicating that a video with a resolution greater than the preset resolution is not supported by the peer. The server then sends the changed second encoding state to the peer so that the peer changes the corresponding third encoding state. If the feedback information indicates that a first video with a resolution greater than the preset resolution cannot be sent, the server does not change the second encoding state.
[0025] If the feedback information received by the display device indicates that a first video with a resolution greater than a preset resolution can be sent, the first video is sent to the other end;
[0026] If the feedback information received by the display device indicates that a first video with a resolution greater than a preset resolution cannot be sent, then the first video is converted into a second video; and the second video is sent to the other end.
[0027] In some embodiments, the controller is further configured to:
[0028] When the display device sends a first video with a resolution greater than a preset resolution and receives an instruction to turn off the camera or leave the video call, it sends a second request to the server. The second request is used to change the second encoding state so that the server changes the second encoding state from an unencoding state to an encoding state according to the second request, and sends the changed second encoding state to the peer so that the peer changes the third encoding state according to the changed second encoding state.
[0029] In some embodiments, the controller is further configured to:
[0030] When the resolution of the first video is greater than the preset resolution, the available uplink bandwidth is greater than the preset bandwidth, the first encoding state is an unencoding state, and when the second video is sent, the server sends a changed second encoding state, wherein the changed second encoding state is an encoding state. Based on the changed second encoding state, the first encoding state is changed, and the first video is sent to the other end.
[0031] Secondly, a video transmission method is provided, including:
[0032] In response to receiving a command to join a video call, the system receives the first video captured by the camera.
[0033] If the resolution of the first video is greater than the preset resolution, then the available uplink bandwidth is obtained, wherein the preset resolution represents the preset resolution requirement value of the video in the video call link;
[0034] If the available uplink bandwidth is detected to be no greater than a preset bandwidth, the image in the first video is cropped to convert the first video into a second video, and the second video is sent to the other end, wherein the resolution of the second video is equal to a preset resolution, the preset bandwidth is the bandwidth required to send the first video, and the other end refers to other terminal devices in the video call.
[0035] In some embodiments, the step of converting the first video into a second video includes:
[0036] Identify the human image in the first video frame;
[0037] Based on the boundary points of the human figure, a first preset area is determined;
[0038] If the resolution of the first preset region is not greater than the preset resolution, then the image in the first video frame is cropped to obtain a fourth video frame, so that the resolution of the fourth video frame is equal to the preset resolution.
[0039] Based on the fourth video frame, a second video is generated. In the display device and video transmission method provided in the above embodiments, the method determines whether the first video needs to be processed to obtain a second video including a clear portrait, based on the resolution of the first video captured by the display device and the available uplink bandwidth, thereby improving the user experience. The method includes: in response to receiving an instruction to join a video call, receiving a first video captured by a camera; if the resolution of the first video is greater than a preset resolution, obtaining the available uplink bandwidth, wherein the preset resolution represents a pre-set resolution requirement value for the video in the video call link; if the available uplink bandwidth is detected to be less than the preset bandwidth, cropping around the portrait in the first video to convert the first video into a second video, and sending the second video to the other end, wherein the resolution of the second video is equal to the preset resolution, the preset bandwidth is the bandwidth required to send the first video, and the other end refers to other terminal devices in the video call. Attached Figure Description
[0040] Figure 1 An operational scenario between a display device and a control device according to some embodiments is illustrated;
[0041] Figure 2 A hardware configuration block diagram of a control device 100 according to some embodiments is shown;
[0042] Figure 3 A hardware configuration block diagram of a display device 200 according to some embodiments is shown;
[0043] Figure 4 A software configuration diagram of a display device 200 according to some embodiments is shown;
[0044] Figure 5 An exemplary flowchart of a video transmission method provided according to some embodiments is shown;
[0045] Figure 6 An exemplary schematic diagram of a user interface provided according to some embodiments is shown;
[0046] Figure 7 A schematic diagram of yet another user interface provided according to some embodiments is shown as an example;
[0047] Figure 8 An exemplary schematic diagram of a first video frame provided according to some embodiments is shown;
[0048] Figure 9 An exemplary schematic diagram of yet another first video frame provided according to some embodiments is shown;
[0049] Figure 10 An exemplary schematic diagram of yet another user interface provided according to some embodiments is shown;
[0050] Figure 11 A flowchart of yet another video transmission method provided according to some embodiments is shown as an example. Detailed Implementation
[0051] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0052] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0053] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0054] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0055] The display device provided in this application can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.
[0056] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.
[0057] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.
[0058] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.
[0059] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.
[0060] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.
[0061] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.
[0062] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0063] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0064] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.
[0065] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0066] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.
[0067] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.
[0068] The user interface can be used to receive control signals from the control device 100 (such as an infrared remote control).
[0069] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0070] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.
[0071] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.
[0072] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0073] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can execute operations related to the object selected by the user command.
[0074] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.
[0075] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0076] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0077] See Figure 4 In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Layer"), and the kernel layer.
[0078] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs, system settings programs, or clock programs that come with the operating system; they may also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0079] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0080] like Figure 4As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: ActivityManager, which interacts with all activities running in the system; LocationManager, which provides access to system location services for system services or applications; PackageManager, which retrieves various information related to application packages currently installed on the device; NotificationManager, which controls the display and clearing of notification messages; and WindowManager, which manages icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0081] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling display window changes (e.g., shrinking the display window, shaking the display, distorting the display, etc.).
[0082] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0083] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0084] The terminal devices participating in a video call include the local end and the peer end, and there can be multiple peer ends. The local end refers to the local terminal device, and the peer end refers to other terminal devices in the video call.
[0085] In this embodiment of the application, the local and remote ends participating in the video call can be the same type of terminal device, or they can be different types of devices. In one example, the local end can be a display device, and the remote end can be a mobile terminal. In another example, both the local and remote ends are display devices.
[0086] In this embodiment of the application, the video call can exist in a variety of scenarios. In addition to the video conferencing mentioned above, it can also be applied to scenarios such as online education, online medical care, financial services, and social games.
[0087] During video conferencing, a common problem is that images containing human figures displayed on the screen are often blurry. Therefore, how to ensure users see clearer images of people during video conferencing has become a pressing issue for those skilled in the art.
[0088] To address the aforementioned technical problems, embodiments of this application provide a video transmission method. In this method, based on the resolution of a first video frame in a first video captured by a display device and the available uplink bandwidth, it is determined whether the first video needs to be processed to obtain a second video that includes a clear image of a person to the greatest extent possible, thereby improving the user experience.
[0089] To clearly describe the video transmission method in the embodiments of this application, the local end is taken as the subject of the description. In fact, the methods performed by the local end and the remote end are the same, and the descriptions of the local end and the remote end are also the same.
[0090] Figure 5 An exemplary flowchart of a video transmission method provided according to some embodiments is shown, the method including steps S100-S700.
[0091] S100. When the local terminal receives an instruction to join a video call, it receives a first video captured by the camera, wherein the first video includes at least one first video frame.
[0092] In this embodiment, the camera on the local end can be configured as a 4K camera, capable of shooting 4K video. The camera can also be configured as a non-4K camera, capable of shooting video with a resolution other than 4K. For example, the camera can be a 1080P camera, capable of shooting video with a resolution of 1080P. In some embodiments, when the local end receives a command to join a video call, the camera is activated and begins shooting video.
[0093] In some embodiments, joining a video call can mean that there is currently no video call, but the local device initiates and joins the video call. For example, Figure 6 The illustration shows a schematic diagram of a user interface according to some embodiments, in which a user can select users they wish to participate in a video call. Figure 6 The screen displays a video call initiation control 601. The user selects the video call initiation control via the control device, inputs a command to join the video call, and simultaneously sends an invitation to other users' terminal devices. For example, this end is the initiator of the video call; at this time, the user interface can display... Figure 6 The user interface shown.
[0094] In other embodiments, joining a video call can also refer to joining an existing video call initiated by the other end. For example, Figure 7 An exemplary schematic diagram of yet another user interface provided according to some embodiments is shown. Figure 7 The screen displays a "Join Video Call" control 701. The user selects the "Join Video Call" control via the control device and enters the command to join the video call. Once the other end has initiated a video call, the local end is invited to join, and the user interface displays... Figure 7 The user interface shown.
[0095] S200: Detect whether the resolution of the first video is greater than the preset resolution.
[0096] In this embodiment of the application, the preset resolution represents the maximum preset resolution requirement value of the video in the video call link.
[0097] In some embodiments, the preset maximum resolution requirement for the video in the video call link is determined through negotiation by the terminal devices in the video call. In other embodiments, it may also be pre-configured by the operator. In still other embodiments, it may be the maximum image and video resolution supported when multiple video streams exist in the video call link. Specifically, when the local device plays multiple videos simultaneously, the maximum resolution of the multiple videos it can handle is simultaneously a preset resolution, meaning the local device can play multiple videos at preset resolutions simultaneously. For example, the preset resolution may be 1080P.
[0098] S300. If the resolution of the first video is not greater than the preset resolution, the local end directly sends the first video to the remote end, and the remote end receives and plays the first video.
[0099] S400. If the resolution of the first video is greater than the preset resolution, then obtain the available uplink bandwidth.
[0100] The available uplink bandwidth refers to the bandwidth currently available on the display device for sending the first video. There is also a corresponding available downlink bandwidth, which refers to the bandwidth used to download video captured by the other end. For example, the available uplink bandwidth for sending 4K video could be 3 Mbps. When the available uplink bandwidth is greater than the preset bandwidth, the local end can send the first video. When the available uplink bandwidth is not greater than the preset bandwidth, the local end cannot send the first video.
[0101] S500. Detect whether the available uplink bandwidth is greater than the preset bandwidth, where the preset bandwidth is the bandwidth required to send the first video.
[0102] In this embodiment of the application, since the available uplink bandwidth will affect the uploading of the first video on the local end, the available uplink bandwidth is detected.
[0103] S600. If the available uplink bandwidth is detected to be no greater than the preset bandwidth, the first video is converted, the second video is determined, and the second video is sent to the other end.
[0104] In this embodiment, if the available uplink bandwidth is not greater than the preset bandwidth, it indicates that the local end cannot send the first video. Therefore, this embodiment needs to process the first video to determine the second video. The specific steps for determining the second video are described below.
[0105] In some embodiments, the local device also receives a third video sent by the peer device and controls the display to play the third video. In this embodiment, the display device can simultaneously capture the first video and receive and play the third video captured by the peer device. For example, in addition to displaying the first video captured by the local device, the user interface of the local device can also display multiple third videos captured by the peer device.
[0106] In some embodiments, the step of determining the second video includes:
[0107] Identify the human image in the first video frame. In this embodiment, when the available uplink bandwidth is not greater than the preset bandwidth, the first video cannot be sent directly. Therefore, the first video needs to be processed to determine the second video. In this embodiment, in order to ensure the effect of the video call, the human image in each frame of the first video is identified to ensure that the video sent to the other end includes a relatively clear human image, thereby improving the effect of the video call.
[0108] The first preset area is determined based on the boundary points of the human figure.
[0109] In some embodiments, the step of determining a first preset region based on the boundary points of the human image includes:
[0110] A coordinate system is established with the top-left corner of the first video frame as the origin to determine the coordinates of the boundary points of the human image. Specifically, the boundary points include the top, bottom, left, and right endpoints of the human image. In some embodiments, the human image includes images corresponding to the user's head and neck. For example, Figure 8 An exemplary diagram of a first video frame provided according to some embodiments is shown. Figure 8The image displays the first video frame, where point A is the top endpoint of the image, point B is the bottom endpoint, point C is the left endpoint, and point D is the right endpoint. The first preset area 801 is a rectangle enclosed by the boundary points.
[0111] If the resolution of the first preset region is not greater than the preset resolution, then the image in the first video frame is cropped to obtain a fourth video frame, so that the resolution of the fourth video frame is equal to the preset resolution.
[0112] In some embodiments, the step of cropping around the human image in the first video frame to obtain a fourth video frame includes: determining the center point of a first preset region; determining a second preset region based on the center point of the first preset region, wherein the center point of the second preset region coincides with the center point of the first preset region, and the resolution of the second preset region is a preset resolution; and cropping the second preset region to obtain the fourth video frame. For example, Figure 9 An exemplary schematic diagram of another first video frame provided according to some embodiments is shown, where point E is the center point of the first preset region 901, the center point is used as the center point of the second preset region, the second preset region 902 is cropped out, and the second preset region is used as the fourth video frame. Figure 10 An exemplary illustration shows yet another user interface diagram provided according to some embodiments, in Figure 10 The system displays a fourth video frame, and while displaying the fourth video frame, the third video sent by the other end can also be displayed on the right side of the user interface. Figure 10 The screen shows three third videos sent by the other end.
[0113] A second video is generated based on the fourth video frame. In this embodiment, the first video includes first video frames, and a fourth video frame is determined for each first video frame. All fourth video frames are combined to form the second video.
[0114] Compared to methods that directly compress 4K video to 1080P video, where the maximum width of a person's head in 4K video is 500px and the maximum width after compression to 1080P is 250px, in this embodiment, the maximum width of a person's head in 4K video remains 500px after conversion to 1080P video, and the image clarity is not reduced.
[0115] In some embodiments, the method further includes: if the resolution of the first preset region is greater than the preset resolution, then blurring the non-human image region of the first video frame to obtain a fifth video frame, so that the storage space occupied by the fifth video frame is less than the storage space occupied by the first video frame.
[0116] Thus, by blurring the non-human image area, the color difference of pixels in the blurred non-human image area becomes smaller, resulting in a smaller file size for the fifth video frame, meaning the storage space occupied by the fifth video frame is smaller than that occupied by the first video frame. The method for blurring the non-human image area of the first video frame in this embodiment is not limited.
[0117] In this embodiment, if the resolution of the first preset region is greater than the preset resolution, it means that the human image is too large and cannot be fully displayed within the preset resolution region. Therefore, in order to avoid abruptly segmenting the human image, the non-human image region in the first video frame is directly blurred to obtain the fifth video frame.
[0118] A second video is generated based on the fifth video frame, wherein the bandwidth required to transmit the second video is no greater than a preset bandwidth. In this embodiment, since the first video obtained by this end includes at least one first video frame, when the resolution of the first preset region is greater than the preset resolution, all first video frames in the first video are blurred, and each first video frame will obtain a corresponding fifth video frame. The fifth video frames are arranged in the order of the first video frames to form the second video. In this way, the bandwidth required to transmit the second video composed of the fifth video frames is reduced compared to the bandwidth required to transmit the first video, so the transmission rate of the second video can be accelerated. Furthermore, during blurring, only non-human image parts are blurred, and human image parts in the first video are not blurred, so that the image clarity in the second video is still the same as that in the first video, improving the user experience.
[0119] In other embodiments, the method further includes: if the resolution of the first preset region is greater than the preset resolution, then filling the non-human image region of the first video frame with monochrome values to obtain a sixth video frame.
[0120] In this embodiment, filling the non-human image area of the first video frame with monochrome values can make the color difference of the pixels in the non-human image area zero, thereby reducing the size of the sixth video frame, that is, reducing the storage space occupied by the sixth video frame compared to the storage space occupied by the first video frame.
[0121] A second video is generated based on the sixth video frame, wherein the bandwidth required to transmit the second video is no greater than a preset bandwidth.
[0122] In this way, the bandwidth required to transmit the second video, which consists of the sixth video frame, is reduced compared to the bandwidth required to transmit the first video. Therefore, the transmission rate of the second video can be accelerated. Furthermore, during the blurring process, only the non-human parts are blurred, and the human parts in the first video are not blurred. This ensures that the clarity of the human figures in the second video is still the same as that in the first video, thus improving the user experience.
[0123] In some embodiments, before blurring or filling the non-human image area of the first video frame with monochrome values, the first video frame can be cropped according to the human image area in the first video frame to obtain a seventh video frame. The resolution of the non-human image area in the seventh video frame is lower than the resolution of the non-human image area in the first video frame. Blurring or filling the non-human image area with monochrome values is then performed on the seventh video frame. This reduces the area of the blurred or filled non-human image area and increases the area occupied by the human image in the second video, improving the user's viewing experience.
[0124] S700: If the available uplink bandwidth is greater than the preset bandwidth, then determine the first encoding state.
[0125] Due to limitations in preset resolution, for example, while current display devices' built-in cameras already support 4K video capture and can shoot high-resolution videos, in multi-person real-time video calls, the clarity of the played video often suffers from limitations in the bandwidth required for 4K video transmission and the decoding capabilities of the display device. Decoding capability refers to the ability of a display device to decode only one 4K video stream. For instance, when display device A plays a 4K video it has acquired, it cannot play a 4K video sent by display device B.
[0126] Therefore, to avoid the problem of display devices being unable to play multiple 4K videos normally, in some embodiments, while ensuring that the display device can play multiple videos normally, the display device can accept multiple 1080P videos to a maximum extent. The display device has no pressure when playing multiple 1080P videos. Therefore, in some embodiments, all 4K videos captured by all display devices during the video call are directly compressed into 1080P videos before transmission, that is, all are compressed to the preset resolution. This can ensure that the display device can play multiple videos normally. However, compared with 4K videos, the clarity of 1080P videos is significantly reduced, which reduces the effect of video calls for users.
[0127] For example, four users are having a video call through their respective display devices, and the user interface displays the videos captured by the four display devices. For instance, the user interface displays the video corresponding to display devices A and D. Although display devices A and D all have the capability to capture 4K video, only the window corresponding to display device A, which initiates the video call, displays a 4K video. The windows corresponding to the other display devices B and D display videos that are not 4K in resolution, resulting in low video clarity.
[0128] In this embodiment of the application, due to the limited decoding capability of the display device, it can only decode one video with a resolution greater than the preset resolution. The video with a resolution greater than the preset resolution can be a 4K video. When the local end shoots a 4K video and sends it to the other end for playback, if the other end also shoots a 4K video, the local end cannot play the 4K video sent by the other end.
[0129] Therefore, to avoid this situation, a first encoding state is set to indicate whether the peer supports videos with a resolution greater than the preset resolution. The first encoding state includes an encodeable state and a non-encodeable state. The encodeable state indicates that the first video is supported by the peer; in this case, the peer can play the first video captured by the camera with a resolution greater than the preset resolution after encoding by this end. The non-encodeable state indicates that the first video is not supported by the peer; in this case, the peer cannot play the first video captured by the camera with a resolution greater than the preset resolution after encoding by this end. Thus, by using the first encoding state, this avoids sending videos with a resolution greater than the preset resolution to the peer, preventing multiple videos with resolutions greater than the preset resolution from existing simultaneously in a video conference.
[0130] In some embodiments, the first encoding state can be obtained when the local end receives an instruction to join a video call. In some embodiments, it can also be obtained after detecting that the resolution of the first video is greater than a preset resolution.
[0131] In one example, the initial value of the first encoding state is an encodeable state. When the local end initiates a video call, and receives an instruction to join the video call, the first encoding state obtained at this time is an encodeable state. In another example, if the local end is joining an existing video call, and the other end has already sent a video with a resolution greater than a preset value, then when the local end receives an instruction to join the video call, the first encoding state obtained is an unencodeable state.
[0132] In some embodiments, the first encoded state is obtained from a server, which stores a second encoded state. When the local end obtains the first encoded state, the server sends the stored second encoded state as the first encoded state to the local end.
[0133] In some embodiments, if the first encoding state is an encodeable state, a first request is sent to the server. This first request requests feedback indicating whether feedback information for a first video with a resolution greater than a preset resolution can be sent, so that the server determines and sends the feedback information based on the second encoding state. The second encoding state refers to the encoding states stored in the server.
[0134] In some embodiments, the second encoding state includes an encoding state and an unencoding state.
[0135] When the second encoding state in the server is an encodeable state, the feedback information determined based on this second encoding state indicates that a first video with a resolution greater than a preset resolution can be sent. When the second encoding state in the server is a non-encodeable state, the feedback information determined based on this second encoding state indicates that a first video with a resolution greater than a preset resolution cannot be sent.
[0136] In some embodiments, when the local end and the remote end first enter a video conference, the first encoding state they obtain may both be in an encodeable state. This can lead to a situation where, when the uplink available bandwidth of both the local end and the remote end exceeds a preset bandwidth, both ends need to send videos with a resolution greater than the preset resolution. Since the terminal device has limited decoding capabilities for videos with resolutions greater than the preset resolution, it cannot simultaneously decode multiple videos with resolutions greater than the preset resolution. Therefore, in this embodiment, the local end sends a first request to the server to confirm whether it is possible to send a first video with a resolution greater than the preset resolution based on the second encoding state in the server. If the feedback information sent by the server indicates that the first video with a resolution greater than the preset resolution can be sent, then the first video with a resolution greater than the preset resolution can be sent. If the feedback information sent by the server indicates that the first video with a resolution greater than the preset resolution cannot be sent, then the first video with a resolution greater than the preset resolution cannot be sent.
[0137] In some embodiments, the local end sends a first request. If the feedback indicates that a first video with a resolution greater than a preset resolution can be sent, the server changes the second encoding state from an encodeable state to an unencodeable state and sends the changed second encoding state to the peer end to change the third encoding state in the peer end. If the feedback indicates that a first video with a resolution greater than the preset resolution cannot be sent, the second encoding state is not changed.
[0138] The third encoding state of the receiving end includes an encoding state and a non-encoding state. In this embodiment, when a modified second encoding state is received, the third encoding state is changed to the modified second encoding state. For example, if the original third encoding state is an encoding state, and the received modified second encoding state is a non-encoding state, then the third encoding state is changed to a non-encoding state.
[0139] The method for obtaining the third encoding state of the peer end is the same as that for the first encoding state of the local end. When the peer end is in the encodeable state and wants to send a third video with a resolution greater than the preset resolution, it also needs to send a request to the server. The server determines whether the peer end can send a third video with a resolution greater than the preset resolution based on the second encoding state. It is worth noting that the methods executed by the local end and the peer end are the same in this embodiment of the application. However, if the local end initiates a video conference, the time when it starts executing the above methods will be earlier than the time when the peer end starts executing the methods.
[0140] In some embodiments, after both the local and remote ends join the video conference, their respective encoding states are both encodeable; that is, the local end's first encoding state is encodeable, and the remote end's third encoding state is encodeable. When the local end sends a first request, the server changes the second encoding state from encodeable to non-encodeable based on feedback information. At this time, after the remote end sends its first request, the server determines the feedback information to send to the remote end based on the changed second encoding state. This feedback information indicates that video with a resolution greater than a preset resolution cannot be sent. This ensures that only one video with a resolution greater than the preset resolution will ever exist in the video conference.
[0141] In some embodiments, when the local end receives feedback indicating that a video with a resolution greater than a preset resolution can be sent, it sends the first video to the other end. When the local end receives feedback indicating that a video with a resolution greater than the preset resolution cannot be sent, it converts the first video into a second video and sends the second video to the other end.
[0142] In some embodiments, when the local end sends a first video including a resolution greater than a preset resolution and receives an instruction to turn off the camera or leave the video call, it sends a second request to the server. The second request is used to change the second encoding state so that the server changes the second encoding state from an unencoding state to an encoding state according to the second request, and sends the changed second encoding state to the peer end to change the third encoding state in the peer end.
[0143] In this way, the other end can send a third video with a resolution greater than the preset resolution based on the modified third encoding state.
[0144] In one example, the participants in the video call include display device A (AD). Currently, display device A sends a first video with a resolution greater than a preset resolution. When display device A receives an instruction to turn off its camera or leave the video call, its second encoding state changes from an unencoding state to an encoding state, and it notifies display device BD (BD) through the server. Display device BD then changes its third encoding state from an unencoding state to an encoding state.
[0145] In some embodiments, when the resolution of the first video is greater than a preset resolution, the available uplink bandwidth is greater than a preset bandwidth, the first encoding state is an unencoding state, and when the second video is sent, the local end receives the changed second encoding state sent by the server, wherein the changed second encoding state is an encoding state; the local end changes the first encoding state according to the changed second encoding state and sends the first video to the other end.
[0146] In this embodiment, if the peer sends a third video with a resolution greater than a preset resolution, the local end can only send a second video with a resolution no greater than the preset resolution. When the peer receives an instruction to turn off the camera or exit the video call, it sends a request to change the second encoding state to the server. The server can send the changed second encoding state to the local end. Upon receiving the changed second encoding state, the local end changes the corresponding first encoding state and sends a first video with a resolution greater than the preset resolution to the peer still in the video call.
[0147] In some embodiments, the server includes media services and signaling services.
[0148] In some embodiments, the terminal devices participating in a video call include display devices AC, both of which can capture 4K video. Display device A is the local end, display device B is the remote end a, and display device C is the remote end b.
[0149] After the local end initiates a video call, it invites the other end, a and b, to join the video call. Subsequently, the other end, a and b, join the video call.
[0150] In some embodiments, the preset resolution represents the maximum video resolution supported when there are multiple video streams in the video call link. Higher video call clarity leads to a better user experience, and image resolution is one of the means to improve clarity.
[0151] In some embodiments, because the camera's hardware capabilities allow the captured image to have a resolution greater than a preset resolution, compression is generally used for image transmission in existing video calls. However, video call scenarios focus on the interaction between people. Therefore, this application uses a method of cropping the portrait area to at least partially replace compression, avoiding the loss of pixel data in the portrait area during compression, and allowing more details of the person to be presented in the image that better meets the preset resolution.
[0152] In some embodiments, the human image area may occupy a significant portion of the camera's capture area, resulting in the area containing the entire human image being unable to be cropped or failing to meet the preset resolution requirement (i.e., exceeding the preset resolution) even after cropping. Commonly used methods in the prior art include: 1. compressing the cropped human image area to reduce the resolution of the obtained image; 2. reducing the cropping range of the human image area, sacrificing some human features to reduce the resolution of the obtained image. In this application, while ensuring the clarity of the human image in video calls, the integrity of the human image must also be guaranteed. To resolve this contradiction, when the cropped human image area does not meet the preset resolution requirement, this application blurs the area outside the human image. Blurring reduces the RGB data of pixels; therefore, under the same resolution, reducing the RGB data of pixels reduces the bandwidth required to transmit the blurred video compared to the unblurred original video.
[0153] In some embodiments, during a video call, different channels may use the same resolution requirement and the same preset resolution. In some embodiments, different channels may have different resolution requirements. The main window display area is large and requires a high resolution, while the auxiliary window display area is smaller than the main window display area and its resolution may be lower than the main window resolution. Therefore, the preset resolution may include a first preset resolution and a second preset resolution.
[0154] The following scenario illustrates a method for video transmission.
[0155] The scenario is as follows: This device initiates the video call. Devices A and B accept the invitation and join the video call successively. The initial encoding status obtained by this device, device A, and device B after entering the video call is all in an encodeable state. The resolution of the first video captured by the cameras of this device, device A, and device B is greater than the preset resolution, and the available uplink bandwidth is greater than the preset bandwidth.
[0156] Figure 11 A flowchart illustrating another video transmission method according to some embodiments is provided. The method includes steps S1101-S1135.
[0157] S1101. This end acquires the first video captured by the camera. S1102. The peer a acquires video A captured by its own camera. S1103. The peer b acquires video B captured by its own camera.
[0158] S1104. If the local end detects that the resolution of the first video is greater than the preset resolution, then obtain the available uplink bandwidth.
[0159] S1105. When it is detected that the available uplink bandwidth is greater than the preset bandwidth, a first encoding state is determined, wherein the first encoding state is determined from the second encoding state of a video with a resolution greater than the preset resolution obtained from the signaling service, and the first encoding state is an encodingable state.
[0160] S1106. If the first encoding state is an encodeable state, the local end sends a first request to the signaling service. The first request is used to inquire whether it is possible to send a video with a resolution greater than the preset resolution.
[0161] S1107. The signaling service determines feedback information based on the second encoding state, and S1108. sends the feedback information to the local end. S1109. If the signaling service detects that the feedback information indicates that video with a resolution greater than a preset resolution can be sent, it changes the second encoding state from an encodeable state to an unencodeable state. S1110. The signaling service sends the changed second encoding state to peer a, and S1111. It sends the changed second encoding state to peer b. S1112. Peer a changes the third encoding state based on the received changed second encoding state. S1113. Peer b also changes the third encoding state based on the received changed second encoding state.
[0162] S1114. If the feedback information received by this end indicates that a first video with a resolution greater than a preset resolution can be sent, the first video is sent to the media service. S1115. The media service sends the first video to peer a, and S1116. The first video is sent to peer b. S1117. Peer a plays the first video, and S1118. Peer b plays the first video.
[0163] In this embodiment, after peer a joins a video call, S1119, peer a detects that the resolution of video A is greater than a preset resolution, and then obtains the available uplink bandwidth; and S1120, when the available uplink bandwidth of peer a is detected to be greater than the preset bandwidth, the preset bandwidth corresponding to peer a is the bandwidth required to send video A, and the third encoding state of the video with a resolution greater than the preset resolution is determined. Since peer a will receive a third encoding state that is in an encoding state after entering the video conference, S1121, when the third encoding state is an encoding state, peer a sends a first request to the signaling service. The first request is used to inquire whether it is possible to send a video with a resolution greater than the preset resolution. S1122, the signaling service determines feedback information based on the current second encoding state. At this time, the second encoding state has changed from an encoding state to an unencoding state due to the first request sent by the peer, so the feedback information indicates that a video with a resolution greater than the preset resolution cannot be sent. S1123, the feedback information is sent to peer a. S1124, peer a determines a third video based on the feedback information. The resolution of the third video is not greater than the preset resolution. S1125. The peer sends a third video converted from video A to the media service. The resolution of the third video is no greater than a preset resolution. S1126. The media service sends the third video to this end. S1127. The third video is sent to peer b. S1128. This end plays the third video. S1129. Peer b plays the third video.
[0164] After peer b joins the video conference, it detects that the resolution of video B is greater than a preset resolution. It then obtains the available uplink bandwidth. If the available uplink bandwidth for peer b is greater than the preset bandwidth, it determines the third encoding state of the video with a resolution greater than the preset resolution. Since peer b receives a third encoding state that is encodeable after joining the video conference, when the third encoding state is encodeable, peer b sends a first request to the server. This first request inquires whether a video with a resolution greater than the preset resolution can be sent. The signaling service determines feedback information based on the current second encoding state. At this time, the second encoding state has changed from encodeable to non-encodeable due to the first request sent by the local end, so the feedback information indicates that a video with a resolution greater than the preset resolution cannot be sent. Feedback information is sent to peer b. Based on the feedback information, peer b does not send video B. It sends a third video converted from video B, the resolution of which is not greater than the preset resolution. Peer b sends the third video to the local end and also to peer a. The local end plays the third video. Peer a plays the third video. It should be noted that the steps performed by peer b and peer a are the same; however, due to resolution limitations in the illustration, they are not shown here. Figure 11 The text in the annotation refers to the steps performed on the peer b.
[0165] S1130. When the local end receives an instruction to turn off the camera or leave the video call, it sends a second request to the signaling service. S1131. Based on the second request, the signaling service changes the second encoding state from an unencoding state to an encoding state. S1132. The changed second encoding state is sent to the peer a. S1133. Peer a determines a third encoding state based on the changed second encoding state; the changed third encoding state is an encoding state. S1134. The changed second encoding state is sent to peer b. S1135. Peer b determines a third encoding state based on the changed second encoding state; the changed third encoding state is an encoding state. Subsequent steps are the same as those performed after determining the third encoding state, as mentioned above, and will not be repeated.
[0166] This application embodiment also provides a display device, including:
[0167] A monitor is used to display the user interface.
[0168] User interface, used to receive input signals;
[0169] The controllers, which are connected to the display and the user interface respectively, are configured as follows:
[0170] In response to receiving a command to join a video call, the system receives the first video captured by the camera.
[0171] If the resolution of the first video is greater than the preset resolution, then the available uplink bandwidth is obtained, wherein the preset resolution represents the preset resolution requirement value of the video in the video call link;
[0172] If the available uplink bandwidth is detected to be no greater than a preset bandwidth, the image in the first video is cropped to convert the first video into a second video, and the second video is sent to the other end, wherein the resolution of the second video is equal to a preset resolution, the preset bandwidth is the bandwidth required to send the first video, and the other end refers to other terminal devices in the video call.
[0173] In the display device and video transmission method provided in the above embodiments, the method determines whether the first video needs to be processed to obtain a second video including a clear human image based on the resolution of the first video captured by the display device and the available uplink bandwidth, thereby improving the user experience. The method includes: in response to receiving an instruction to join a video call, receiving the first video captured by a camera; if the resolution of the first video is greater than a preset resolution, obtaining the available uplink bandwidth, wherein the preset resolution represents a pre-set maximum resolution requirement value for the video in the video call link; if the available uplink bandwidth is detected to be less than the preset bandwidth, cropping around the human image in the first video to convert the first video into a second video, and sending the second video to the other end, wherein the resolution of the second video is equal to the preset resolution, the preset bandwidth is the bandwidth required to send the first video, and the other end refers to other terminal devices in the video call.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0175] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A display device, characterized in that, include: A monitor is used to display the user interface. User interface, used to receive input signals; The controllers, which are connected to the display and the user interface respectively, are configured as follows: In response to receiving a command to join a video call, the system receives the first video captured by the camera. If the resolution of the first video is greater than the preset resolution, then the available uplink bandwidth is obtained, wherein the preset resolution represents the maximum resolution requirement value of the video in the video call link. If the available uplink bandwidth is detected to be no greater than a preset bandwidth, the image in the first video is cropped to convert the first video into a second video, and the second video is sent to the other end. The resolution of the second video is equal to the preset resolution. The preset bandwidth is the bandwidth required to send the first video. The other end refers to other terminal devices in the video call. If the detected uplink available bandwidth is greater than a preset bandwidth, a first encoding state is determined. When the first encoding state is an encodeable state, the encodeable state indicates that a video with a resolution greater than a preset resolution is supported by the peer. A first request is sent to the server. If the feedback information received by the display device indicates that a first video with a resolution greater than the preset resolution can be sent, the first video is sent to the peer. The first request is used to request feedback information indicating whether a first video with a resolution greater than the preset resolution can be sent, so that the server determines and sends feedback information to the display device according to the second encoding state. The second encoding state refers to the encoding state stored in the server. If the feedback information indicates that a first video with a resolution greater than the preset resolution can be sent, the server changes the second encoding state from an encodeable state to an unencodeable state and sends the changed second encoding state to the peer, so that the peer changes the corresponding third encoding state. If the feedback information indicates that a first video with a resolution greater than the preset resolution cannot be sent, the server does not change the second encoding state. The unencodeable state indicates that a video with a resolution greater than the preset resolution is not supported by the peer.
2. The display device according to claim 1, characterized in that, The controller is also configured to receive a third video sent by the peer and control the display to play the third video.
3. The display device according to claim 1, characterized in that, The first video includes a first video frame; the controller, which performs the step of cropping around a person in the first video to convert the first video into a second video, is further configured to: Identify the human image in the first video frame; Based on the boundary points of the human figure, a first preset area is determined; If the resolution of the first preset region is not greater than the preset resolution, then the image in the first video frame is cropped to obtain a fourth video frame, so that the resolution of the fourth video frame is equal to the preset resolution. A second video is generated based on the fourth video frame.
4. The display device according to claim 3, characterized in that, The controller, which performs the step of cropping around the image in the first video to convert the first video into a second video, is further configured to: If the resolution of the first preset region is greater than the preset resolution, then the non-human image region of the first video frame is blurred to obtain the fifth video frame, so that the storage space occupied by the fifth video frame is less than the storage space occupied by the first video frame. A second video is generated based on the fifth video frame, wherein the bandwidth required to transmit the second video is no greater than a preset bandwidth.
5. The display device according to claim 1, characterized in that, The controller is further configured to send the first video to the other end if the resolution of the first video is not greater than a preset resolution.
6. The display device according to claim 1, characterized in that, The controller is also configured to: When the display device sends a first video with a resolution greater than a preset resolution and receives an instruction to turn off the camera or leave the video call, it sends a second request to the server. The second request is used to change the second encoding state so that the server changes the second encoding state from an unencoding state to an encoding state according to the second request, and sends the changed second encoding state to the peer so that the peer changes the third encoding state according to the changed second encoding state.
7. The display device according to claim 6, characterized in that, The controller is also configured to: When the resolution of the first video is greater than the preset resolution, the available uplink bandwidth is greater than the preset bandwidth, the first encoding state is an unencoding state, and when the second video is sent, the server sends a changed second encoding state, wherein the changed second encoding state is an encoding state. Based on the changed second encoding state, the first encoding state is changed, and the first video is sent to the other end.
8. A video transmission method, characterized in that, include: In response to receiving a command to join a video call, the system receives the first video captured by the camera. If the resolution of the first video is greater than the preset resolution, then the available uplink bandwidth is obtained, wherein the preset resolution represents the preset resolution requirement value of the video in the video call link; If the available uplink bandwidth is detected to be no greater than a preset bandwidth, the image in the first video is cropped to convert the first video into a second video, and the second video is sent to the other end. The resolution of the second video is equal to the preset resolution. The preset bandwidth is the bandwidth required to send the first video. The other end refers to other terminal devices in the video call. If the detected uplink available bandwidth is greater than a preset bandwidth, a first encoding state is determined. When the first encoding state is an encodeable state, the encodeable state indicates that a video with a resolution greater than a preset resolution is supported by the peer. A first request is sent to the server. If the feedback information received by the display device indicates that a first video with a resolution greater than the preset resolution can be sent, the first video is sent to the peer. The first request is used to request feedback information indicating whether a first video with a resolution greater than the preset resolution can be sent, so that the server determines and sends feedback information to the display device according to the second encoding state. The second encoding state refers to the encoding state stored in the server. If the feedback information indicates that a first video with a resolution greater than the preset resolution can be sent, the server changes the second encoding state from an encodeable state to an unencodeable state and sends the changed second encoding state to the peer, so that the peer changes the corresponding third encoding state. If the feedback information indicates that a first video with a resolution greater than the preset resolution cannot be sent, the server does not change the second encoding state. The unencodeable state indicates that a video with a resolution greater than the preset resolution is not supported by the peer.
9. The method according to claim 8, characterized in that, The steps for converting the first video into the second video include: Identify the human image in the first video frame; Based on the boundary points of the human figure, a first preset area is determined; If the resolution of the first preset region is not greater than the preset resolution, then the image in the first video frame is cropped to obtain a fourth video frame, so that the resolution of the fourth video frame is equal to the preset resolution. A second video is generated based on the fourth video frame.