Video conference display method, device, terminal and storage medium

By using multiple display areas for splicing and displaying in video conferences, the problem of cumbersome video operations for viewing specific participants is solved, and the search efficiency is improved.

CN114845083BActive Publication Date: 2025-08-26MIGU CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210501797.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-08-26
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

In video conferencing, it is cumbersome to view videos of specific participants and is inefficient in finding them.

Method used

By splicing and displaying it in multiple display areas of the user display page, highlighting the videos of the target participants, simplifying user operations.

Benefits of technology

It improves the convenience of locating specific participants and simplifies user operation procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114845083B_ABST
    Figure CN114845083B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a video conference display method, device, terminal, and storage medium, relating to the field of video processing technology, to solve the problem of cumbersome operations when locating a speaker in a video conference. The method includes: obtaining a target video of a first participant in a video conference, wherein the video conference includes multiple participants, and the first participant is one of the multiple participants; utilizing multiple display areas included in a user display page of the video conference to splice and display the target video, wherein each display area in the user display page corresponds to a participant. In this manner, the first participant can be highlighted, thereby improving the convenience of locating the first participant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing, and in particular to a video conference display method, device, terminal and storage medium. Background Art

[0002] With the development of technology, remote work is becoming more and more common, making video conferencing essential. Currently, in video conferencing, when there are many participants, if a user needs to view the video of a specific participant, such as the current speaker or department head, they must search for each participant's video one by one, which is cumbersome and inefficient. Summary of the Invention

[0003] The embodiments of the present invention provide a video conference display method, device, terminal and storage medium to solve the problems in the prior art of cumbersome operation and low search efficiency when viewing the video of a specific participant in a video conference.

[0004] The embodiment of the present invention is implemented as follows:

[0005] In a first aspect, an embodiment of the present invention provides a video conference display method, comprising:

[0006] Obtaining a target video of a first participant in a video conference, where the video conference includes multiple participants, and the first participant is one of the multiple participants;

[0007] The target video is spliced ​​and displayed using a plurality of display areas included in the user display page of the video conference, wherein each display area in the user display page corresponds to a participant.

[0008] In a second aspect, an embodiment of the present invention provides a video conference display device, including:

[0009] An acquisition module, configured to acquire a target video of a first participant in a video conference, where the video conference includes multiple participants, and the first participant is one of the multiple participants;

[0010] The display module is used to display the target video by splicing using multiple display areas included in the user display page of the video conference, and each display area in the user display page corresponds to a participant.

[0011] In a third aspect, an embodiment of the present invention further provides a terminal comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the video conferencing display method as described in the first aspect.

[0012] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the video conference display method described in the first aspect are implemented.

[0013] In an embodiment of the present invention, a target video of a first participant in a video conference is obtained, where the video conference includes multiple participants, and the first participant is one of the multiple participants; the target video is spliced ​​and displayed using multiple display areas included in a user display page of the video conference, where each display area in the user display page corresponds to a participant. Through the above method, the terminal displays the video of the first participant on the user display page, and splices and displays the target video using the display areas included in the user display page to highlight the first participant. The user does not need to manually search and locate the first participant, which simplifies user operations and improves the convenience of locating the first participant. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flow chart of a video conference display method provided by an embodiment of the present invention;

[0015] Figure 2a This is one of the video conference display schematic diagrams provided by an embodiment of the present invention;

[0016] Figure 2b This is the second video conference display diagram provided by an embodiment of the present invention;

[0017] Figure 2c This is the third video conference display diagram provided by an embodiment of the present invention;

[0018] Figure 3 is a structural diagram of a video conferencing display device provided by an embodiment of the present invention;

[0019] Figure 4 is a structural diagram of a terminal provided by an embodiment of the present invention;

[0020] Figure 5 is another structural diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a video conference display method provided by an embodiment of the present invention. Figure 1 As shown, this embodiment provides a video conference display method, which is applied to a terminal and includes the following steps:

[0023] Step 101: Acquire a target video of a first participant in a video conference, where the video conference includes multiple participants and the first participant is one of the multiple participants.

[0024] The first participant may be the current speaker among the multiple participants, or may be the first marked user. For example, a participant is marked based on user settings, and the marked participant is the first marked user, which is not limited here.

[0025] Optionally, the video conference includes U pages, each page includes at most N display areas, U is a positive integer, U is a positive integer, N is a positive integer greater than 1, and each participant corresponds to a display area.

[0026] In the case where the first participant is the current speaker, if there are at least two participants speaking at the same time, the speaker who speaks the loudest among the at least two participants can be selected as the first participant, or the speaker corresponding to the video stream first obtained by the terminal among the at least two participants can be selected as the first participant. For example, participant A and participant B speak at the same time, the server obtains the videos of participant A and participant B, and sends the videos to the terminal. The terminal receives the video of participant A first, then the terminal selects participant A as the first participant.

[0027] Due to the different sizes of the terminal display screens, the display areas of the participants that can be displayed on each page may also be different. The number of display areas on each page can be adaptively set according to the size of the terminal display screen, or can be set by the user, or the default setting can be used, which is not limited here. For example, the display area of ​​each page can display the display areas of up to 4 or 6 participants, and the multiple display areas are arranged in the form of multiple rows and columns. Taking 4 display areas as an example, the 4 display areas can be arranged in 2 rows and 2 columns. Generally, the number of participants may not be an integer multiple of the maximum display area that can be displayed on each page. In this case, the last page in the U pages will include less than N display areas. For example, if there are 22 participants and each page displays the display areas of 4 participants, then the video conference has 6 pages. Among the first 5 pages, each page includes 4 display areas, and the last page includes 2 display areas.

[0028] Optionally, in order to achieve a better display effect when splicing and displaying the video of the first participant on the display area included in the page, when the number of display areas included on the page is less than N, the default window can be used to fill the insufficient display area. For example, in the above example, the last page includes two display areas, and these two display areas are displayed in the first row. Then, two default windows are added in the second row to form a 2-row 2-column display area arrangement on the last page. In this way, the arrangement style of the display areas on the last page is the same as that of the display areas on other pages, both arranged in 2 rows and 2 columns. The size of the default window is the same as that of the other display areas.

[0029] Optionally, when the number of display areas included in a page is less than N, the size of the display areas included in the page is adjusted. For example, the size of the display areas can be increased. For example, in the above example, if the page displays a maximum of 4 windows, and the 4 windows are arranged in 2 rows and 2 columns when displayed, the last page includes 2 display areas. If not adjusted, these 2 display areas are displayed in the first row of 2 rows and 2 columns. The sizes of these 2 display areas are adjusted. For example, the widths of the 2 display areas are kept unchanged, and the lengths are adjusted to 2 times. The 2 resized display areas can be displayed across the 2 rows before adjustment, so that when the video of the first participant is displayed in a spliced ​​manner, a better display effect is achieved.

[0030] The server can set an order for the participants in the video conference, for example, according to the order in which each participant joins the conference. This order can be used as an identifier for the participant, and the identifier is sent to the terminal. When the terminal displays these multiple parameter participants, it can sort them according to the participant identifiers and display them in this order on pages 1 to U. For example, if the number of participants is 8 and each page includes 4 display areas, the display areas of the 4 participants ranked first are displayed on page 1. The display order can be from left to right and from top to bottom, which is not limited here. The display areas of the 4 participants ranked last are displayed on page 2.

[0031] Step 102: Utilize the multiple display areas included in the user display page of the video conference to splice and display the target video, where each display area in the user display page corresponds to one participant.

[0032] The target video is spliced ​​and displayed using the display area included in the user display page, and the user display page is one of the U pages.

[0033] The splicing display can be understood as that at the same time, each display area included in the user display page displays a partial picture of the same video frame of the target video, and each display area is spliced ​​into a complete picture of the video frame, that is, each display area is used to perform a pseudo full-screen display of the video frame, such as Figure 2a The figure shows a schematic diagram of the page display when no one is speaking. In the figure, reference numeral 1 indicates a display area, reference numeral 11 indicates an area in the display area that displays the video of the participants, and reference numeral 22 indicates an area in the display area that displays the personal information of the participants. Figure 2b The figure shows a schematic diagram of the page display when someone is speaking. The speaker is participant A, and the video of participant A is spliced ​​and displayed in four display areas.

[0034] Optionally, when the terminal uses the display area included in the user display page to splice and display the target video, it can also simultaneously display the personnel information of the participants corresponding to the display area in the display area, for example, the nickname, status information (whether the microphone is turned on, etc.), gender, position, whether there are special marks, etc. of the participants in the video conference can be displayed in the middle area or other areas of the display area, such as Figure 2a The “Participant A”, “Participant B”, etc. shown in the figure are the personnel information of the participants.

[0035] Optionally, if the user display page includes a specific display area (see the description below for details), the specific display area does not participate in the display of the partial picture of the video frame. Accordingly, the partial picture that should have been displayed in the specific display area will be displayed by default, that is, the picture spliced ​​together by each display area is not the complete picture of the video frame, and some partial pictures are displayed by default.

[0036] The user display page is the page where the display area corresponding to the first participant is located. For example, the display area corresponding to the first participant is on page 2, then page 2 is the user display page. If the page currently displayed by the terminal is page 1, the terminal can splice and display the target video on the display area included in page 2.

[0037] Alternatively, the user display page is the page currently displayed by the terminal. For example, if the page currently displayed by the terminal is page 1, then page 1 is the user display page, and the terminal can splice and display the target video on the display area included in page 1.

[0038] The terminal can be a mobile phone, tablet computer, laptop computer, desktop computer, smart watch, etc., without limitation here.

[0039] In this embodiment, a target video of the first participant in a video conference is obtained, where the video conference includes multiple participants, and the first participant is one of the multiple participants. The target video is spliced ​​and displayed using multiple display areas included in a user display page of the video conference, where each display area in the user display page corresponds to a participant. In this manner, the terminal displays the video of the first participant on the user display page, and splices and displays the target video using the display areas included in the user display page to highlight the first participant. The user does not need to manually search and locate the first participant, which simplifies user operations and improves the convenience of locating the first participant.

[0040] It should be noted that when the terminal executes the method in this application, the terminal can obtain the video of each participant, for example, each participant has turned on the camera, or the terminal can obtain the video of each current speaker, for example, each participant must turn on the camera at least before speaking, or, the video of the participant obtained by the terminal is the default video, for example, the current speaker has not turned on the camera, the terminal plays the default video and adds a watermark of the current speaker's personal information in each video frame of the default video.

[0041] In one embodiment of the present application, the step of splicing and displaying the target video using the multiple display areas included in the user presentation page of the video conference includes:

[0042] Splitting the target video according to the position of each display area on the user display page to obtain split videos corresponding to each display area;

[0043] The split video corresponding to each display area is used to replace the initial video corresponding to the display area, and the initial video of the display area is the video of the participant corresponding to the display area.

[0044] Specifically, each display area corresponds to a participant. Initially, the display area displays the video of the corresponding participant, i.e., the initial video. When the display areas are used to splice and display the target video, the target video is split according to the position of each display area on the user display page. The split videos corresponding to each display area are obtained, and the split videos are used to replace the initial video corresponding to the display area. For example, the split video corresponding to display area 1 replaces the initial video of display area 1, and the split video corresponding to display area 2 replaces the initial video of display area 2.

[0045] At the same time, each display area included in the user display page displays a partial picture of the same video frame of the target video. Each display area can be spliced ​​into a complete picture of the video frame, that is, each display area is used to perform a pseudo full-screen display of the video frame, such as Figure 2a The figure shows a schematic diagram of the page display when no one is speaking. In the figure, number 1 indicates a display area. Figure 2b The figure shows a schematic diagram of the page display when someone is speaking. The speaker is participant A, and the video of participant A is spliced ​​and displayed in four display areas.

[0046] In one embodiment of the present application, the multiple conference participants include a first marked user;

[0047] The replacing the initial video corresponding to each display area with the split video corresponding to each display area comprises:

[0048] In a case where the user display page does not include a display area corresponding to the first marked user, the split video corresponding to each display area is used to replace the initial video corresponding to the display area.

[0049] The first marking user may be set by the user. For example, the user clicks a marking function on a certain display area to mark the participant corresponding to the display area.

[0050] In one embodiment of the present application, the multiple conference participants include a first marked user;

[0051] The step of replacing the initial video corresponding to each display area with the split video corresponding to each display area includes:

[0052] In a case where the user display page includes a display area corresponding to the first marked user, replacing the initial video corresponding to the first display area with the split video corresponding to the first display area, and continuing to display the initial video corresponding to the second display area in the second display area;

[0053] The first display area is an area included in the user display page except the second display area, and the second display area is a display area corresponding to the first marked user.

[0054] The first marking user can be set by the user. For example, the user clicks the marking function of a certain display area to mark the participant corresponding to the display area. The display area of ​​the marked participant (this area becomes the second display area) does not participate in the pseudo-full-screen display. The second display area always displays the video corresponding to the participant. The first display area splices and displays the target video. During the splicing display, the partial picture corresponding to the second display area is default, that is, not displayed.

[0055] For example, if the user display page includes 2 rows and 2 columns, and the participant corresponding to the display area in the 1st row and 1st column is marked, then the display area displays the video of the corresponding participant, and the other three display areas display the target video in a spliced ​​manner. When the spliced ​​display is displayed, the video frame of the target video can be split into four parts, which correspond to 2 rows and 2 columns. The split video in the 1st row and 1st column should originally be displayed in the second display area, but because the second display area in the user display page is the display area corresponding to the first marked user, this position cannot display the part of the 1st row and 1st column of the video frame, so this part will not be displayed.

[0056] In one embodiment of the present application, the multiple conference participants include a first marked user, and the method further includes:

[0057] When the terminal meets a preset condition, displaying the video corresponding to the first marked user in a target window, wherein the target window is displayed in a floating manner on the display interface of the terminal;

[0058] The preset condition includes one of the following:

[0059] The video conference is in background running state;

[0060] The page currently displayed by the terminal is a first page, wherein the video conference includes M pages, the first page is any page of the M pages except the second page, the second page includes a display area corresponding to the first marked user, and M is an integer greater than or equal to 2;

[0061] The terminal receives a page turning operation. For example, the currently displayed page is the second page. The page turning operation triggers the terminal to switch the displayed second page to the first page.

[0062] Specifically, the first marking user can be set by the user. For example, the user clicks on the marking function of a certain display area to mark the participant corresponding to the display area. The marked participant has a higher priority. Even if the terminal is not currently displaying the user display page, the video of the marked participant will still be displayed in the target window, and the target window will be displayed floating on the display interface of the terminal. For example, participant A in page 1 is marked, and the terminal currently displays page 2, or the video conference is running in the background. In this case, the terminal still displays the video of participant A. In this way, the user can mark the participant who needs special attention, so as to quickly pay attention to the participant, simplify user operations, and improve user experience.

[0063] In one embodiment of the present application, after splicing and displaying the target video using the multiple display areas included in the user presentation page of the video conference, the method further includes:

[0064] receiving a first operation on a third display area, where the third display area is any display area included in the user display page;

[0065] In response to the first operation, the video of the participant corresponding to the third display area is displayed in the third display area.

[0066] Specifically, the first operation can be a click operation, a slide operation, etc. For example, if the user clicks the third display area, the terminal displays the video of the corresponding participant in the third display area. For example, the terminal is currently displaying page 1, which includes display areas corresponding to participant A, participant B, participant C, and participant D. The user clicks the display area corresponding to participant B, and the terminal displays participant B's video in the display area.

[0067] Furthermore, after displaying the video of the conference participant corresponding to the third display area in the third display area in response to the first operation, the method further includes:

[0068] receiving a second operation on the third display area;

[0069] In response to the second operation, the target video is spliced ​​and displayed again using the display area included in the user display page.

[0070] Specifically, the second operation may be a click operation, a slide operation, etc. For example, if the user clicks on the display area of ​​participant B again, the display area of ​​participant B will again participate in the splicing display of the target video.

[0071] In the above, the user can view the video of the participant of interest through the first operation, and switch to pseudo full-screen display through the second operation, so that the user can flexibly switch between the pseudo full-screen of the first participant and the video of the participant, simplifying user operations and improving operational efficiency.

[0072] In one embodiment of the present application, if the first participant is the current speaker, after splicing and displaying the target video using the multiple display areas included in the user presentation page of the video conference, the method further includes:

[0073] Receive page turning operation;

[0074] In response to the page turning operation, if the current speaker is the first participant, switching the user display page to a third page, the third page being the page displayed by the terminal based on the page turning operation;

[0075] The target video is spliced ​​and displayed using the display area included in the third page.

[0076] Specifically, when there are a large number of participants in a video conference and one page cannot carry and display the display areas corresponding to all participants, the display areas corresponding to multiple participants need to be displayed in pages, for example, divided into 2 pages or more than 2 pages for display. Users can turn pages through page turning operations, which can be click operations, sliding operations or voice input operations, etc., which are not limited here.

[0077] Based on the page turning operation, the terminal displays the third page, which is the currently displayed page. If the current speaker is the first participant, the target video is spliced ​​and displayed using the display area included in the first page.

[0078] While the first participant is speaking, the user can flip the page currently displayed on the terminal and display a pseudo-fullscreen of the target video on a third page. This means that the target video is spliced ​​and displayed using the display area included in the third page. This allows the user to search for information about other participants while the first participant is speaking. Furthermore, the user can intuitively identify the current speaker as the first participant from the video display on the third page, making it easier for users to access information during video conferences.

[0079] In the above embodiment, when the terminal responds to the page turning operation, the current speaker does not change and remains the first participant. In another case, when the terminal responds to the page turning operation, the current speaker changes from the first participant to the second participant. The specific process is as follows:

[0080] After splicing and displaying the target video using the multiple display areas included in the user display page of the video conference, the method further includes:

[0081] Receive page turning operation;

[0082] In response to the page turning operation, if the current speaker is a second participant, switching the user display page to display a fourth page, where the second participant is a participant other than the first participant among the multiple participants, and the fourth page is a page where a display area corresponding to the second participant is located, or the fourth page is a page displayed by the terminal based on the page turning operation;

[0083] The video of the second participant is spliced ​​and displayed using the display area included in the fourth page.

[0084] Specifically, when there are a large number of participants in a video conference and one page cannot carry and display the display areas corresponding to all participants, the display areas corresponding to these multiple participants need to be displayed in pages, for example, divided into 2 pages or more than 2 pages for display. The user can turn the page through a second operation. The page turning operation can be a click operation, a sliding operation, or a voice input operation, etc., which is not limited here.

[0085] In response to a page-turning operation, in one case, if the current speaker becomes the second participant, the second participant's video can be displayed in pseudo-full screen on the fourth page, which is the page where the second participant's corresponding display area is located. For example, if the second participant's corresponding display area is on page 3, the user display page is page 1, and the terminal's current display page is the user display page, in response to the page-turning operation, the terminal displays page 3 and uses the display area included in page 3 to splice the second participant's video. In this case, the terminal's current display page is determined based on the page where the second participant's corresponding display area is located.

[0086] In another scenario, if the current speaker becomes the second participant, the second participant's video can be displayed in a spliced ​​format on the fourth page, which is the page displayed based on the page turning operation. For example, the user display page is page 1, the terminal's current display page is the user display page, and the page displayed based on the page turning operation is page 2. In response to the page turning operation, the terminal displays page 2 and uses the display area included in page 2 to splice the second participant's video. In this case, the terminal's current display page is determined based on the page turning operation.

[0087] In the above embodiment, when the current speaker changes, the video of the current speaker is always displayed in pseudo-full screen in the current display page of the terminal. Participants can intuitively know who the current speaker is through the current display page of the terminal, thereby improving the convenience of information acquisition.

[0088] During a video conference, a large number of users may have their cameras turned on. Mobile devices have limited page capacity, accommodating up to four users. If a single page cannot display all participants, the page will be split into pages. In a multi-person meeting, this can result in multiple, or even dozens, pages.

[0089] Since user behavior is unpredictable, we don't know which page the user will turn to or which participants (also called attendees) to view. It is possible that during a meeting, due to the large number of people, it is impossible to find the core speaker immediately. The video conference display method provided in this application can help users quickly know the current speaker.

[0090] The following describes the video conference display method provided by this application by taking a mobile phone as an example.

[0091] The process of the video conference display method is as follows:

[0092] In video conferencing, the screen size of a mobile phone is limited. Assume that the screen can display the display area of ​​four users at a time, such as Figure 2a When there are more than four participants, a second page will appear. The structure of this page is the same as the first page, for example, both are displayed in 2 rows and 2 columns. Figure 2c As shown in the figure, the fifth and sixth users can be switched between multiple pages by sliding. In this example, each user has turned on the camera. The terminal can determine the current page based on the sliding event. The page is tentatively set as page x, and x starts at 0. The number of users on this page is x multiplied by 4 to Math.min(total number of users - 1, (x * 4) + 3). The terminal pre-acquires the order of multiple participating users from the server, for example, the order in which multiple participating users access the video conference. The order of each participating user can be determined based on the participating user's identifier (peerID). For example, the participating user with identifier 1 is ranked first, the participating user with identifier 2 is ranked second, and so on.

[0093] There are two display schemes: custom and automatic.

[0094] Solution 1: Automatically move the page to highlight the speaker

[0095] Assume that there are 20 people currently participating and one page can carry the screens of four users. Therefore, the video conference consists of five pages. If the user slides the page to page 3, but a user on the first page starts speaking at this time, the terminal monitors the speaking user and obtains the user's peerID. Through calculation, it can obtain the page the current user is on and the various statuses of each user on the page.

[0096] When the user slides to a certain page, it is determined whether there is a speaker on the current page. If there is a speaker on the current page, the speaker's video is spliced ​​and displayed in the display area included in the current page. That is, the display area included in the current page plays part of the video frame in the video stream respectively, and the whole splicing is a complete face, such as Figure 2b shown.

[0097] Figure 2a After the first user speaks, the application layer terminal receives the audio packet and splices the video of the first user in the four display areas of the page.

[0098] This process can be divided into two cases. Case 1: When a participant (also called a user) speaks on the page, the speaker's video is spliced ​​and displayed using the display area included in the page; if more than two people speak, the display is switched in the order of speaking. If no participant speaks and the terminal's business end does not receive any audio packets, the normal four-person display will be restored; the user slides the page, and the speaker's video is spliced ​​and displayed on each page;

[0099] In the second case, when the display area corresponding to the speaker is not in the page currently displayed by the terminal, the terminal automatically jumps to the page where the display area corresponding to the speaker is located, and uses the page to splice and display the speaker's video core.

[0100] The technical implementation of this solution is as follows:

[0101] Step 1: When a user logs in to a video conference, the server stores the user's peerID in the local database. This peerID belongs only to this user, and a list is generated for this peerID, which contains attributes such as whether the user has turned on the camera and microphone. Each user's peerID and related information will be included in this list.

[0102] Step 2: At the terminal application layer, the video of the conference participants is displayed in a predetermined manner, with a maximum of four display areas displayed on one page. If there are multiple conference participants (more than four), the video images of the conference participants are displayed on the terminal user interface in a predetermined manner, with four display areas displayed on each page, according to the first solution (automatically moving the key display person).

[0103] Step 3: When the terminal application layer monitors the page turning operation, the number of users on the page displayed after the page turning is calculated as x multiplied by 4 to Math.min(total number of people - 1, (x*4)+3), where x represents the current page, and the page starts from page 0. This formula is used to obtain the users on this page in the list, and at the same time obtain the status of the microphone and camera.

[0104] When someone speaks, the terminal can receive the user with the peer ID who is speaking through the signaling sent by the server, and splice the obtained video stream into each window, displaying a part of it respectively, hiding the previous video streams of other users. This effect is achieved by refreshing the data.

[0105] Solution 2: Customize the display of key people

[0106] Users can set a star function for a specific participant. The star function has two functions. First, if a user sets a star for a user on a certain page, the user cannot be replaced by speakers on other pages on that page. Second, even if the user turns the page, the interface of the starred participant will become suspended and follow the page turning. Or, if the video conference switches to the background, the video of the starred participant will always be displayed in a suspended state on the terminal display interface.

[0107] Step 1: The star in the display area of ​​the starred participant is lit, the peerID user is marked in the user list, and an importance attribute is added to the user. Users with this attribute are not included in the speaker video splicing, that is, the user's video is not hidden but displayed in the display area.

[0108] Then, the list of other users who spoke on the page is retrieved to see whether the video stream can be spliced ​​and presented. If so (the other users who spoke are not star users), the video of the speaking user will be presented in a "pseudo-full screen mode" in other windows on the current page except the window of the star user.

[0109] Step 2: If you turn to other pages, the terminal can start a foreground service and embed the video stream of the star user in the service. In this way, when the user turns the page or switches to the background, the video of the star user will always be displayed on the user interface, allowing the user to always pay attention to the participants they selected.

[0110] Step 3: Monitor the page through the terminal application layer to determine whether the currently displayed page is the page where the display area corresponding to the star user is located. If so, the target display area disappears; if not, the target display area appears.

[0111] The solution in this application can help users better locate the speaker in a multi-person multi-video conference, thereby improving convenience and user experience.

[0112] See also Figure 3 , Figure 3 FIG. 1 is a structural diagram of a video conferencing display device provided by an embodiment of the present invention. Figure 3 As shown, the video conference display device 300 includes:

[0113] An acquisition module 301 is configured to acquire a target video of a first participant in a video conference, where the video conference includes multiple participants, and the first participant is one of the multiple participants;

[0114] The first display module 302 is configured to display the target video in a spliced ​​manner using a plurality of display areas included in a user display page of the video conference, where each display area in the user display page corresponds to a participant.

[0115] Optionally, the first participant is a speaker in the video conference or the first participant is a first marked user.

[0116] Optionally, the first display module 302 includes:

[0117] An acquisition submodule is configured to split the target video according to the position of each display area on the user display page, and obtain a split video corresponding to each display area;

[0118] The replacement submodule is used to replace the initial video of the corresponding display area with the split video corresponding to each display area, where the initial video of the display area is the video of the participant corresponding to the display area.

[0119] Optionally, the multiple conference participants include a first marked user;

[0120] The replacement submodule is used to replace the initial video of the corresponding display area with the split video corresponding to each display area when the user display page does not include the display area corresponding to the first marked user.

[0121] Optionally, the multiple conference participants include a first marked user;

[0122] The replacement submodule is configured to, when the user display page includes a display area corresponding to the first marked user, replace the initial video corresponding to the first display area with the split video corresponding to the first display area, and continue to display the initial video corresponding to the second display area in the second display area;

[0123] The first display area is an area included in the user display page except the second display area, and the second display area is a display area corresponding to the first marked user.

[0124] Optionally, the device further includes:

[0125] A second display module is configured to display the video corresponding to the first marked user in a target window when the terminal meets a preset condition, and the target window is displayed in a floating manner on the display interface of the terminal;

[0126] The preset condition includes one of the following:

[0127] The video conference is in background running state;

[0128] The page currently displayed by the terminal is a first page, wherein the video conference includes M pages, the first page is any page of the M pages except the second page, the second page includes a display area corresponding to the first marked user, and M is an integer greater than or equal to 2;

[0129] The terminal receives a page turning operation.

[0130] Optionally, the device further includes:

[0131] a receiving module, configured to receive a first operation on a third display area, where the third display area is any display area included in the user display page;

[0132] The response module is used to respond to the first operation and display the video of the participant corresponding to the third display area in the third display area.

[0133] The video conference display device 300 can achieve Figure 1 The various processes implemented in the illustrated method embodiments achieve the same beneficial effects, and to avoid repetition, they will not be described again here.

[0134] Figure 4 A schematic diagram of the hardware structure of a terminal provided for implementing an embodiment of the present invention is shown as follows: Figure 4 As shown, the terminal 900 includes but is not limited to: a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, a processor 910, and a power supply 911. Those skilled in the art will understand that Figure 4 The terminal structure shown in the figure does not constitute a limitation of the terminal. The terminal may include more or fewer components than shown, or combine certain components, or arrange the components differently. In the embodiments of the present invention, the terminal includes but is not limited to a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle terminal, a wearable device, and a pedometer.

[0135] The radio frequency unit 901 is configured to obtain a target video of a first participant in a video conference, where the video conference includes multiple participants and the first participant is one of the multiple participants;

[0136] The display unit 906 is configured to display the target video in a spliced ​​manner using a plurality of display areas included in the user display page of the video conference, where each display area in the user display page corresponds to a participant.

[0137] Optionally, the first participant is a speaker in the video conference or the first participant is a first marked user.

[0138] Optionally, the processor 910 is configured to split the target video according to the position of each display area on the user display page to obtain a split video corresponding to each display area;

[0139] The split video corresponding to each display area is used to replace the initial video of the corresponding display area, where the initial video of the display area is the video of the participant corresponding to the display area.

[0140] The multiple conference participants include a first marked user;

[0141] The replacing the initial video of each display area with the split video corresponding to the display area includes:

[0142] In a case where the user display page does not include the display area corresponding to the first marked user, the split video corresponding to each display area is used to replace the initial video of the corresponding display area.

[0143] Optionally, the multiple conference participants include a first marked user;

[0144] Processor 910 is configured to, when the user display page includes a display area corresponding to the first marked user, replace an initial video corresponding to the first display area with a split video corresponding to the first display area, and continue to display the initial video corresponding to the second display area in the second display area;

[0145] The first display area is an area included in the user display page except the second display area, and the second display area is a display area corresponding to the first marked user.

[0146] Optionally, the multiple participants include a first marked user, and the display unit 906 is configured to display a video corresponding to the first marked user in a target window when the terminal meets a preset condition, and the target window is displayed in a floating manner on the display interface of the terminal;

[0147] The preset condition includes one of the following:

[0148] The video conference is in background running state;

[0149] The page currently displayed by the terminal is a first page, wherein the video conference includes M pages, the first page is any page of the M pages except the second page, the second page includes a display area corresponding to the first marked user, and M is an integer greater than or equal to 2;

[0150] The terminal receives a page turning operation.

[0151] Optionally, the user input unit 907 is configured to receive a first operation on a third display area, where the third display area is any display area included in the user display page;

[0152] The display unit 906 is configured to display, in response to the first operation, the video of the conference participant corresponding to the third display area in the third display area.

[0153] The terminal 900 can implement each process implemented by the terminal in the aforementioned embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.

[0154] It should be understood that in this embodiment of the present invention, the RF unit 901 can be used to receive and transmit signals during information transmission or calls. Specifically, it receives downlink data from the base station and transmits it to the processor 910 for processing; in addition, it transmits uplink data to the base station. Generally, the RF unit 901 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like. Furthermore, the RF unit 901 can communicate with the network and other devices via a wireless communication system.

[0155] The terminal provides users with wireless broadband Internet access through the network module 902, such as helping users to send and receive emails, browse web pages, and access streaming media.

[0156] The audio output unit 903 can convert audio data received by the RF unit 901 or the network module 902 or stored in the memory 909 into an audio signal and output it as sound. In addition, the audio output unit 903 can also provide audio output related to specific functions performed by the terminal 900 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 903 includes a speaker, a buzzer, a receiver, etc.

[0157] The input unit 904 is used to receive audio or video signals. The input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes image data of a still picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frames can be displayed on the display unit 906. The image frames processed by the graphics processor 9041 can be stored in the memory 909 (or other storage medium) or transmitted via the radio frequency unit 901 or the network module 902. The microphone 9042 can receive sound and can process such sound into audio data. The processed audio data can be converted into a format that can be sent to a mobile communication base station via the radio frequency unit 901 in the case of a telephone call mode.

[0158] The terminal 900 also includes at least one sensor 905, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 9061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 9061 and / or the backlight when the terminal 900 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used to identify the terminal posture (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 905 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.

[0159] The display unit 906 is used to display information input by the user or information provided to the user. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0160] The user input unit 907 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the terminal. Specifically, the user input unit 907 includes a touch panel 9071 and other input devices 9072. The touch panel 9071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel 9071). The touch panel 9071 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal caused by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into contact point coordinates, which are then sent to the processor 910, which receives the command sent by the processor 910 and executes it. In addition, the touch panel 9071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 9071, the user input unit 907 may also include other input devices 9072. Specifically, other input devices 9072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0161] Furthermore, the touch panel 9071 may be overlaid on the display panel 9061. When the touch panel 9071 detects a touch operation on or near it, it transmits the information to the processor 910 to determine the type of touch event. Subsequently, the processor 910 provides corresponding visual output on the display panel 9061 according to the type of touch event. Figure 4 In the embodiment, the touch panel 9071 and the display panel 9061 are two independent components to realize the input and output functions of the terminal. However, in some embodiments, the touch panel 9071 and the display panel 9061 can be integrated to realize the input and output functions of the terminal, which is not limited here.

[0162] The interface unit 908 is an interface for connecting external devices to the terminal 900. For example, the external devices may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 908 may be used to receive input (e.g., data information, power, etc.) from the external device and transmit the received input to one or more components within the terminal 900, or may be used to transmit data between the terminal 900 and the external device.

[0163] Memory 909 can be used to store software programs and various data. Memory 909 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, memory 909 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0164] Processor 910 is the terminal's control center, connecting all components of the terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 909 and accessing data stored in memory 909, it executes various terminal functions and processes data, thereby providing overall terminal monitoring. Processor 910 may include one or more processing units; preferably, processor 910 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 910.

[0165] The terminal 900 may also include a power supply 911 (such as a battery) for supplying power to various components. Preferably, the power supply 911 may be logically connected to the processor 910 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0166] In addition, the terminal 900 includes some functional modules not shown, which will not be described in detail here.

[0167] like Figure 5 As shown, the embodiment of the present invention further provides a terminal 1000, including a processor 1002, a memory 1001, a computer program stored in the memory 1001 and capable of running on the processor 1002, and the computer program is executed by the processor 1002 to achieve the above Figure 1 The various processes of the illustrated embodiments can achieve the same technical effects, and will not be described again here to avoid repetition.

[0168] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, which, when executed by a processor, implements the above Figure 1The various processes of the video conferencing display method embodiment shown in the figure can achieve the same technical effect. To avoid repetition, they are not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0169] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0171] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A video conference display method, applied to a terminal, characterized in that: include: Obtaining a target video of a first participant in a video conference, where the video conference includes multiple participants, the first participant being a speaker in the video conference or the first participant being a first marked user, where the first marked user is a participant among the multiple participants who is marked based on user settings; Utilizing a plurality of display areas included in a user display page of the video conference, the target video is spliced ​​and displayed, each display area in the user display page corresponding to a participant; After splicing and displaying the target video using the multiple display areas included in the user display page of the video conference, the method further includes: Receiving a first operation on a third display area, where the third display area is any display area included in the user display page, wherein the first operation includes a click operation and a slide operation; In response to the first operation, the video of the participant corresponding to the third display area is displayed in the third display area.

2. The method according to claim 1, characterized in that The method of utilizing the multiple display areas included in the user display page of the video conference to splice and display the target video includes: Splitting the target video according to the position of each display area on the user display page to obtain split videos corresponding to each display area; The split video corresponding to each display area is used to replace the initial video of the corresponding display area, where the initial video of the display area is the video of the participant corresponding to the display area.

3. The method according to claim 2, characterized in that The step of replacing the initial video of each display area with the split video corresponding to the display area includes: When the first participant is the first marked user and the user display page does not include the display area corresponding to the first marked user, the split video corresponding to each display area is used to replace the initial video of the corresponding display area.

4. The method according to claim 2, characterized in that The step of replacing the initial video of each display area with the split video corresponding to the display area includes: When the first participant is a first marked user and the user display page includes a display area corresponding to the first marked user, the initial video corresponding to the first display area is replaced with the split video corresponding to the first display area, and the initial video corresponding to the second display area continues to be displayed in the second display area; The first display area is an area included in the user display page except the second display area, and the second display area is a display area corresponding to the first marked user.

5. The method according to claim 1, wherein The method further comprises: If the first participant is a first marked user and the terminal meets a preset condition, displaying the video corresponding to the first marked user in a target window, and the target window is displayed in a floating manner on the display interface of the terminal; The preset condition includes one of the following: The video conference is in background running state; The page currently displayed by the terminal is a first page, wherein the video conference includes M pages, the first page is any page of the M pages except the second page, the second page includes a display area corresponding to the first marked user, and M is an integer greater than or equal to 2; The terminal receives a page turning operation.

6. A video conference display device, characterized in that: include: an acquisition module, configured to acquire a target video of a first participant in a video conference, the video conference including multiple participants, the first participant being a speaker in the video conference or the first participant being a first marked user, the first marked user being a participant among the multiple participants who is marked based on user settings; A first display module is configured to display the target video in a spliced ​​manner using a plurality of display areas included in a user display page of the video conference, wherein each display area in the user display page corresponds to a participant; a receiving module configured to receive a first operation on a third display area after the first display module displays the target video in a spliced ​​manner using multiple display areas included in the user display page of the video conference, the third display area being any one of the display areas included in the user display page, wherein the first operation includes a click operation and a slide operation; The response module is used to respond to the first operation and display the video of the participant corresponding to the third display area in the third display area.

7. A terminal, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the video conference display method according to any one of claims 1 to 5 when executed by the processor.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the video conference display method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Information display method and device and electronic equipment

    CN112312224A