Video conference system and picture presentation method

By using multiple cameras and microphone arrays in the video conferencing system and combining with the conference host to determine the position of the spokesperson, the automatic cropping and display of the picture is achieved, solving the problem of unclear and frequent switching of the picture during discussion among multiple people, and improving the automation and interactivity of the system.

CN120343196APending Publication Date: 2025-07-18YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510616207.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video conferencing system is difficult to ensure clear picture presentation and effective speaker recognition during multi-person discussions, resulting in frequent picture switching, affecting the fluency and interactivity of the meeting.

Method used

Multiple cameras are used to divide the conference venues, combine the microphone array to collect speech information and position information, judge the conference area where the speaker's location belongs to through the conference host, and crop the screen information and send it to the close-up screen of the display.

Benefits of technology

It has achieved the improvement of the automation and fluency of the video conferencing system, and can accurately capture and display spokespersons, reduce screen switching, and improve the interactivity and viewing experience of the conference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343196A_ABST
    Figure CN120343196A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent conference systems, in particular to a video conference system and a picture presentation method, a conference site is divided into a corresponding number of conference areas according to the number of cameras, and the video conference system further comprises a plurality of cameras used for shooting picture information of the corresponding conference areas; the microphone array is used for collecting speaking information and position information of a spokesman; the conference host is used for judging a conference area to which the position of the spokesman belongs according to the speaking information and the position information to obtain judgment result information; and cutting the picture information of the conference area according to the judgment result information to obtain close-up picture information sent to the display screen, so that the video conference system can automatically capture a specific picture, and the automation degree and the smoothness of the video conference system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent conference systems, and particularly to a video conference system and a method for presenting images. Background Art

[0002] In the current working environment, especially in the conference scenario, remote collaboration and video conferencing have become increasingly common. With the continuous development of technology, video conference systems have gradually become an important tool for enterprises, organizations, and teams to communicate and make decisions. Most modern conference systems usually adopt a single camera or a simple intercom mode for video display and speaker capture. As the scale of the conference expands and the number of participants increases, the single camera and basic intercom mode often have many limitations during multi-person discussions. When multiple participants speak simultaneously, it is difficult for traditional systems to ensure clear image presentation and effective speaker recognition, resulting in frequent image switching and affecting the fluency and viewing experience of the conference.

[0003] The current conference systems can only display the latest speaker or the images of a limited number of participants. For multiple parties involved in the discussion, it is impossible to obtain an all-round perspective. In a complex multi-person discussion environment, the speeches of the participants alternate frequently, resulting in the inability to accurately capture and display the images of each speaker, thus affecting the interactivity and communication effect of the conference. The above problems need to be solved. Summary of the Invention

[0004] In order to enable the video conference system to automatically capture specific images and improve the automation degree and fluency of the video conference system, the present application provides a video conference system and a method for presenting images, and adopts the following technical solutions:

[0005] In a first aspect, the present application provides a video conference system. The conference venue is divided into corresponding numbers of conference areas according to the number of cameras, and further includes:

[0006] A plurality of cameras for shooting the image information of the corresponding conference areas;

[0007] A microphone array for collecting the speech information and position information of the speaker;

[0008] A conference host for judging the conference area to which the speaker's position belongs according to the speech information and position information to obtain judgment result information; and cropping the image information of the conference area according to the judgment result information to obtain the close-up image information for sending to the display screen.

[0009] Preferably, it further includes:

[0010] A data input device is used to input information about the size of the meeting venue, the position information of the cameras, and the camera angle information, and allocate the corresponding meeting area picture information for the cameras according to the meeting venue size information, the camera coordinate information, and the camera angle information.

[0011] Preferably, the data input device is further used to input meeting type information, the number of cameras is greater than or equal to 2, and the camera height is a first threshold value.

[0012] Preferably, when the meeting type information is a horizontal meeting, there is a field of view overlap between two adjacent cameras; when the meeting type information is a non-horizontal meeting, the field of view of the left camera covers at least all the portraits on the right half of the meeting room, and the field of view of the right camera covers at least all the portraits on the left half of the meeting room.

[0013] Preferably, each camera and the microphone array are an integrated device, and the integrated device has at least one camera and one microphone array.

[0014] In a second aspect, the present application provides a picture presentation method, including:

[0015] Judging the meeting area to which the speaker's position belongs according to the speech information and the position information, and obtaining judgment result information;

[0016] Cropping the meeting area picture information according to the judgment result information to obtain the close-up picture information for sending to the display screen.

[0017] Preferably, it further includes:

[0018] Obtaining the number of participants information, and when the number of participants is greater than a first preset value, outputting the first initial picture information for displaying the first preset value of the number.

[0019] Preferably, it further includes:

[0020] Obtaining the number of participants information, and when the number of participants is less than the first preset value, outputting the second initial picture information for displaying the number of participants.

[0021] Preferably, it further includes:

[0022] When the number of speakers in the close-up picture information is zero and the number of participants information is greater than zero, adjusting the close-up picture information according to the first initial picture information or the second initial picture information.

[0023] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is set to execute a picture presentation method as described above when running.

[0024] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by the present application at least include:

[0025] In the present application, the meeting venue is divided into meeting areas by cameras. The cameras capture the picture information of the meeting areas and send it to the meeting host. The microphone array collects the speech information and position information of the speakers and sends it to the meeting host. The meeting host analyzes the real-time information obtained, determines which meeting area captured by which camera the speaker's position belongs to, crops the picture information of the corresponding meeting area captured by the camera according to the judgment result to obtain close-up picture information, and transmits the close-up picture information to the display screen for real-time display, so that the video conferencing system can automatically capture specific pictures, improving the automation degree and fluency of the video conferencing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the modules of a video conferencing system according to an embodiment of the present application.

[0027] Figure 2 It is a schematic diagram of the structure when the meeting type information is a horizontal meeting according to an embodiment of the present application.

[0028] Figure 3 It is a schematic diagram of the structure when the meeting type information is a non-horizontal meeting according to an embodiment of the present application.

[0029] Figure 4 It is a schematic diagram of the flow of a method for presenting pictures according to an embodiment of the present application.

[0030] Description of the reference numerals:

[0031] 1. Camera; 2. Microphone array; 3. Meeting host. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following is a further detailed description of the present application in conjunction with Figures 1-4 The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting.

[0033] Referring to Figure 1 , a video conferencing system involved in the present application divides the meeting venue into corresponding numbers of meeting areas according to the number of cameras, and further includes:

[0034] A plurality of cameras for capturing the picture information of the corresponding meeting areas;

[0035] A microphone array for collecting the speech information and position information of the speakers;

[0036] A conference host is used to determine the conference area to which the speaker's position belongs based on the speech information and location information, and obtain the judgment result information; and crop the conference area screen information according to the judgment result information to obtain the close-up screen information to be sent to the display screen.

[0037] Specifically, in the multi-camera tracking solution of this application, multiple cameras are set, and different capture areas are specified for each camera. At the beginning of the meeting, several participants who appear first are displayed, and a window is provided below the screen to display the panoramic view of the meeting room. When someone is speaking, the information obtained by multiple cameras is used to determine the position of the speaker, determine which camera's capture area the speaker is in, and then further crop the accurate portrait and replace the portrait of the speaker who has not spoken for the longest time in the current screen. This application can display the close-up screens of several recent speakers at the same time, so the problem of multiple switches of a single screen can be avoided in a multi-person discussion scenario.

[0038] In this application, the conference venue is divided into conference areas by cameras. The conference area screen information captured by the cameras is sent to the conference host, and the speech information and location information of the speaker are collected by the microphone array and sent to the conference host. The conference host analyzes the real-time information obtained, determines which conference area the speaker's position belongs to the one captured by which camera, crops the conference area screen information captured by the corresponding camera according to the judgment result to obtain the close-up screen information, and transmits the close-up screen information to the display screen for real-time display, so that the video conferencing system can automatically capture specific screens, improving the automation degree and fluency of the video conferencing system.

[0039] Among them, the cameras mainly do not process the conference area screen information. Each conference area is matched with a corresponding number of cameras to comprehensively shoot the scenes within their respective areas. The cameras provide real-time video stream data for subsequent processing and display. Through the coverage of multiple cameras, the situation of each conference area can be comprehensively displayed, ensuring that no speaker's screen is missed. The screen information provides materials for subsequent screen cropping and close-up display.

[0040] The speech information includes the audio data of the speaker; through the collaborative work of multiple microphones, the microphone array can accurately identify the direction and source of the sound. And through the analysis of the audio data, the microphone array can determine the position of the speaker. Cooperating with the positioning technology, it can track the position of the speaker in real time and further provide the information of the conference area where the speaker is located for the conference host.

[0041] Based on the speech information and location information obtained from the microphone array, the conference host determines the specific conference area where the current speaker is located. The host combines this information with the video data of each conference area to decide which pictures need to be cropped and displayed. Through the real-time processing of audio and video, it is possible to automatically switch and crop the displayed pictures according to the speaker's position and voice information. For example, when a certain speaker starts speaking, the host will automatically switch to a close-up picture of the conference area where the speaker is located and send it to the display screen. The automated operation improves the interactivity of the conference and the viewing experience of the audience.

[0042] As one of the implementation manners, it further includes:

[0043] A data input device, which is used to input the conference venue size information, camera position information, and camera angle information, and allocate the picture information corresponding to the camera for each conference area according to the conference venue size information, camera coordinate information, and camera angle information.

[0044] Specifically, the data input device in the embodiment of the present application receives the specific size information of the conference venue, such as length, width, height, etc. By receiving this data, it plans the coverage range and layout of the cameras, and then clarifies the size and structure of the venue. Only in this way can the system reasonably arrange and match the installation positions and angles corresponding to the cameras, so that each area can be photographed.

[0045] The camera coordinate information indicates the specific position of each camera, which is the key data for the camera to cover the correct area. For example, if the camera is located in a corner of the conference venue or on a specific hanger, the coordinate information helps the system determine the precise position of each camera, thereby affecting the cropping and display of the picture information.

[0046] The camera angle information is the shooting direction or shooting range of each camera. The lens angle of each camera determines which area can be photographed and the size of the photographed picture. By accurately inputting the angle information of each camera, the system can better allocate the picture area of each conference area, so that the perspectives of different cameras do not overlap or miss shooting.

[0047] The height of the camera should be higher than the desktop. The specific distance above the desktop is determined according to the average height of the items on the desktop and the lens protection requirements, usually ranging from 5 cm to 20 cm, to avoid the items blocking the shooting perspective. Secondly, it is necessary to ensure that the shooting object is within the optimal longitudinal viewing angle of the camera to avoid the influence of camera distortion; therefore, it is necessary to ensure that the shooting object is within the optimal longitudinal viewing angle both when standing and sitting. Based on the longitudinal viewing angle characteristics of the camera, with the center point of the lens as the vertex, the visible area expands outward at an angle in the vertical direction. To ensure that the shooting object can be within the optimal longitudinal viewing angle in different postures such as standing and sitting, according to the trigonometric function relationship and the size of the meeting room, and considering that the visual distortion of looking up is much greater than that of looking down, there is a recommended installation height range. In some embodiments, the optimal recommended deployment range is obtained as the minimum height greater than 1 m and the maximum height less than 2.5 m.

[0048] Combined with the size of the meeting venue, the coordinates and angles of the cameras, the data input device helps the system reasonably allocate the meeting areas responsible for each camera, so that each meeting area has at least one camera within its field of view, and automatically determines the boundaries of the areas it shoots according to the shooting angles and positions of the cameras.

[0049] As one of the implementation manners, the data input device is also used to input meeting type information. When the meeting type information is a horizontal meeting, the number of cameras is greater than or equal to 2, the minimum height of the cameras is greater than 1 m, the maximum height is less than 2.5 m, and there is an overlap in the fields of view of two adjacent cameras. In this example, the preferred distance between the cameras is greater than 3 m.

[0050] As one of the implementation manners, when the meeting type information is a non-horizontal meeting, the field of view of the left camera covers at least all the portraits on the right half of the meeting room, and the field of view of the right camera covers at least all the portraits on the left half of the meeting room.

[0051] Refer to Figure 2 , specifically, the embodiments of the present application process the meeting type with a horizontal layout. A meeting room with a horizontal layout requires at least two cameras. A horizontal meeting is specifically a meeting layout in which the participants are arranged horizontally at intervals in a straight line or curve arrangement. The device deployment height is 1 m - 2.5 m, the distance between the devices should be more than 3 m, and the device deployment height is selected according to the personnel positions. The ratio of the distance between the person and the device to the height of the device from the ground should avoid exceeding 4:3. That is, if the person is 2 m away from the device, the height of the device from the ground should not be lower than 1.5 m as much as possible to avoid too low shooting angles of the camera, resulting in limited dynamic tracking of the camera. The interval between the meeting room seats should be greater than 0.5 m, and 1 m is optimal. Too close seat distances may cause offsets when tracking audio or overlapping of the portrait frames of the speakers.

[0052] Among them, horizontal meetings usually require multiple cameras for comprehensive coverage to ensure that no matter where the participants are, the cameras can capture the speaker's image in real time. Multiple cameras are set at intervals to complement each other and reduce the loss of important details caused by the limited angle of a single camera.

[0053] The minimum height of the camera is greater than 1m to avoid the portrait being blocked by the table and the portraits in front; the maximum height is less than 2.5m to avoid an obvious downward viewing angle during shooting, which may cause facial distortion in the portrait image. To ensure that the camera can capture the expressions and movements of the participants from an appropriate angle, too low an angle will affect the picture quality, while too high an angle will prevent the camera from capturing details or even cause visual errors. Also, it is set to match the general height of people and the height of the seats.

[0054] The cameras are kept 3m apart, which can reduce the overlap of the shooting angles of different cameras, improve the clarity of the image and the overall picture coverage. Keeping a large distance can also avoid interference and visual overlap between multiple cameras, reducing misjudgment and picture distortion.

[0055] To ensure a reasonable shooting angle for the camera, the ratio of the height of the device from the ground to the distance from the person should not exceed 4:3. If the height of the camera from the ground is too low or the distance between the participant and the camera is too close, the camera may not be able to effectively track the speaker's movements, or even lose the speaker's target. Therefore, a certain ratio should be maintained between the height of the device from the ground and the distance from the person to avoid a decline in picture quality.

[0056] The seat spacing affects the tracking of audio and the framing of portraits. If the seat spacing is set too close, it will increase the tracking difficulty of the camera, especially during audio tracking, resulting in confusion or misidentification of the sound source. Also, when the seats are too close, the portrait frames of multiple participants overlap, affecting the accuracy of image recognition and tracking. Therefore, the seat spacing should be appropriate and maintain a certain distance to ensure that the portrait frames of each speaker in the picture are clear and independent.

[0057] As one of the implementation methods, the data input device is also used to input meeting type information. When the meeting type information is a non-horizontal meeting, such as a vertical meeting room, the number of cameras is greater than 3, the height of the cameras is greater than 2m, and the first cameras on both sides are set facing each other at the midpoint of the meeting venue, preferably at a 45° angle.

[0058] Refer to Figure 3, specifically, the embodiments of the present application process the meeting type of a conventional round table. Usually, participants are seated on three sides of a rectangular table. A meeting room with this layout requires at least three cameras, which are respectively deployed at the left, middle, and right positions on the side without people. The deployment height of the cameras is 2m or more. The cameras on both sides are deployed at a 45° angle opposite to the midpoint of the coverage range and exceed the first participant seat in the coverage range to obtain a better portrait shooting angle. The position indicated by the arrow of the middle camera is the recommended deployment position of the right camera in the meeting room, and the camera deployment position can be adjusted by ±15° according to the actual situation.

[0059] Among them, setting at least three cameras is to ensure that all participants can be comprehensively covered. Especially in a traditional round-table meeting, multiple cameras can provide a wider perspective and avoid blind spots in the perspective caused by relying on a single camera.

[0060] Placing the cameras at the 45° angle opposite positions on both sides is to avoid a single and limited perspective when shooting directly at the meeting table in the traditional way. The 45° angle deployment can cover a larger shooting range, increase the shooting of the sides and inclined planes, and help show the interaction status of the participants. By setting it before the first participant seat in the coverage range, the camera enables the shooting range to cover more key participants and avoid perspective dead ends.

[0061] As one of the implementation manners, each camera and the microphone array are an integrated device, and the integrated device has at least one camera and one microphone array.

[0062] As one of the implementation manners, the number of people in the picture information of the meeting area is less than or equal to 6 people.

[0063] Specifically, the camera in the embodiments of the present application can shoot images with a maximum resolution of 4K, and the panoramic horizontal field of view angle is 110°, and the vertical field of view angle is 78°.

[0064] The maximum number of participants that each camera can cover is 6 people. If a camera covers more than 6 participants, some randomly selected participants may not be framed, and the number of deployed cameras needs to be increased. For example, there are five cameras, one at one end of the meeting table to shoot the participants at the other end, and there are more than 6 and less than 13 participants on both sides of the meeting table. Two cameras are set on each side, and each camera shoots at most 6 people on the other side of the meeting table.

[0065] Refer to Figure 4 , a method for presenting a picture is provided for the embodiments of the present application, including:

[0066] Judging the meeting area to which the speaker's position belongs according to the speech information and the position information to obtain judgment result information;

[0067] Crop the meeting area screen information according to the judgment result information to obtain the close-up screen information for sending to the display screen.

[0068] Specifically, the screen presentation method of the present application is configured in the conference host. By obtaining the information collected by the camera and the microphone array for judgment and analysis to obtain the judgment result information, when there is a speaker speaking, the meeting area screen information is cropped, so as to obtain the close-up screen information of the speaker in close-up, and then transmitted to the display screen for real-time playback.

[0069] In the embodiment of the present application, during the process of the meeting, when a participant makes a speech, at this time, it is necessary to crop the meeting area screen information. In the embodiment of the present application, the meeting room is divided into multiple areas, each area corresponds to a camera, a regional coordinate mapping table is established, the boundary coordinates of each area are stored, the audio signal is captured by the microphone array of the camera, the sound source angle is calculated using techniques such as beamforming, combined with the coordinate information filled in during deployment, the precise position of the speaker is determined, the area to which the speaker belongs is determined according to the speaker coordinates, the video stream of the corresponding camera is called, and cropping is performed using methods such as the OpenCV library, so that the speaker is centered and displayed, and the cropped close-up screen is output to the display screen.

[0070] As one of the implementation manners, it further includes:

[0071] Obtain the number of participants information. When the number of participants is greater than the first preset value, output the first initial screen information for displaying the number of the first preset value.

[0072] Specifically, the embodiment of the present application uses a face detection algorithm to count the number of participants in real time. When the meeting room is empty, the panoramic screen of the panoramic camera's fixed-focus lens is displayed. When there are people in the meeting room, set the first preset value to 4 for example. When it is detected that the number of people is greater than 4, display the pictures of the first 4 people recognized, using a four-screen layout, and each screen displays a participant.

[0073] As one of the implementation manners, it further includes:

[0074] Obtain the number of participants information. When the number of participants is less than the first preset value, output the second initial screen information for displaying the number of participants.

[0075] Specifically, the embodiment of the present application uses a face detection algorithm to count the number of participants in real time. When the meeting room is empty, the panoramic screen of the panoramic camera's fixed-focus lens is displayed. When there are people in the meeting room, when it is detected that the number of people ≤ 4, adjust the display layout according to the actual number of people, display the pictures of the first four participants who appear, and display single-screen, two-screen or three-screen pictures when there are less than four people.

[0076] As one of the implementation manners, it further includes:

[0077] When the number of speakers in the close-up screen information is zero and the number of participants information is greater than zero, the close-up screen information is adjusted according to the first initial screen information or the second initial screen information.

[0078] Specifically, in the embodiment of the present application, it is judged whether there is a speaker currently through audio signal analysis, and a speech detection threshold is set, such as continuous silence for 2 seconds. When there is no speaker, the current screen layout is maintained, and the initial screen mode is selected according to the number of participants, such as the first or second initial screen. When it is detected that there is no speaker, the initial screen generation logic is called to regenerate the screen layout according to the current number of people.

[0079] In the present application, when someone is speaking, multiple devices capture audio signals through the built-in microphone array to obtain the sound source angle of the speaker, then confirm the coordinates of the speaker together according to the coordinate information filled in during deployment, and crop and display a suitable portrait screen from the camera responsible for this area. When the person shown in the screen leaves, if there are still people not in the screen in the conference room, then one person is selected to be shown; if there are no people not in the screen in the conference room, the screen layout changes and one is reduced.

[0080] In the embodiment of the present application, when the person moves within the coverage area of the camera, the screen will be adjusted in real time to ensure that the person is centered. When the person moves out of the current camera coverage area, the person will be lost immediately, and another person not in the screen will replace him. If there are less than four people in the conference room, then the person in motion will disappear and reappear after moving across the camera.

[0081] The present application realizes accurate speaker positioning and screen cropping, provides flexible screen layout adjustment ability, maintains the screen stability when there is no speaker, and improves the intelligent level of the conference system.

[0082] The embodiment of the present application provides a screen presentation device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the screen presentation method as described above.

[0083] The embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the screen presentation method as described above when running.

[0084] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and products can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0085] In several embodiments provided by the present application, it should be understood that the disclosed methods, systems, devices, and program products can be implemented in other ways.

[0086] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0087] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A video conferencing system, characterized in that, Divide the meeting venue into corresponding numbers of meeting areas according to the number of cameras, and further include: Several cameras for shooting the picture information of the corresponding meeting areas; A microphone array for collecting the speech information and position information of the speaker; A meeting host for judging the meeting area to which the speaker's position belongs according to the speech information and position information to obtain judgment result information; and cropping the meeting area picture information according to the judgment result information to obtain the picture information to be sent to the display screen.

2. The video conferencing system according to claim 1, wherein Further include: A data input device for inputting the meeting venue size information, camera position information and camera angle information, and allocating the corresponding meeting area picture information of the camera according to the meeting venue size information, camera coordinate information and camera angle information.

3. The video conferencing system according to claim 2, characterized in that, The data input device is further used for inputting the meeting type information, the number of cameras is greater than or equal to 2, and the camera height is the first threshold.

4. The video conferencing system according to claim 3, wherein, When the meeting type information is a horizontal meeting, there is a field of view overlap between two adjacent cameras; When the meeting type information is a non-horizontal meeting, the field of view of the left camera covers at least all the portraits on the right half of the meeting room, and the field of view of the right camera covers at least all the portraits on the left half of the meeting room.

5. The video conferencing system according to claim 1, characterized in that, Each camera and the microphone array are an integrated device, and the integrated device has at least one camera and one microphone array.

6. A method for presenting a screen, characterized in that, Include: Judging the meeting area to which the speaker's position belongs according to the speech information and position information to obtain judgment result information; Cropping the meeting area picture information according to the judgment result information to obtain the close-up picture information to be sent to the display screen.

7. The method for presenting a screen according to claim 6, wherein Further include: Obtaining the number of participants information, and when the number of participants is greater than the first preset value, outputting the first initial picture information for displaying the first preset value of quantity.

8. The screen presentation method according to claim 7, wherein Further include: Obtaining the number of participants information, and when the number of participants is less than the first preset value, outputting the second initial picture information for displaying the number of participants.

9. The method for presenting a screen according to claim 8, wherein Further include: When the number of speakers in the close-up picture information is zero and the number of participants information is greater than zero, adjusting the close-up picture information according to the first initial picture information or the second initial picture information.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program is set to execute the picture presentation method according to any one of claims 6-9 when running.