Video conference picture switching method, system, equipment and medium

Through the combination of microphone array and camera, the speaker's position is determined and the screen switching is optimized, which solves the visual fatigue caused by frequent lens movement in traditional video conferencing and improves the conference experience.

CN120547296APending Publication Date: 2025-08-26YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510616204.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In traditional video conferencing, the visual fatigue and dizziness caused by frequent and rapid movement of the camera lens and zooming are not effective in meeting experience, especially when multiple speakers take turns speaking.

Method used

The voice signal is obtained through the microphone array, the speaker's position is determined, and the camera is controlled to capture the speaker's picture based on the sound source angle and the picture portrait angle captured by the camera; when the angle angle changes, insert the middle transition screen or switch directly to avoid large-scale movement of the lens.

Benefits of technology

It effectively reduces the frequent and rapid movement of the camera lens, improves the visual experience and comfort of the conference participants, and reduces visual fatigue and dizziness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547296A_ABST
    Figure CN120547296A_ABST
Patent Text Reader

Abstract

The invention discloses a video conference picture switching method, system and device and a medium, and the method comprises the steps: obtaining a voice signal in a conference room through a microphone array, and carrying out the preprocessing of the voice signal; determining the position of the speaker based on the sound source angle of the preprocessed voice signal and the image portrait angle captured by the camera; based on the position of the speaker, controlling a first camera to capture the portrait of the speaker, and outputting the picture of the speaker; when other speakers appear, setting a virtual vertex position and calculating an included angle formed between the position of the current speaker and the position of the previous speaker; and controlling the first camera to select different picture switching modes to output pictures of the speaker according to the included angle. According to the invention, visual fatigue and dizziness caused by frequent and rapid movement of a camera lens and rapid movement of a picture during zooming can be effectively avoided, and the visual experience and comfort of conference participants are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video image processing technology, and in particular to a method, system, device and medium for switching images in a video conference. Background Art

[0002] With the development of video conferencing, it has become an important tool for remote communication and collaboration. In video conferencing, the camera's voice tracking mode, a mature technology, can track speakers in the conference room in real time and automatically adjust the lens to capture a clear image of the speaker. However, traditional voice tracking mode directly displays the camera lens's framing process when switching speakers. For two speakers far apart, or when there are multiple speakers taking turns speaking in the conference room, the camera lens needs to frequently and rapidly move and zoom to capture the image. The rapid changes in the conference image cause visual fatigue and dizziness in viewers, further reducing the conference experience. Summary of the Invention

[0003] The main purpose of this application is to overcome the shortcomings and deficiencies of the existing technology and provide a video conferencing screen switching method, system, equipment and medium, which effectively avoids the visual fatigue and dizziness caused by frequent and rapid movement of the camera lens and rapid movement of the screen during zooming, and improves the visual experience and comfort of the conference participants.

[0004] In order to achieve the above objectives, this application adopts the following technical solutions:

[0005] In a first aspect, the present application provides a method for switching screens in a video conference, comprising the following steps:

[0006] Acquire voice signals in the conference room through a microphone array and pre-process the voice signals;

[0007] Determine the speaker's position based on the sound source angle of the pre-processed speech signal and the portrait angle of the image captured by the camera;

[0008] Based on the position of the speaker, controlling the first camera to capture the speaker's portrait and outputting the speaker's image;

[0009] When another speaker appears, a virtual vertex is set and the angle formed between the position of the current speaker and the position of the previous speaker is calculated;

[0010] The first camera is controlled to select different image switching modes according to the included angle to output the image of the speaker.

[0011] As a preferred technical solution, the preprocessing includes noise reduction, de-echo and gaining of the speech signal.

[0012] As a preferred technical solution, the first camera selects different screen switching modes according to the angle, including:

[0013] If the angle is greater than a preset angle value, the first camera is controlled to switch to the current speaker's image, and a preset intermediate transition image is inserted to replace each frame in the switching process until the first camera captures the speaker's portrait and is in the center of the conference video screen, and then the speaker's image captured by the first camera is output;

[0014] If the angle is smaller than the preset angle value, the first camera is controlled to turn directly to the current speaker's screen.

[0015] As a preferred technical solution, when there are two speakers taking turns speaking, the first camera or the second camera is selected to capture the portrait of the first speaker or the second speaker according to the positions of the speakers, thereby obtaining the first speaker image and the second speaker image;

[0016] The first speaker's screen and the second speaker's screen are spliced ​​and displayed.

[0017] As a preferred technical solution, the method of selecting the first camera or the second camera to capture the portrait of the first speaker or the second speaker to obtain the first speaker image and the second speaker image includes:

[0018] Determining the position distances between the first speaker, the second speaker, and the first camera, and the second camera, respectively;

[0019] If the distance between the first speaker or the second speaker and the second camera is within a range preset by the second camera, capturing the portrait of the first speaker or the second speaker using the second camera, and outputting the image of the first speaker or the second speaker using the second camera;

[0020] Otherwise, the first camera is used to capture the first speaker portrait or the second speaker portrait, and the first camera selects different screen switching methods to output the first speaker screen or the second speaker screen according to the angle formed between the positions of the first speaker and the second speaker.

[0021] As a preferred technical solution, the second camera is a panoramic camera.

[0022] As a preferred technical solution, it also includes:

[0023] When there are multiple speakers speaking within a preset time, the first camera or the second camera captures the speaker portraits and outputs the images of the speakers speaking within the preset time.

[0024] As a preferred technical solution, when a new speaker joins, the new speaker's screen is output on the screen by splicing with the existing speaker's screen.

[0025] In a second aspect, the present application provides a video conferencing screen switching system, which is applied to the aforementioned video conferencing screen switching method, including a preprocessing module, a position determination module, a screen output module, an angle calculation module, and a switching mode selection module;

[0026] The preprocessing module is used to obtain the voice signal in the conference room through the microphone array and preprocess the voice signal;

[0027] The position determination module is used to determine the position of the speaker based on the sound source angle of the pre-processed voice signal and the portrait angle of the image captured by the camera;

[0028] The output image module is used to control the first camera to capture the speaker's portrait based on the speaker's position and output the speaker's image;

[0029] The angle calculation module is used to calculate the angle formed between the position of the current speaker and the position of the previous speaker with the camera lens as the vertex position when another speaker appears;

[0030] The switching mode module is used to control the first camera to select different image switching modes according to the angle to output the image of the speaker.

[0031] In a third aspect, the present application provides an electronic device, comprising:

[0032] at least one processor; and a memory communicatively coupled to the at least one processor;

[0033] The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the video conference screen switching method.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the video conference screen switching method.

[0035] In summary, compared with the prior art, the effective effects brought about by the technical solution provided by this application include at least:

[0036] This application proposes a method for switching screens in a video conference. The method uses a microphone array to obtain voice signals in a conference room and pre-process the voice signals. The method determines the speaker's position based on the sound source angle of the processed voice signal and the angle of the image captured by the camera. The method controls a first camera to capture the speaker's image and output the speaker's image. When another speaker appears, the method uses the camera lens as the vertex position to calculate the angle between the current speaker's position and the previous speaker's position. The method controls the first camera to select different screen switching methods based on the angle to output the speaker's image. Selecting a screen switching method based on the angle formed by the change in the speaker's position effectively avoids visual fatigue and dizziness caused by frequent and rapid camera lens movement and rapid image movement during zooming, thereby improving the visual experience and comfort of conference participants. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 A flowchart of a method for switching screens in a video conference provided in one embodiment of the present application;

[0039] Figure 2 A block diagram of a video conferencing screen switching system provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0041] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0042] This application adopts the following technical solutions:

[0043] See also Figure 1 In one embodiment of the present application, a method for switching screens in a video conference is provided, comprising the following steps:

[0044] S1. Acquire a voice signal in a conference room through a microphone array and pre-process the voice signal.

[0045] Furthermore, the present application obtains the speaker's voice through a microphone array in a multi-person conference room, and sends the voice signal to the host signal processing module in the video conferencing system, which performs pre-processing such as noise reduction, de-echo and voice gain.

[0046] S2. Determine the speaker's position based on the sound source angle of the preprocessed voice signal and the portrait angle of the image captured by the camera.

[0047] The pre-processed voice signal is sent to the host judgment module in the video conferencing system, and the position of the speaker is determined by the sound source angle of the pre-processed voice signal and the portrait angle of the image captured by the camera.

[0048] Specifically, when someone speaks in a meeting, the microphone array can calculate the angle of the current sound source through sound source positioning technology. The camera will determine whether there is a human figure in that direction based on the audio direction information. If there is a human figure, and the angle of the human figure and the angle of the sound source are within a certain error range, and the sound source is confirmed to be the person speaking, the regional position of the speaker can be determined. Then, based on the size of the head frame of the human figure captured by the camera, the distance between the human figure and the camera can be judged, and the specific position of the speaker can be determined.

[0049] S3. Based on the position of the speaker, control the first camera to capture the speaker's portrait and output the speaker's image.

[0050] After determining the position of the speaker, the first camera is controlled to capture the speaker's portrait and output a close-up image of the speaker; and according to the size of the portrait in the image, the magnification of the lens is adjusted to enlarge the portrait to an appropriate proportion.

[0051] The first camera of this application is a PTZ (Pan-Tilt-Zoom) camera with pan / tilt / zoom functions. Pan refers to the horizontal rotation function of the camera, allowing the camera to rotate horizontally to cover a wider area. Tilt refers to the vertical rotation function of the camera, which allows the camera to rotate vertically to adjust the camera's viewing angle; zoom refers to the optical magnification function of the camera, allowing the camera to adjust the focal length, thereby enlarging or reducing the view of the object being photographed.

[0052] S4. When another speaker appears, set a virtual vertex and calculate the angle between the current speaker's position and the previous speaker's position. The virtual vertex is a point on the conference room's central axis or a point on the camera's focal line. The virtual vertex is preferably the location of the camera lens.

[0053] S5. Control the first camera to select different image switching modes according to the included angle to output the image of the speaker.

[0054] Furthermore, when other speakers appear, the position of the current speaker is still determined by the sound source angle and the portrait angle of the image captured by the camera; at this time, the position of the current speaker, the position of the previous speaker and the camera lens will form an angle, therefore, with the camera lens as the vertex position, calculate the angle formed between the position of the current speaker and the position of the previous speaker, and determine whether the angle is greater than the preset angle value.

[0055] If the angle is greater than the preset angle value, the first camera is controlled to switch to the current speaker's screen, and a preset black image is inserted to replace each frame in the switching process until the first camera captures the speaker's portrait and is in the center of the conference video screen, and then the speaker's screen captured by the first camera is output.

[0056] If the angle is smaller than the preset angle value, the first camera is controlled to directly switch to the current speaker's screen;

[0057] Furthermore, the preset angle value in the present application is less than 50°, and preferably may be 25°.

[0058] Specifically, when the angle between the current speaker's position and the previous speaker's position is less than 25°, the first camera's lens only needs to rotate slightly or even not at all to track the current new speaker. The lens will not rotate significantly, and the image content will not change rapidly. Therefore, the first camera's lens swing (i.e., the process of PTZ capturing and selecting speakers, moving from one speaker to another) is retained.

[0059] When the angle between the position of the current speaker and the position of the previous speaker is greater than 25°, the lens of the first camera needs to rotate significantly (i.e., perform a large PTZ, i.e., pan / tilt, rotation, and zoom operation) to track the new speaker. Large lens rotation will cause the content of the picture to change rapidly and significantly. This rapid switching of pictures may cause visual discomfort or dizziness to the viewer, especially when switching speakers quickly. Therefore, when the lens of the first camera of this application starts PTZ operation (i.e., starts rotating to track the new speaker), a black picture will be inserted to replace each frame of the picture during the lens rotation process. The black picture is inserted as a transition for picture switching, making the picture switching process smoother and avoiding the dizziness caused by large-angle lens rotation and large changes in picture content.

[0060] Furthermore, in the present application, when there are two speakers taking turns speaking (i.e., intercom mode), the first camera or the second camera is selected to capture the portrait of the first speaker or the second speaker according to the position of the speaker, and the first speaker screen and the second speaker screen are obtained; the first speaker screen and the second speaker screen are spliced ​​and displayed.

[0061] Specifically, selecting the first camera or the second camera to capture the portrait of the first speaker or the portrait of the second speaker to obtain the first speaker image and the second speaker image includes the following steps:

[0062] Determining the position distances between the first speaker, the second speaker, and the first camera, and the second camera, respectively;

[0063] If the distance between the first speaker or the second speaker and the second camera is within a range preset by the second camera, capturing the portrait of the first speaker or the second speaker using the second camera, and outputting the image of the first speaker or the second speaker using the second camera;

[0064] Otherwise, the first camera is used to capture the first speaker portrait or the second speaker portrait, and the first camera selects different screen switching methods to output the first speaker screen or the second speaker screen according to the angle formed between the positions of the first speaker and the second speaker.

[0065] Furthermore, there may be multiple first cameras, each of which may be responsible for a different shooting area, and the displayed images may be output individually or after being spliced ​​and output.

[0066] Furthermore, the second camera in the present application is a panoramic camera, wherein the first camera is responsible for capturing the portrait of the speaker at a long distance, and the second camera is responsible for capturing the portrait of the speaker at a close distance.

[0067] Furthermore, the panoramic camera is a fixed camera, and the panoramic effect is achieved by combining multiple fixed cameras.

[0068] This application takes into account the position distance between the speaker and the camera. When the speaker is within the preset range of a camera, that camera is used for image capture first, thereby effectively utilizing the camera resources in the conference room, avoiding unnecessary camera switching and image adjustment, and improving the overall system operation efficiency and stability.

[0069] When two or more speakers take turns speaking, the display screen is dynamically adjusted according to the changes in the speakers in the conference room, including: when A is speaking, a close-up picture of A is displayed; when B is speaking, a combined picture of A and B is displayed; when A stops speaking and B continues speaking, a close-up picture of B is displayed; and when no one is speaking, the display returns to the camera's AF picture.

[0070] Among them, there are two optional display modes for the combined picture: portrait frame effect and stitching effect.

[0071] Portrait framing effect: When the first speaker speaks, the camera directly captures and frames the first speaker's portrait, outputting the first speaker's image. When the second speaker speaks, the camera captures and frames the first and second speakers, and outputs the first and second speakers' images. When the third speaker speaks, the camera captures and frames the third speaker, and simultaneously determines the positional relationship of the three speakers. If the three speakers are close enough to be clearly displayed in the same frame, they are displayed simultaneously. If the first speaker is farther away, the second and third speakers' portraits are displayed together, framed. If the second and third speakers are farther away, the camera automatically switches to a spliced ​​effect. If one speaker speaks continuously for 5 seconds without anyone else speaking, the camera returns to framing the speaker. If the speaker is silent for 7 seconds, the camera returns to the AF mode. AF (Auto Framing) refers to automatic portrait framing, an AI function used by the camera to output images during video conferencing. The automatic portrait framing function allows the lens to automatically adjust the angle and magnification according to the number of participants, outputting a picture with appropriate size and reasonable layout that includes everyone in the meeting room.

[0072] Stitching effect: The process of capturing and selecting speakers is blocked. When the first speaker speaks, the first speaker is first displayed using a fade-in / fade-out method. When the second speaker speaks, the first and second speakers' split screens are displayed by panning in from the left or right, depending on their relative positions. When the third speaker speaks, the first speaker is replaced with a fade-in / fade-out effect. If one speaker speaks continuously for 5 seconds without anyone else speaking, the other speaker's split screen is destroyed by a push-out from the left or right method. If the speaker does not speak for 7 seconds, the screen returns to the AF mode.

[0073] In this application, when the device switches between the gimbal lens and the electronic lens or when switching within the electronic lens, because the camera movement cannot be controlled, cutting the image by cropping will result in a fast jumping feeling, so the picture adds a fade-in and fade-out transition effect; fade-in and fade-out means gradually increasing the transparency at the current end frame, that is, the transparency goes from 0 to 1; gradually reducing the transparency at the start frame of the next frame, that is, the transparency goes from 1 to 0; the two frames of special effects are connected to form a fade-in and fade-out effect.

[0074] The present application also includes: when multiple speakers speak within a preset time, the first camera or the second camera captures the speaker's portrait, and outputs the image of the speaker who spoke within the preset time. When a new speaker joins, the new speaker's image is output on the screen by splicing it with the existing speaker's image.

[0075] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the described order of actions, because according to this application, certain steps can be performed in other orders or simultaneously.

[0076] Based on the same idea as the video conferencing screen switching method in the above embodiment, the present application also provides a video conferencing screen switching system, which can be used to execute the above video conferencing screen switching method. For ease of explanation, the structural diagram of an embodiment of a video conferencing screen switching system only shows the parts related to the embodiment of the present application. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the system, and may include more or fewer components than shown in the diagram, or combine certain components, or arrange the components differently.

[0077] See also Figure 2 In another embodiment of the present application, a video conference screen switching system is provided, which includes a pre-processing module 101, a position determination module 102, a screen output module 103, an angle calculation module 104, and a switching mode selection module 105;

[0078] The preprocessing module 101 is used to obtain the voice signal in the conference room through the microphone array and preprocess the voice signal;

[0079] The position determination module 102 is configured to determine the position of the speaker based on the sound source angle of the pre-processed voice signal and the portrait angle of the image captured by the camera;

[0080] The output image module 103 is used to control the first camera to capture the speaker's portrait based on the speaker's position and output the speaker's image;

[0081] The angle calculation module 104 is used to calculate the angle between the position of the current speaker and the position of the previous speaker using the camera lens as the vertex position when another speaker appears;

[0082] The switching mode module 105 is used to control the first camera to select different image switching modes according to the angle to output the image of the speaker.

[0083] It should be noted that a video conferencing screen switching system of the present application corresponds one-to-one to a video conferencing screen switching method of the present application. The technical features and beneficial effects described in the embodiment of the above-mentioned video conferencing screen switching method are all applicable to the embodiment of a video conferencing screen switching system. For specific contents, please refer to the description in the embodiment of the method of the present application. No further details will be given here. This is hereby declared.

[0084] In addition, in the implementation of a video conferencing screen switching system in the above embodiment, the logical division of each program module is only an example. In actual applications, the above functions can be assigned to different program modules as needed, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation. That is, the internal structure of the video conferencing screen switching system is divided into different program modules to complete all or part of the functions described above.

[0085] In another embodiment, an electronic device for implementing a video conferencing screen switching method is provided, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, a video conferencing screen switching method of any embodiment of the present application is implemented.

[0086] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more module elements may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device.

[0087] The device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The device may include, but is not limited to, a processor and a memory.

[0088] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, connecting various parts of the entire device using various interfaces and lines.

[0089] The memory can be used to store the computer programs and / or modules, and the processor realizes various functions of the device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; in addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0090] Accordingly, the present application also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a video conference screen switching method as described in any one of the above embodiments.

[0091] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0092] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0093] The above embodiments are preferred implementation modes of the present application, but the implementation modes of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present application should be considered as equivalent replacement methods and are included in the scope of protection of the present application.

Claims

1. A video conference screen switching method, characterized in that: The steps include: Acquire voice signals in the conference room through a microphone array and pre-process the voice signals; Determine the speaker's position based on the sound source angle of the pre-processed speech signal and the portrait angle of the image captured by the camera; Based on the position of the speaker, controlling the first camera to capture the speaker's portrait and outputting the speaker's image; When another speaker appears, a virtual vertex is set and the angle formed between the position of the current speaker and the position of the previous speaker is calculated; The first camera is controlled to select different image switching modes according to the included angle to output the image of the speaker.

2. A video conference screen switching method according to claim 1, characterized in that: The preprocessing includes performing noise reduction, de-echo and gain on the speech signal.

3. A video conference screen switching method according to claim 1, characterized in that: The first camera selects different image switching modes according to the angle, including: If the angle is greater than a preset angle value, the first camera is controlled to switch to the current speaker's image, and a preset intermediate transition image is inserted to replace each frame in the switching process until the first camera captures the speaker's portrait and is in the center of the conference video screen, and then the speaker's image captured by the first camera is output; If the angle is smaller than the preset angle value, the first camera is controlled to turn to the current speaker's screen.

4. The method for switching screens in a video conference according to claim 1, wherein: When there are two speakers who take turns speaking, the first camera or the second camera is selected to capture the portrait of the first speaker or the second speaker according to the positions of the speakers, thereby obtaining the first speaker image and the second speaker image; The first speaker's screen and the second speaker's screen are spliced ​​and displayed.

5. A video conference screen switching method according to claim 4, characterized in that: The selecting the first camera or the second camera to capture the portrait of the first speaker or the portrait of the second speaker to obtain the first speaker image and the second speaker image includes: Determining the position distances between the first speaker, the second speaker, and the first camera, and the second camera, respectively; If the distance between the first speaker or the second speaker and the second camera is within a range preset by the second camera, capturing the portrait of the first speaker or the second speaker using the second camera, and outputting the image of the first speaker or the second speaker using the second camera; Otherwise, the first camera is used to capture the first speaker portrait or the second speaker portrait, and the first camera selects different screen switching methods to output the first speaker screen or the second speaker screen according to the angle formed between the positions of the first speaker and the second speaker.

6. The second camera according to claim 5, characterized in that The second camera is a panoramic camera.

7. A video conference screen switching method according to claim 5, characterized in that: Also includes: When there are multiple speakers speaking within a preset time, the first camera or the second camera captures the speaker portraits and outputs the images of the speakers speaking within the preset time.

8. A video conference screen switching method according to claim 7, characterized in that: When a new speaker joins, the new speaker's screen is output on the screen by splicing with the existing speaker's screen.

9. A video conference screen switching system, characterized in that: A video conference screen switching method applied to any one of claims 1-7, comprising a pre-processing module, a position determination module, a screen output module, an angle calculation module, and a switching mode selection module; The preprocessing module is used to obtain the voice signal in the conference room through the microphone array and preprocess the voice signal; The position determination module is used to determine the position of the speaker based on the sound source angle of the pre-processed voice signal and the portrait angle of the image captured by the camera; The output image module is used to control the first camera to capture the speaker's portrait based on the speaker's position and output the speaker's image; The angle calculation module is used to calculate the angle formed between the position of the current speaker and the position of the previous speaker with the camera lens as the vertex position when another speaker appears; The switching mode module is used to control the first camera to select different image switching modes according to the angle to output the image of the speaker.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute a video conference screen switching method as described in any one of claims 1-8.