Picture-in-picture display method and device for conference video

Through the combination of voice and gesture recognition, the picture-in-picture display is flexibly triggered, which solves the problem of single picture display mode in the existing technology, and realizes picture-in-picture display with multiple configuration requirements, adapting to multiple conference modes.

CN120343301APending Publication Date: 2025-07-18YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510616208.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing picture-in-picture technology, the screen display method is single, and different sizes cannot be displayed in different positions, and it cannot adapt to the needs of multiple conference modes.

Method used

Through the combination of voice recognition and gesture recognition, the picture-in-picture display is flexibly triggered, and the display of the first and second screens is determined based on the picture-in-picture configuration parameters, supporting a variety of different configuration requirements.

Benefits of technology

It improves the flexibility of triggering method for picture-in-picture display, can adapt to a variety of different configuration requirements, and meets the display needs of multiple conference modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343301A_ABST
    Figure CN120343301A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and discloses a picture-in-picture display method and device for a conference video, and the method comprises the steps: receiving a picture-in-picture configuration parameter and a picture-in-picture starting instruction; obtaining conference video data and corresponding voice data, and extracting a current video frame; performing voice recognition on the voice data to obtain a voice recognition result; performing gesture recognition on the current video frame to obtain a gesture recognition result; judging whether the voice recognition result and / or the gesture recognition result meet a preset triggering condition or not; and if yes, determining and displaying a first picture and a second picture of the current video frame according to the picture-in-picture configuration parameters. The flexibility of the triggering mode can be improved, and picture-in-picture display with various different configuration requirements is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a method and device for displaying a picture-in-picture of a conference video. Background Art

[0002] Picture-in-picture is to display two pictures on the same screen, that is, to insert a sub-picture on the main picture being viewed normally, so as to view the two pictures simultaneously. In the existing picture-in-picture technology processing, it is possible to support the simultaneous display of different pictures, but the method behavior is relatively single. For example, a close-up picture and a panoramic shot picture can only be displayed at a fixed position, or the picture content of the two pictures cannot be selected, and different sizes cannot be displayed at different positions and used in different conference modes. Summary of the Invention

[0003] This application provides a method and device for displaying a picture-in-picture of a conference video, which can improve the flexibility of the triggering method and support the picture-in-picture display with various different configuration requirements.

[0004] In a first aspect, an embodiment of this application provides a method for displaying a picture-in-picture of a conference video, including:

[0005] Receiving picture-in-picture configuration parameters and a picture-in-picture start instruction;

[0006] Obtaining conference video data and corresponding voice data, and extracting the current video frame;

[0007] Performing speech recognition on the voice data to obtain a speech recognition result;

[0008] Performing gesture recognition on the current video frame to obtain a gesture recognition result;

[0009] Judging whether the speech recognition result and / or the gesture recognition result meet a preset trigger condition; if so, determining and displaying a first picture and a second picture of the current video frame according to the picture-in-picture configuration parameters.

[0010] Further, the picture-in-picture configuration parameters include a first picture type, a second picture position, and a second picture display ratio.

[0011] Further, the first picture and the second picture are different; at least one of them is a close-up portrait picture.

[0012] Further, the person in the close-up portrait picture is the trigger person of the speech recognition result or the gesture recognition result that meets the preset trigger condition.

[0013] Further, the method further includes:

[0014] Obtaining the current video working mode; determining the preset trigger condition according to the current video working mode.

[0015] Further, the method further includes:

[0016] Determine whether the speech recognition result or the gesture recognition result meets the preset trigger condition of any video working mode; if so, use the video working mode corresponding to the satisfied preset trigger condition as the target working mode and switch;

[0017] After switching to the target working mode, generate and display a switching setting reminder;

[0018] In response to a setting instruction, display a parameter configuration interface; receive the current configuration parameters of the target working mode;

[0019] Determine and display the first picture and the second picture of the current video frame according to the current configuration parameters.

[0020] Further, the method further includes:

[0021] If the number of preset trigger conditions satisfied by the speech recognition result or the gesture recognition result is greater than 1, determine the target working mode and switch among the video working modes corresponding to the satisfied preset trigger conditions according to the preset mode priority.

[0022] Further, the method further includes:

[0023] Store the received current configuration parameters and the corresponding target working mode in a parameter setting database;

[0024] After generating the switching setting reminder, in response to a display instruction, determine and display the first picture and the second picture of the current video frame according to the current configuration parameters corresponding to the target working mode in the parameter setting database.

[0025] Further, the method further includes:

[0026] Receive a configuration parameter modification instruction; update the parameter setting database according to the configuration parameter modification instruction.

[0027] In a second aspect, an embodiment of the present application provides a picture-in-picture display device for conference videos, including:

[0028] A receiving module, configured to receive picture-in-picture configuration parameters and a picture-in-picture start instruction;

[0029] An obtaining module, configured to obtain conference video data and corresponding voice data, and extract the current video frame;

[0030] A speech recognition module, configured to perform speech recognition on the voice data to obtain a speech recognition result;

[0031] A gesture recognition module, configured to perform gesture recognition on the current video frame to obtain a gesture recognition result;

[0032] A determination module, configured to determine whether a speech recognition result and / or a gesture recognition result meets a preset trigger condition; if so, determine and display a first picture and a second picture of a current video frame according to picture-in-picture configuration parameters.

[0033] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it performs the steps of a method for displaying a picture-in-picture of a conference video according to any one of the above embodiments.

[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a method for displaying a picture-in-picture of a conference video according to any one of the above embodiments.

[0035] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0036] A method for displaying a picture-in-picture of a conference video provided by an embodiment of the present application first triggers the display of the picture-in-picture through two methods of speech recognition and gesture recognition, improving the flexibility of the trigger method; secondly, after determining to trigger the picture-in-picture, the first picture and the second picture are determined according to the received picture-in-picture configuration parameters, supporting the display of picture-in-picture with a variety of different configuration requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart of a method for displaying a picture-in-picture of a conference video provided by an exemplary embodiment of the present application.

[0038] Figure 2 It is a structural diagram of a device for displaying a picture-in-picture of a conference video provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0040] Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0041] Please refer to Figure 1 , an embodiment of the present application provides a method for displaying a picture-in-picture of a conference video, including:

[0042] Step S11, receiving picture-in-picture configuration parameters and a picture-in-picture start instruction.

[0043] Specifically, the Picture-in-Picture start instruction is an instruction to enable the Picture-in-Picture function. That is, after receiving the Picture-in-Picture start instruction, the configuration interface is displayed, and the Picture-in-Picture configuration parameters input by the user in the configuration interface are received.

[0044] Step S12: Obtain the conference video data and the corresponding voice data, and extract the current video frame.

[0045] Among them, the conference video data can be obtained through the camera in the conference room, specifically a panoramic camera. The conference video data includes multiple video frames, and the current video frame is the video frame at the most recent moment among all the video frames.

[0046] Step S13: Perform speech recognition on the voice data to obtain the speech recognition result.

[0047] Among them, speech recognition is to identify whether someone is speaking, and whether there is a human voice is used as the speech recognition result.

[0048] Step S14: Perform gesture recognition on the current video frame to obtain the gesture recognition result.

[0049] Step S15: Determine whether the speech recognition result and / or the gesture recognition result meet the preset trigger conditions; if so, determine and display the first picture and the second picture of the current video frame according to the Picture-in-Picture configuration parameters.

[0050] Among them, the Picture-in-Picture configuration parameters may include the first picture type, the second picture position, and the second picture display ratio.

[0051] Specifically, the first picture is the main picture. The first picture type can be a close-up picture or a panoramic picture, or the picture content after automatically framing the portrait in the panoramic picture. The close-up picture can be an adjusted PTZ picture (the PTZ picture is a fixed picture manually controlled by the camera before the start of the meeting), the speaker picture determined and framed through voice tracking, the picture of the speaker tracked in the follow-up mode, etc.; the second picture is the small picture, which is generally defaulted to the panoramic camera picture.

[0052] It can be considered that the first picture and the second picture are different; at least one of them is a close-up portrait picture; further, the portrait in the close-up portrait picture is the trigger person of the speech recognition result or the gesture recognition result that meets the preset trigger conditions.

[0053] It should be noted that the extraction of the above first picture type is a known technology, such as the close-up positioning of the speaker picture, the portrait tracking and positioning, and the framing, etc., which will not be elaborated here.

[0054] The position of the second screen can be selected at the lower right, upper left, lower left, etc. of the entire display interface, and its screen edge coincides with the screen edge of the entire display image; the display ratio of the second screen is adjustable, and 1 / 4 or 1 / 9 is preferably selected. This ratio refers to 1 / 4 or 1 / 9 of the entire display area of the second screen, and the output size of the second screen is reduced proportionally for display.

[0055] When the preset trigger condition is that someone is speaking, or the gesture is five fingers spread, and the voice recognition result is that there is a human voice or the gesture recognition result is five fingers spread, then process the current video frame according to the picture-in-picture configuration parameters, scale the original video frame as the second screen, and set it to the corresponding position. Extract the content corresponding to the first screen type from the current video frame and display it as the main screen, that is, fill the entire display area; when receiving a click instruction in the display area of the second screen by the user, the display parameters of the first screen and the second screen can be exchanged, that is, the second screen is displayed as the main screen filling the display area, and the first screen is scaled and placed according to the relevant parameters of the second screen in the picture-in-picture configuration parameters.

[0056] It should be noted that the above-mentioned gesture of five fingers spread for recognition is only a preferred embodiment. When the preset trigger condition is gesture recognition, other triggering gestures with strong recognition ability can be selected, such as making a fist, etc.

[0057] A method for picture-in-picture display of a conference video provided in the above embodiment triggers the picture-in-picture display through voice recognition and gesture recognition in two ways, improving the flexibility of the triggering method; secondly, after determining to trigger the picture-in-picture, determine the display of the first screen and the second screen according to the received picture-in-picture configuration parameters, supporting picture-in-picture displays with various different configuration requirements.

[0058] Furthermore, in some embodiments, the preset trigger condition can be set with both a voice trigger condition and a gesture trigger condition at the same time, and the corresponding voice recognition and gesture recognition can also be set for sequential recognition. That is, in the received picture-in-picture configuration parameters, the user can set the recognition priority sorting option, and the recognition order of the two recognition results of voice and gesture is executed according to the recognition priority sorting option in the received picture-in-picture configuration parameters. Among them, the recognition priority sorting option can be set to: priority recognition, simultaneous recognition, and prohibited recognition, where priority recognition is restricted to choose one of the two.

[0059] In some embodiments, the method further includes:

[0060] Obtain the current video working mode; determine the preset trigger condition according to the current video working mode.

[0061] Specifically, video conferences usually have multiple working modes, such as PTZ mode (portrait close-up mode), follow-up mode, etc.

[0062] This application can set preset trigger conditions for picture-in-picture display in different video working modes.

[0063] For example, the preset trigger condition for the PTZ mode is that someone is speaking. Then, when the video conference is working in the PTZ mode and the voice recognition result is that there is human voice, the first picture and the second picture are determined and displayed according to the picture-in-picture configuration parameters.

[0064] Among them, the second picture in the PTZ mode is displayed as a close-up picture of the condition triggerer, and the displayed condition triggerer can be one or more.

[0065] It should be noted that in the specific implementation process, the recognition algorithm for the conference video data can be set according to the preset trigger conditions of each existing video working mode. For example, the preset trigger condition for the follow-up mode is a voice command, and the corresponding voice recognition needs to be accurate to text conversion and analysis of the received audio.

[0066] Furthermore, the trigger gestures in the preset trigger conditions can also include special gestures and other gestures.

[0067] In addition to the control of the picture-in-picture configuration parameters, when a special gesture recognition is triggered, the second picture can be preferably set as the picture in the follow-up mode, that is, the second picture shows a close-up of the gesture recognizer and follows until the second end gesture is recognized; when other gestures are triggered, the second picture is preferably set as the picture in the portrait close-up mode, and can be optionally displayed simultaneously with the portrait close-up of the voice recognition.

[0068] The above embodiments can enable each video conference working mode to have a corresponding and different picture-in-picture preset trigger method, further improving the flexibility of the picture-in-picture display.

[0069] In some embodiments, the method further includes:

[0070] Step S21, determining whether the voice recognition result or the gesture recognition result meets the preset trigger condition of any video working mode; if so, using the video working mode corresponding to the satisfied preset trigger condition as the target working mode and switching.

[0071] Furthermore, if the number of preset trigger conditions satisfied by the voice recognition result or the gesture recognition result is greater than 1, the target working mode is determined and switched among the video working modes corresponding to each satisfied preset trigger condition according to the preset mode priority.

[0072] Specifically, the recognition result of each recognition algorithm can be matched with the preset trigger conditions of each video working mode. If there is a matching preset trigger condition but the corresponding video working mode is not the current video working mode, a switch is made.

[0073] When there are multiple preset trigger conditions that match the recognition result, then according to the preset mode priority, compare which mode has the highest priority among the various video working modes that meet the preset trigger conditions, and use it as the target working mode.

[0074] It should be noted that the preset mode priority here does not conflict with the recognition result priority mentioned above. The recognition result priority is for recognition in the case where there are multiple preset trigger conditions for one mode, while the preset mode priority here is to select which mode to switch and the picture-in-picture display when the recognition result meets the preset trigger conditions of multiple video working modes.

[0075] Step S22, after switching to the target working mode, generate and display a switching setting reminder.

[0076] Step S23, in response to the setting instruction, display the parameter configuration interface; receive the current configuration parameters of the target working mode.

[0077] Step S24, determine and display the first picture and the second picture of the current video frame according to the current configuration parameters.

[0078] Specifically, the present application also supports allowing the user to set the picture-in-picture configuration parameters in real time when switching the video working mode, that is, after switching to the target working mode, a switching setting reminder is generated. There are two options under this reminder, corresponding to the setting instruction and the display instruction respectively. If the user clicks the setting option, the parameter configuration interface is displayed, allowing the user to operate on the corresponding picture-in-picture function configuration items in different modes, and perform picture-in-picture display according to the current configuration parameters of the target working mode.

[0079] The above embodiments support both allowing the user to set the picture-in-picture configuration parameters of the current video working mode before the meeting starts, and allowing the user to configure the current configuration parameters of the target working mode in real time when switching the video working mode.

[0080] For the end of the picture-in-picture display, it can be by receiving the user's picture-in-picture end instruction, or corresponding preset end conditions can be set for the preset trigger conditions. When it is recognized that the current video frame meets the preset end conditions, the original current video frame is displayed. For example, when a preset cancellation gesture is detected in the current video frame in the follow-up shooting mode, the picture-in-picture function is exited; or it can start timing when it is detected that the current video frame does not meet any preset trigger conditions. When the recognition results of a preset number of consecutive video frames do not meet any preset trigger conditions, the picture-in-picture function is automatically closed. For example, in the voice tracking mode, if there is no human voice, no human figure, and no gesture in the video frames for 10 consecutive seconds, and it does not meet any preset trigger conditions, the picture-in-picture function is exited.

[0081] Further, the present application can also store the received current configuration parameters and the corresponding target working mode in the parameter setting database; after generating the switching setting reminder, in response to the display instruction, determine and display the first picture and the second picture of the current video frame according to the current configuration parameters corresponding to the target working mode in the parameter setting database.

[0082] Specifically, there are two options for the switching setting reminder displayed after switching to the target working mode, and the other is the display instruction. If the user clicks the display option, the picture-in-picture display is performed according to the current configuration parameters in the parameter setting database.

[0083] In the specific implementation process, relatively important or serious meetings may not support the user to manually set the picture-in-picture display in the new mode when switching the video working mode. At this time, the historical data recorded in the database can be used for display.

[0084] In some embodiments, the method further includes:

[0085] Receiving a configuration parameter modification instruction; updating the parameter setting database according to the configuration parameter modification instruction.

[0086] Specifically, the configuration parameter modification instruction is input by the user through the APP, control tablet, remote control, etc. The user can modify the picture-in-picture configuration parameters of each video working mode before the meeting, so that there is no need to manually set the picture-in-picture display parameters when the video working mode is automatically switched during the meeting.

[0087] Please refer to Figure 2 , another embodiment of the present application provides a picture-in-picture display device for conference videos, including:

[0088] A receiving module 101, configured to receive picture-in-picture configuration parameters and a picture-in-picture start instruction.

[0089] An obtaining module 102, configured to obtain conference video data and corresponding voice data, and extract the current video frame.

[0090] A voice recognition module 103, configured to perform voice recognition on the voice data to obtain a voice recognition result.

[0091] A gesture recognition module 104, configured to perform gesture recognition on the current video frame to obtain a gesture recognition result.

[0092] A judgment module 105, configured to judge whether the voice recognition result and / or the gesture recognition result meet a preset trigger condition; if so, determine and display the first picture and the second picture of the current video frame according to the picture-in-picture configuration parameters.

[0093] In some embodiments, the device further includes:

[0094] A condition determination module, configured to obtain the current video working mode; and determine a preset trigger condition according to the current video working mode.

[0095] In some embodiments, the device further includes a mode switching module, configured to:

[0096] Determine whether the voice recognition result or the gesture recognition result meets the preset trigger condition of any video working mode; if so, use the video working mode corresponding to the satisfied preset trigger condition as the target working mode and switch; after switching to the target working mode, generate and display a switching setting reminder; in response to a setting instruction, display a parameter configuration interface; receive the current configuration parameters of the target working mode; and determine and display the first picture and the second picture of the current video frame according to the current configuration parameters.

[0097] In some embodiments, the device further includes a priority judgment module, configured to:

[0098] When the number of preset trigger conditions satisfied by the voice recognition result or the gesture recognition result is greater than 1, determine the target working mode and switch among the video working modes corresponding to the satisfied preset trigger conditions according to the preset mode priority.

[0099] In some embodiments, the device further includes a database setting module, configured to:

[0100] Store the received current configuration parameters and the corresponding target working mode into a parameter setting database; after generating the switching setting reminder, in response to a display instruction, determine and display the first picture and the second picture of the current video frame according to the current configuration parameters corresponding to the target working mode in the parameter setting database.

[0101] For the specific limitations of the picture-in-picture display device for a conference video provided in this embodiment, reference may be made to the embodiment of the picture-in-picture display method for a conference video in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned picture-in-picture display device for a conference video can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0102] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor executes the steps of a method for displaying a picture-in-picture of a conference video as described in any of the above embodiments.

[0103] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of the method for displaying a picture-in-picture of a conference video in the foregoing text, and details are not described herein again.

[0104] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a method for displaying a picture-in-picture of a conference video as described in any of the above embodiments. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the method for displaying a picture-in-picture of a conference video in the foregoing text, and details are not described herein again.

[0105] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0106] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0107] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for displaying a picture-in-picture of a conference video, characterized in that, including: Receiving picture-in-picture configuration parameters and a picture-in-picture start instruction; Obtaining conference video data and corresponding voice data, and extracting the current video frame; Performing speech recognition on the voice data to obtain a speech recognition result; Performing gesture recognition on the current video frame to obtain a gesture recognition result; Judging whether the speech recognition result and / or the gesture recognition result meet a preset trigger condition; If so, determining and displaying a first picture and a second picture of the current video frame according to the picture-in-picture configuration parameters.

2. The method for displaying a picture-in-picture of a conference video according to claim 1, wherein The picture-in-picture configuration parameters include a first picture type, a second picture position, and a second picture display ratio.

3. The method for displaying a picture-in-picture of a conference video according to claim 1, wherein The first picture and the second picture are different; at least one of them is a close-up portrait picture.

4. The method for displaying a picture-in-picture of a conference video according to claim 3, wherein, The person in the close-up portrait picture is the trigger person of the speech recognition result or the gesture recognition result that meets the preset trigger condition.

5. The method for displaying a picture-in-picture of a conference video according to claim 1, characterized in that, It also includes: Obtaining the current video working mode; determining the preset trigger condition according to the current video working mode.

6. The method for displaying a picture-in-picture of a conference video according to claim 5, wherein It also includes: Judging whether the speech recognition result or the gesture recognition result meets any of the preset trigger conditions of the video working mode; If so, taking the video working mode corresponding to the satisfied preset trigger condition as the target working mode and switching; After switching to the target working mode, generating and displaying a switching setting reminder; Responding to a setting instruction to display a parameter configuration interface; Receiving the current configuration parameters of the target working mode; Determining and displaying a first picture and a second picture of the current video frame according to the current configuration parameters.

7. The method for displaying a picture-in-picture of a conference video according to claim 6, wherein It also includes: If the number of preset trigger conditions satisfied by the speech recognition result or the gesture recognition result is greater than 1, determining and switching the target working mode among the video working modes corresponding to the various satisfied preset trigger conditions according to the preset mode priority.

8. The method for displaying a picture-in-picture of a conference video according to claim 6, wherein, It also includes: Storing the received current configuration parameters and the corresponding target working mode in a parameter setting database; After generating the switching setting reminder, in response to a display instruction, determining and displaying a first picture and a second picture of the current video frame according to the current configuration parameters corresponding to the target working mode in the parameter setting database.

9. The method for displaying a picture-in-picture of a conference video according to claim 8, wherein It also includes: Receiving a configuration parameter modification instruction; updating the parameter setting database according to the configuration parameter modification instruction.

10. A picture-in-picture display device for conference videos, characterized in that, including: A receiving module, used for receiving picture-in-picture configuration parameters and a picture-in-picture start instruction; An obtaining module, used for obtaining conference video data and corresponding voice data, and extracting the current video frame; A speech recognition module, used for performing speech recognition on the voice data to obtain a speech recognition result; A gesture recognition module, used for performing gesture recognition on the current video frame to obtain a gesture recognition result; A judging module, used for judging whether the speech recognition result and / or the gesture recognition result meet a preset trigger condition; If so, determining and displaying a first picture and a second picture of the current video frame according to the picture-in-picture configuration parameters.