Intelligent directing method, device and system for video conferencing
By identifying the key parts of participants in video conferences and building virtual positioning frames, and switching image modes based on speaking events, the problem of incomplete images in video conferences is solved, and the director's composition effect and user experience are improved.
Patent Information
- Application Number
- CN202210917222.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-08-01
AI Technical Summary
In existing video conferencing, the broadcasting capabilities of the venue's video images are insufficient, resulting in incomplete images of participants or speakers and poor composition ratios, affecting the user experience.
By identifying the key parts of the participants in the video image, building a virtual positioning frame, and combining it with the speaking event signal, it can flexibly switch between panoramic video and video close-up images, and adjust the composition to create the best panoramic or close-up picture.
It improves the directing composition effect and picture quality of video conferences, and enhances the user experience.
Smart Images

Figure CN115499615B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to an intelligent directing method, device, and system for video conferencing. Background Art
[0002] With the development of society, the application of video conferencing is becoming more and more extensive, and at the same time, the requirements for audio and video quality, data collaboration and sharing, flexibility and ease of use, and manageability are becoming more and more stringent.
[0003] Currently, during a video conference, the video images of the conference venue where the participants are located can be played and displayed. However, the video image playback mode of the video conference is too simple, and the picture construction of the played video image is poor, and it may be impossible to present the target participants in a complete and appropriate picture. Summary of the Invention
[0004] The purpose of one or more embodiments of this specification is to provide an intelligent directing method, device and system for video conferencing, so as to flexibly and intelligently switch between panoramic video images and video close-up images, and to adjust the construction of the best panoramic picture or the best close-up picture with a virtual positioning frame, thereby improving the directing composition effect and picture quality of the video conference, and thus improving the user experience.
[0005] To solve the above technical problems, one or more embodiments of this specification are implemented as follows:
[0006] First, an intelligent directing method for a video conference is proposed, including:
[0007] Obtain video images of at least two conference sites participating in the video conference;
[0008] Identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part;
[0009] Determining whether a first signal triggered by a participant's speech event is detected within a set time;
[0010] If the first signal is not detected within the set time, the virtual positioning frames of the multiple participants in each video image are adjusted as a whole in a closest framing mode, and transmitted to the target terminal for playback;
[0011] If a first signal is detected within a set time, a virtual positioning frame of the currently speaking participant is determined based on the first signal, a first video close-up image is constructed based on the virtual positioning frame and the current posture of the currently speaking participant, and the target terminal is triggered to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0012] Secondly, an intelligent video conferencing director device is proposed, comprising:
[0013] An acquisition module, configured to acquire video images of at least two conference sites participating in the video conference;
[0014] an identification module for identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part;
[0015] A judgment module, configured to judge whether a first signal triggered by a participant's speech event is detected within a set time;
[0016] a broadcast director module, configured to adjust the video image as a whole using a closest framing method by framing the virtual positioning frames of the multiple participants in each video image if the first signal is not detected within a set time, and transmit the video image to the target terminal for playback; and
[0017] If a first signal is detected within a set time, it is used to determine a virtual positioning frame of the currently speaking participant based on the first signal, and to construct a first video close-up image based on the virtual positioning frame and the current posture of the currently speaking participant, and to trigger the target terminal to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0018] Thirdly, an intelligent video conferencing directing system is proposed, including:
[0019] A target terminal located in each of the at least two conference venues, and the intelligent video conference director device located in the main conference venue according to the second aspect;
[0020] The intelligent directing device is used to execute the intelligent directing method described in the first aspect based on the acquired video images, so as to intelligently direct the video conference at the target terminal.
[0021] In a fourth aspect, an electronic device is provided, comprising:
[0022] processor; and
[0023] A memory arranged to store computer-executable instructions, which, when executed, cause the processor to execute the intelligent directing method for video conferencing described in the first aspect.
[0024] In the fifth aspect, a computer-readable storage medium is proposed, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple applications, the electronic device executes the intelligent directing method for video conferencing described in the first aspect.
[0025] It can be seen from the technical solutions provided by one or more embodiments of the above specification that, based on the virtual positioning frame constructed and bound to the participants in the acquired video image, combined with whether a speaking event is detected, it is selected to play the panoramic video image determined by the most recent framing method, or to switch to the virtual positioning frame of the currently speaking participant and the video close-up image constructed with the current posture, so that flexible and intelligent switching can be performed between the panoramic video image and the video close-up image, and the virtual positioning frame can be used to adjust and construct the best panoramic picture or the best close-up picture, thereby improving the director's composition effect and picture quality of the video conference, and thereby improving the user experience during the video conference. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the description of one or more embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 This is a scene diagram of the intelligent directing system for video conferencing provided in an embodiment of this specification.
[0028] Figure 2 This is one of the step diagrams of an intelligent directing method for video conferencing provided in an embodiment of this specification.
[0029] Figure 3a This is a schematic diagram of the principle of constructing a binding virtual positioning frame in the processing background provided by an embodiment of this specification.
[0030] Figure 3b and Figure 3c They are schematic diagrams of different playback windows displayed on a target terminal provided by an embodiment of this specification.
[0031] Figure 4 This is the second step diagram of a method for intelligent directing of a video conference provided by an embodiment of this specification.
[0032] Figure 5a-5f They are respectively the panoramic video image or the video close-up image displayed on the post-direction playback interface provided in the embodiments of this specification.
[0033] Figure 6 This is a structural diagram of an intelligent directing device for video conferencing provided by an embodiment of this specification.
[0034] Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the one or more embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.
[0036] Taking into account the current video conferencing's insufficient ability to direct the video images of the venue, there may be problems such as incomplete images of participants or speakers in the presented video images, poor composition ratios, etc., resulting in poor directing composition effects and picture quality of the video conference, affecting the user experience. To this end, the embodiment of this specification provides an intelligent directing solution for video conferencing to solve the above-mentioned problems. The idea of the present application solution is: based on the virtual positioning frame constructed and bound to the participants in the acquired video image, combined with whether a speaking event is monitored, to choose to play the panoramic video image determined by the most recent framing method, or to switch to the virtual positioning frame of the currently speaking participant and the video close-up image constructed with the current posture, so that flexible and intelligent switching can be performed between the panoramic video image and the video close-up image, and the virtual positioning frame can be used to adjust and construct the best panoramic picture or the best close-up picture, thereby improving the directing composition effect and picture quality of the video conference, and thus improving the user experience during the video conference.
[0037] Reference Figure 1As shown, it is a scene diagram of the intelligent broadcasting system for video conferencing provided in an embodiment of this specification. Assume that there are two venues participating in the video conference: venue 1 and venue 2; the intelligent broadcasting director 102 of the video conference can be located in venue 1 as a hardware device, and is communicatively connected to the target terminal 104 located in venue 1, and is communicatively connected to the target terminal 106 located in venue 2; both the target terminal 104 and the target terminal 106 can be installed with a client U for accessing the intelligent broadcasting director 102 of the video conference. At venue 1, the camera system in the hardware device in which the intelligent broadcasting director 102 of the video conference is integrated can be used to collect local video images of conference 1. In order to avoid missing images, video images of the venue in different directions can be collected by camera devices at multiple angles in the camera system. At venue 2, the camera system configured by the target terminal 106 itself can be used to collect local video images of venue 2.
[0038] The target terminal 106 sends the collected local video images of the conference room 2 to the intelligent broadcast director device 102 of the video conference. The intelligent broadcast director device 102 of the video conference will combine the locally collected video images to identify the key parts of the participants in the received video images and build a virtual positioning frame bound to the key part information of each participant. In this way, it can analyze and judge whether to play a panoramic video image or a close-up video image based on the signals triggered by different speaking events. In fact, the target terminal 104 and the target terminal 106 can respectively play videos through their own screens, or connect to other display devices for screen projection and playback. This manual does not limit the specific playback device.
[0039] In this way, you can Figure 1 The intelligent directing system for video conferencing shown realizes intelligent directing of video conferencing and can ensure that the video image transmitted to the target terminal is the best panoramic picture or the best close-up picture constructed by adjusting the virtual positioning frame, thereby improving the directing composition effect and picture quality of the video conference, thereby improving the user experience during the video conference.
[0040] It should be understood that in the embodiments of this specification, Figure 1 The intelligent broadcasting director device 102 of the video conference shown can be placed as a server in other spaces outside the conference venue. At this time, the video images obtained by the intelligent broadcasting director device 102 of the video conference can be collected by the camera system configured by the target terminal of each conference venue, that is, the intelligent broadcasting director device 102 of the video conference can be used to provide broadcasting director services instead of providing video image acquisition services. The implementation of other broadcasting director services remains unchanged and can also be referred to. Figure 1 The intelligent broadcast directing solution for video conferencing shown in the figure is implemented.
[0041] Reference Figure 2The figure shows a schematic diagram of the steps of an intelligent directing method for video conferencing provided in an embodiment of this specification. It should be understood that the execution subject of the intelligent directing method for video conferencing can be a hardware device with certain computing and processing capabilities, such as a smart phone, personal computer, wearable device, tablet computer, all-in-one conference machine, etc.; or a software device, such as a combination of software modules integrated on the aforementioned various hardware devices), specifically a server that can provide video conferencing directing services, such as a cloud server or other type of server. The intelligent directing method for video conferencing can include the following steps:
[0042] Step 202: Acquire video images of at least two conference sites participating in the video conference.
[0043] Based on the type of execution subject, the specific implementation of step 202 can be divided into different situations;
[0044] In scenario 1, the execution entity is at the local conference site. The execution entity's local camera system captures the local conference site's video image, while the target terminal's camera system at the other conference site captures the video image of the other conference site. Both the local conference site and the other conference sites are participating in the video conference.
[0045] Case 2: The execution subject is not at the conference site; video images of each conference site are acquired from the camera systems of target terminals at different conference sites.
[0046] It should be understood that in the embodiments of this specification, the camera system used to capture video images can be that of the executing entity or the target terminal. If it is the executing entity, the camera system can be a combination of cameras deployed at multiple locations in the venue, for example, different camera positions are set up at multiple locations and angles in the venue; if it is the target terminal, the camera system can be a camera on a laptop or desktop computer or a mobile phone. In the embodiments of this specification, the camera can be an ordinary camera or a camera with additional image processing capabilities such as a depth camera, and this specification does not limit this.
[0047] Step 204: Identify key part information of each participant in the video image, and based on the identified key part information, construct a virtual positioning frame in the video image that is bound to the key part information of each participant and at least covers the key part.
[0048] In an embodiment of the present specification, when step 204 is executed, the facial information of each participant in the video image can be identified based on face recognition technology; based on the identified facial information, a virtual positioning frame that at least covers the face is constructed for each participant in the video image, and a binding relationship is established between the facial information and the virtual positioning frame of the corresponding face.
[0049] In fact, after acquiring the video image, the key parts of the participants in the video image can be identified and located based on facial recognition or other image processing technologies, and the key part information can be obtained. Then, a virtual positioning frame that at least covers the key part is constructed at the key part, and the key part information is bound to the virtual positioning frame. The key part here can be the figure of the participant, the face of the participant, or the lips of the participant. The embodiments of this specification mainly use the face as an example to describe the key part in detail.
[0050] Reference Figure 3a As shown in FIG, a video image among at least two video images is obtained, and the faces of all participants in the video image can be identified and located according to the face recognition technology. Then, a virtual positioning frame is constructed at each face, that is, Figure 3a The dotted box in the middle actually doesn't need to be displayed on the video image. Instead, a partial image covering the area covered by the virtual positioning frame is determined in the background. The size of the virtual positioning frame is also related to the key area used for positioning. If the key area is the face, the size of the virtual positioning frame should at least enclose the face area; if the key area is the human figure, the size of the virtual positioning frame should at least enclose the area of the participant in the video image; if the key area is the lips, the size of the virtual positioning frame should only enclose the lip area of the participant.
[0051] Step 206 : Determine whether a first signal triggered by a participant's speech event is detected within a set time; if the first signal is not detected within the set time, execute step 208 ; otherwise, execute step 210 .
[0052] It should be understood that each venue is equipped with a sound pickup device to collect audio information from the venue while the camera system captures the video images of the venue. The sound pickup device can be located on the execution entity of the intelligent broadcasting method, on the target terminal, or as a wearable independent sound pickup device (such as a sound card). The sound pickup device can be, for example, a microphone or a sound pickup system composed of microphones.
[0053] The sound pickup device monitors the conference venue in real time to see if a participant is speaking. If so, it generates a first signal based on the audio data collected by the sound pickup device. Specifically, this monitoring and identification can be performed using Voice Activity Detection (VAD). The first signal may include audio data from the participant's speech, as well as echo data.
[0054] In fact, the first signal is not limited to audio data and can also be image data or a simple trigger signal. That is, the event of a participant speaking can be detected not only through audio data collected by the sound pickup device, but also through image data collected by a camera system, for example, by capturing lip movements using a conventional camera system or an infrared camera system. Furthermore, the first signal can be manually triggered by trigger buttons placed near each participant when they speak.
[0055] Therefore, in the embodiments of this specification, there are many ways for the intelligent director device to learn about the speaking events of participants, such as a microphone, an ordinary camera, an infrared camera, or a trigger button, and the implementation method is not limited here.
[0056] The set time may be an effective detection time determined according to the transmission delay without affecting normal playback, or may be a time for direct determination after acquisition, that is, the set time is 0.
[0057] It should be noted that, considering that the participants in the venue may sometimes make short cough sounds or interjections due to uncontrollable physiological reactions such as coughing or sneezing, in order to distinguish them from normal speeches, it can be set that when these events occur, the audio or image collection or the pressing of the trigger button will not be triggered; or, it can be further set that after the first signal is detected and the first signal lasts for a specific period of time, it is defaulted to be triggered as a normal participant speech event.
[0058] Step 208: The virtual positioning frames of the multiple participants in each video image are used as a whole to adjust the video image in a closest framing manner, and the video image is transmitted to the target terminal for playback.
[0059] Specifically, for each video image, the virtual positioning frame of each participant in the video image can be regarded as a whole, the video image can be cropped in a closest framing manner, and the cropped video image can be reduced or enlarged to adjust to the same size as the playback window of the target terminal.
[0060] It should be understood that the target terminal may have at least two playback windows. For example, two venues may be used as reference. Figure 3b As shown, there are two main play windows in the middle of the screen, one main play window plays the video image f1 of the local venue, and the other main play window plays the video image f2 of another venue; or, refer to Figure 3c As shown, there is a main playback window in the center of the screen and a secondary playback window below or above the screen. The main playback window plays video image f1 from another venue, and the secondary playback window plays video image f2 from the local venue. When no participant is speaking, the target terminal plays the video image adjusted in the most recent view mode.
[0061] The closest framing here means that the virtual positioning frames of all participants are regarded as a whole, and then expanded outward from the center of the whole until all participants are surrounded. In this way, the re-determined video image must include all participants, and all participants are presented as a whole as much as possible in the center of the picture, avoiding incomplete images of participants and poor composition.
[0062] Step 210: Determine a virtual positioning frame of the currently speaking participant based on the first signal, construct a first video close-up image based on the virtual positioning frame and the current posture of the currently speaking participant, and trigger the target terminal to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0063] A feasible solution is that when determining the virtual positioning frame of the currently speaking participant based on the first signal, the currently speaking participant can be located in the corresponding video image based on the voice information and / or image information carried in the first signal, and the virtual positioning frame of the currently speaking participant can be determined based on the positioning result.
[0064] Corresponding to the signal type in step 206, if the first signal carries voice information, the participant who is currently speaking can be located according to the echo signal in the first signal through the Direction Of Arrival (DOA) technology, and then the location can be mapped to the virtual positioning frame in the video image. If the first signal carries image information, the participant who is currently speaking can be located in the video image through picture comparison to determine the corresponding virtual positioning frame. If the first signal is a trigger signal, the virtual positioning frame of the participant who is currently speaking can be determined based on the correspondence between the trigger signal and the identification of each participant. In fact, the first signal can also contain several of the above-mentioned voice information, image information and trigger signals at the same time, so that accurate positioning can be achieved through multimodal fusion.
[0065] If the participant speech event is a single participant speech event, then based on the virtual positioning frame and current posture of the currently speaking participant, a partial video image containing at least the key parts of the currently speaking participant is extracted, and based on the partial video image, a first video close-up image containing the currently speaking participant is constructed.
[0066] If the participant speech event is a dialogue speech event between at least two participants, then based on the determined virtual positioning frame and the current posture of the participant corresponding to each virtual positioning frame, local video images containing at least the key parts of the participant are extracted respectively, and the at least two extracted local video images are fused to construct a first video close-up image containing at least two participants in the current dialogue speech.
[0067] A feasible solution, when constructing the first video close-up image based on the virtual positioning frame and the current posture of the participant currently speaking, may include the following steps:
[0068] Step 1: Determine the local video image of the participant currently speaking in the corresponding video image according to the virtual positioning frame.
[0069] Since the virtual positioning frame is not necessarily the entire image of the participant, the human-shaped image of the participant in the video image can be expanded outward from the current virtual positioning frame, i.e., a partial video image. The partial video image includes the entire exposed portion of the participant in the video image.
[0070] However, considering that some participants may overlap and be blocked at certain angles, the following methods can be used to determine the local video image:
[0071] If the determined virtual positioning frame overlaps with other virtual positioning frames, an overall positioning frame is generated based on the other virtual positioning frames that overlap with the virtual positioning frame and the virtual positioning frame, and a partial video image of the participant currently speaking in the corresponding video image is determined based on the video image of the area covered by the overall positioning frame;
[0072] If the determined virtual positioning frame does not overlap with other virtual positioning frames, a partial video image of the participant currently speaking in the corresponding video image is determined based on the video image of the area covered by the virtual positioning frame.
[0073] In this way, the complete human image of the participants can be determined as much as possible through the above method.
[0074] Step 2: Determine the matching style type based on the current posture of the participant currently speaking.
[0075] The postures of the participants can include at least sitting and standing. Among them, standing can further include static standing and dynamic standing. Therefore, different style types can be set in advance for each posture, for example, sitting corresponds to bust style, and standing corresponds to half-length style.
[0076] Step 3: crop and / or expand the local video image according to the determined style type to obtain a close-up image of the participant who is currently speaking.
[0077] In specific implementation, the cropped image can be infinitely digitally zoomed based on the image electronic pan-tilt technology to ensure the display versatility and smooth transition of the video image; when the cropped video image is lower than 1080p, the video image clarity can be enhanced through super-resolution algorithms and other methods, so that the image during the entire directing process is in high-definition state.
[0078] If the current posture of the participant currently speaking is sitting, the style type matching the participant is determined to be a bust style; a partial video image is cropped and / or expanded according to the determined bust style to obtain a close-up image of the bust of the participant currently speaking;
[0079] If the current posture of the participant currently speaking is standing, the style type matching the participant is determined to be a bust style; the local video image is cropped and / or expanded according to the determined bust style to obtain a close-up picture of the bust portrait of the participant currently speaking.
[0080] The bust close-up image is vertically divided into a bust blank area and a bust area; wherein the vertical height ratio of the bust blank area to the bust area is in the range of [1 / 6, 1 / 4]; and / or the bust close-up image is vertically divided into a bust blank area and a bust area; wherein the vertical height ratio of the bust blank area to the bust area is in the range of [1 / 8, 1 / 6].
[0081] Step 4: Select a matching image construction rule based on the current speaking mode, and construct the video image containing the close-up picture into a first video close-up image according to the selected image construction rule.
[0082] If the current speaking mode is a single person speaking, a single close-up is constructed as the first video close-up image based on the image construction rule. If the current speaking mode is a dialogue speaking, multiple combined close-ups are constructed as the first video close-up image based on the image construction rule.
[0083] In an embodiment of the present specification, when the target terminal is triggered to switch the currently played video image to the first video close-up image, the target terminal can be triggered to switch the video image played in the main playback window in the center of the screen to the first video close-up image. The best way is to retain a main playback window in the center of the screen for switching to play the first video image, and play the panoramic video images of the local venue and other venues in the sub-playback window at the bottom or top or other areas.
[0084] It should be understood that in the embodiments of this specification, the main playback window can be larger than the secondary playback window.
[0085] One possible way to do this is to refer to Figure 4 As shown, when playing the first video close-up image, the method further includes:
[0086] Step 212: Determine whether a second signal triggered by a movement event of the currently speaking participant is detected within a set time; if the second signal is not detected within the set time, keep the current playback image; otherwise, execute step 214.
[0087] In the embodiment of the present specification, the monitoring of the movement event of the participant who is currently speaking can be achieved by collecting image information through a camera system based on monitoring the speaking event.
[0088] Step 214: Construct a second video close-up image based on the virtual positioning frame and current posture of the current speaking participant tracked in real time, and switch to the newly constructed second video close-up image, wherein the second video close-up image includes a close-up picture of the current speaking participant during the movement.
[0089] After locating the current speaker through step 206, the position of the virtual positioning frame of the current speaker can be further tracked in real time through the image information in the second signal, and then a second video close-up image can be constructed in combination with the posture of the current speaker when moving, thereby triggering real-time switching to play the latest constructed second video close-up image.
[0090] A feasible method is to filter and process different video image sequences of moving and speaking participants to ensure the smoothness and accuracy of tracking, so that the directing effect will not be interfered with by other people.
[0091] The following example uses two venues (venue A and venue B) to access this video conference.
[0092] After the video conference starts, you can obtain video image 1 of site A and video image 2 of site B. Figure 5a As shown in the figure, the left picture is video image 1 of venue A, and the right picture is video image 2 of venue B. Based on the visual recognition algorithm, the facial image of each participant is identified from video image 1 and video image 2 respectively, and a virtual positioning frame corresponding to each facial image is constructed. In fact, it is equivalent to assigning a tracking ID to each participant, and then adjusting the participant corresponding to the identified facial image to the center of the picture in a closest framing manner. Figure 5bAs shown, the left image shows a panoramic video image played by a target terminal in venue A, and the right image shows a panoramic video image played by a target terminal in venue B. It should be understood that for ease of viewing, the panoramic video image of the opposite venue is generally played in the main playback window. When there is a need for viewing, the panoramic video image of the local venue can be played through a secondary playback window located in another area.
[0093] During a video conference, microphones detect that a participant is speaking in Site A. If the participant is speaking alone, the Direction of Access (DOA) algorithm locates participant M1, the current speaker. If the participant is in conversational mode, the DOA algorithm locates participants M2 and M3, both in the conversation. Participants M2 and M3 can be in the same or different sites.
[0094] In the single-speaking mode, the located participant M1 can be processed in the manner described in step 210 to obtain a close-up image of the participant M1. If the currently speaking participant M1 is sitting, the close-up image can be obtained by referring to Figure 5c As shown, participant M1 is presented in bust style, with a blank area of half head height at the top. The bust is about 2 head heights + half head height. If the current participant M1 who is speaking is standing, then the close-up picture can refer to Figure 5d As shown, participant M1 is presented in a bust style, with a blank area of half a head's height at the top, and the bust is approximately 3 head heights + half a head's height in size.
[0095] In the dialogue mode, the located participants M2 and M3 can be processed in a close-up manner as described in step 210 above, to obtain close-up images of the participants M2 and M3. Figure 5e As shown in the image above, participant M2 is always seated during the conversation, so the close-up image shows a bust-up portrait. Participant M3 is always standing, so the close-up image shows a bust-up portrait. This allows the left-right layout to present the two participants in different styles within the same close-up image, ensuring optimal composition and proper presentation of the speakers.
[0096] During the speech process, whether in single speech mode or dialogue speech mode, the camera can identify whether the current speaker moves, for example, walking from his seat to the screen to give a PPT presentation. When movement is confirmed, the close-up image of the moving participant is continuously tracked. For example, if Figure 5cIf participant M1 stands up and moves during his speech, the close-up image should continue to track the close-up of participant M1 during his speech and movement. It should be understood that the close-up image at this time has switched the close-up style from bust style to half-length style. Figure 5e If the participant M3 moves during the conversation, the close-up image of the participant M3 in the close-up image should maintain the bust style and continue to track the close-up image of the participant M3.
[0097] If in Figure 5c In the single speaker mode shown, if participant M4 from the local venue joins the discussion with participant M3, close-up processing is performed in accordance with the processing method in the dialogue speaking mode, and participant M3 and participant M4 are arranged in the same close-up video.
[0098] If in Figure 5c In the single speaker mode shown in FIG, participant M3 is blocked by participant M5 who has not spoken. Then, referring to Figure 5f , the participants M3 and M5 can be regarded as a whole, the overall positioning frame can be re-determined, and then the video close-up image can be constructed according to the method of step 210. Thus, it is ensured that the presented speakers are reasonable and complete.
[0099] Through the above technical solution, based on the virtual positioning frame constructed and bound to the participants in the acquired video image, combined with whether a speaking event is detected, it is selected to play the panoramic video image determined by the most recent framing method, or switch to the virtual positioning frame of the currently speaking participant and the video close-up image constructed with the current posture, so that flexible and intelligent switching can be performed between the panoramic video image and the video close-up image, and the virtual positioning frame can be used to adjust and construct the best panoramic picture or the best close-up picture, thereby improving the director's composition effect and picture quality of the video conference, thereby improving the user experience during the video conference.
[0100] Example 2
[0101] Reference Figure 6 As shown, an intelligent video conferencing director device 600 provided in an embodiment of this specification includes:
[0102] An acquisition module 602 is configured to acquire video images of at least two conference sites participating in the video conference;
[0103] an identification module 604 for identifying key part information of each participant in the video image, and constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part based on the identified key part information;
[0104] A determination module 606 is configured to determine whether a first signal triggered by a participant's speech event is detected within a set time period;
[0105] The broadcast director module 608 is configured to adjust the virtual positioning frames of the multiple participants in each video image as a whole in a closest framing manner if the first signal is not detected within the set time, and transmit the video image to the target terminal for playback; and
[0106] If a first signal is detected within a set time, a virtual positioning frame of the currently speaking participant is determined based on the first signal, a first video close-up image is constructed based on the virtual positioning frame and the current posture of the currently speaking participant, and the target terminal is triggered to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0107] Optionally, as an embodiment,
[0108] The judgment module 606 is further configured to judge whether a second signal triggered by a movement event of the currently speaking participant is detected within a set time;
[0109] The broadcast director module 608 is further configured to maintain the current playback image if the second signal is not detected within a set time; and
[0110] If a second signal is detected within the set time, it is used to construct a second video close-up image based on the virtual positioning frame and current posture of the current speaking participant tracked in real time, and switch to the latest constructed second video close-up image, wherein the second video close-up image includes a close-up picture of the current speaking participant during the movement.
[0111] In a specific implementation of an embodiment of the present specification, when the director module 608 determines the virtual positioning frame of the currently speaking participant based on the first signal, it is specifically used to locate the currently speaking participant in the corresponding video image based on the voice information and / or image information carried in the first signal, and determine the virtual positioning frame of the currently speaking participant based on the positioning result.
[0112] In another specific implementation of the embodiments of this specification, when the director module 608 constructs the first video close-up image based on the virtual positioning frame and the current posture of the currently speaking participant, it is specifically used to determine the local video image of the currently speaking participant in the corresponding video image based on the virtual positioning frame; determine the matching style type based on the current posture of the currently speaking participant; crop and / or expand the local video image according to the determined style type to obtain a close-up picture of the currently speaking participant; select a matching image construction rule based on the current speaking mode, and construct the video image containing the close-up picture into the first video close-up image according to the selected image construction rule.
[0113] In another specific implementation of the embodiments of this specification, when the director module 608 determines the partial video image of the currently speaking participant in the corresponding video image based on the virtual positioning frame, if the determined virtual positioning frame overlaps with other virtual positioning frames, an overall positioning frame is generated based on the other virtual positioning frames that overlap with the virtual positioning frame and the virtual positioning frame, and the partial video image of the currently speaking participant in the corresponding video image is determined based on the video image of the area covered by the overall positioning frame; if the determined virtual positioning frame does not overlap with other virtual positioning frames, the partial video image of the currently speaking participant in the corresponding video image is determined based on the video image of the area covered by the virtual positioning frame.
[0114] In another specific implementation of the embodiments of this specification, the director module 608 determines a matching style type based on the current posture of the participant currently speaking; crops and / or expands the local video image according to the determined style type to obtain a close-up picture of the participant currently speaking, if the current posture of the participant currently speaking is sitting, the style type matching the participant is determined to be a bust style; crops and / or expands the local video image according to the determined bust style to obtain a close-up picture of the bust of the participant currently speaking; if the current posture of the participant currently speaking is standing, the style type matching the participant is determined to be a bust style; crops and / or expands the local video image according to the determined bust style to obtain a close-up picture of the bust of the participant currently speaking.
[0115] In another specific implementation of the embodiments of the present specification, the bust close-up image is divided vertically into a bust blank area and a bust area; wherein the vertical height ratio of the bust blank area to the bust area is in the range of [1 / 6, 1 / 4], and / or, the bust close-up image is divided vertically into a bust blank area and a bust area; wherein the vertical height ratio of the bust blank area to the bust area is in the range of [1 / 8, 1 / 6].
[0116] In another specific implementation of the embodiments of this specification, when the director module 608 constructs the first video close-up image based on the virtual positioning frame and the current posture of the participant who is currently speaking: if the participant speech event is a single participant speech event, then based on the virtual positioning frame and the current posture of the participant who is currently speaking, a local video image containing at least the key parts of the participant who is currently speaking is extracted, and a first video close-up image containing the participant who is currently speaking is constructed based on the local video image; if the participant speech event is a dialogue speech event of at least two participants, then based on the determined virtual positioning frame and the current posture of the participant corresponding to each virtual positioning frame, local video images containing at least the key parts of the participant are extracted respectively, and the at least two extracted local video images are fused to construct a first video close-up image containing at least two participants who are currently speaking in dialogue.
[0117] In another specific implementation of the embodiments of this specification, the recognition module is specifically used to recognize the facial information of each participant in the video image based on face recognition technology when recognizing the key part information of each participant in the video image, and based on the recognized key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and at least covers the key part; based on the recognized facial information, constructing a virtual positioning frame in the video image that at least covers the face for each participant, and establishing a binding relationship between the facial information and the virtual positioning frame corresponding to the face.
[0118] The intelligent directing device of the video conference can choose to play the panoramic video image determined by the most recent framing method based on the virtual positioning frame constructed and bound to the participants in the acquired video image, combined with whether a speaking event is detected, or switch to the video close-up image constructed with the virtual positioning frame and current posture of the currently speaking participant. In this way, it can flexibly and intelligently switch between the panoramic video image and the video close-up image, and can use the virtual positioning frame to adjust and construct the best panoramic picture or the best close-up picture, thereby improving the directing composition effect and picture quality of the video conference, and thereby improving the user experience during the video conference.
[0119] Example 3
[0120] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this specification. Figure 7At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0121] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0122] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0123] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a translation model compression device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0124] Obtain video images of at least two conference sites participating in the video conference;
[0125] Identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part;
[0126] Determining whether a first signal triggered by a participant's speech event is detected within a set time;
[0127] If the first signal is not detected within the set time, the virtual positioning frames of the multiple participants in each video image are adjusted as a whole in a closest framing mode, and transmitted to the target terminal for playback;
[0128] If a first signal is detected within a set time, a virtual positioning frame of the currently speaking participant is determined based on the first signal, a first video close-up image is constructed based on the virtual positioning frame and the current posture of the currently speaking participant, and the target terminal is triggered to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0129] The above is as in this manual Figure 2 and Figure 4 The methods performed by the apparatus disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in one or more embodiments of this specification can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with one or more embodiments of this specification can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0130] The electronic device may also perform Figure 2 and Figure 4 method, and implement the corresponding device in Figure 2 and Figure 4 The functions of the embodiments shown in this specification will not be described in detail here.
[0131] Of course, in addition to software implementation, the electronic device of the embodiments of this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0132] The embodiment of this specification also proposes a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by a portable electronic device including multiple application programs, can enable the portable electronic device to execute Figure 2 and Figure 4 The method of the embodiment shown is specifically used to perform the following method:
[0133] Obtain video images of at least two conference sites participating in the video conference;
[0134] Identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part;
[0135] Determining whether a first signal triggered by a participant's speech event is detected within a set time;
[0136] If the first signal is not detected within the set time, the virtual positioning frames of the multiple participants in each video image are adjusted as a whole in a closest framing mode, and transmitted to the target terminal for playback;
[0137] If a first signal is detected within a set time, a virtual positioning frame of the currently speaking participant is determined based on the first signal, a first video close-up image is constructed based on the virtual positioning frame and the current posture of the currently speaking participant, and the target terminal is triggered to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the currently speaking participant.
[0138] In short, the above description is only a preferred embodiment of this specification and is not intended to limit the scope of protection of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification shall be included in the scope of protection of this specification.
[0139] The systems, devices, modules, or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0140] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0141] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0142] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0143] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. An intelligent directing method for a video conference, comprising: Obtain video images captured by camera systems of at least two conference sites participating in the video conference; Identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part; Determining whether a first signal triggered by a participant's speech event is detected within a set time; If the first signal is not detected within the set time, the video image is adjusted by taking the virtual positioning frame of the multiple participants in each video image as a whole, expanding outward from the center of the whole until it surrounds all the participants, and transmitting it to the target terminal for playback; If the first signal is detected within the set time, a virtual positioning frame of the participant currently speaking is determined based on the first signal, a partial video image of the participant currently speaking in the corresponding video image is determined based on the virtual positioning frame, a first video close-up image is constructed based on the current posture of the participant currently speaking and the partial video image, and the target terminal is triggered to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up of the participant currently speaking; Among them, determining the local video image of the participant who is currently speaking in the corresponding video image based on the virtual positioning frame includes: if the determined virtual positioning frame overlaps with other virtual positioning frames, generating an overall positioning frame based on other virtual positioning frames that overlap with the virtual positioning frame and the virtual positioning frame, and determining the local video image of the participant who is currently speaking in the corresponding video image based on the video image of the area covered by the overall positioning frame; if the determined virtual positioning frame does not overlap with other virtual positioning frames, determining the local video image of the participant who is currently speaking in the corresponding video image based on the video image of the area covered by the virtual positioning frame.
2. The intelligent broadcast directing method according to claim 1, wherein during the playback of the first video close-up image, the method further comprises: Determine whether a second signal triggered by a movement event of the currently speaking participant is detected within a set time; If the second signal is not detected within the set time, the current playback picture is maintained; If a second signal is detected within the set time, a second video close-up image is constructed based on the virtual positioning frame and current posture of the current speaking participant tracked in real time, and the second video close-up image is switched to the latest constructed second video close-up image, wherein the second video close-up image includes a close-up picture of the current speaking participant during the movement.
3. The intelligent broadcast directing method according to claim 1, wherein determining a virtual positioning frame of the participant currently speaking based on the first signal comprises: Based on the voice information and / or image information carried in the first signal, the participant currently speaking is located in the corresponding video image, and a virtual positioning frame of the participant currently speaking is determined based on the positioning result.
4. The intelligent directing method according to claim 1, wherein the first video close-up image is constructed based on the current posture of the participant currently speaking and the partial video image, comprising: Determine the matching style type based on the current posture of the participant currently speaking; Performing cropping and / or expansion processing on a partial video image according to the determined style type to obtain a close-up image of the participant currently speaking; A matching image construction rule is selected based on the current speaking mode, and the video image containing the close-up picture is constructed into a first video close-up image according to the selected image construction rule.
5. The intelligent directing method according to claim 4, wherein the matching style type is determined based on the current posture of the participant currently speaking; The partial video image is cropped and / or expanded according to the determined style type to obtain a close-up image of the participant currently speaking, including: If the current posture of the participant currently speaking is a sitting posture, then determining that the style type matching the participant is a bust style; Performing cropping and / or expansion processing on a partial video image according to the determined bust style to obtain a close-up bust image of the participant currently speaking; If the current posture of the participant currently speaking is standing, the style type matching the participant is determined to be a bust style; the local video image is cropped and / or expanded according to the determined bust style to obtain a close-up picture of the bust portrait of the participant currently speaking.
6. The intelligent broadcasting method according to claim 5, wherein the bust close-up image is vertically divided into a bust blank area and a bust area; The vertical height ratio of the bust blank area to the bust area is [1 / 6, 1 / 4]; and / or, The bust close-up image is vertically divided into a bust blank area and a bust area; wherein the vertical height ratio of the bust blank area to the bust area is in the range of [1 / 8, 1 / 6].
7. The intelligent directing method as described in any one of claims 1 to 6, wherein if the current speaking mode is a single participant speaking, the first video close-up image includes a close-up picture of the single participant; if the current speaking mode is a dialogue between at least two participants, the first video close-up image includes close-up pictures of at least two participants.
8. The intelligent directing method according to any one of claims 1 to 6, wherein the method further comprises: identifying key part information of each participant in the video image, and constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part based on the identified key part information, comprising: Identifying facial information of each participant in the video image based on face recognition technology; Based on the recognized facial information, a virtual positioning frame that at least covers the face is constructed for each participant in the video image, and a binding relationship is established between the facial information and the virtual positioning frame corresponding to the face.
9. An intelligent video conferencing director device, comprising: An acquisition module, configured to acquire video images captured by camera systems of at least two conference sites participating in the video conference; an identification module for identifying key part information of each participant in the video image, and based on the identified key part information, constructing a virtual positioning frame in the video image that is bound to the key part information of each participant and covers at least the key part; A judgment module, configured to judge whether a first signal triggered by a participant's speech event is detected within a set time; The broadcast director module is configured to adjust the video image by taking a virtual positioning frame of multiple participants in each video image as a whole, expanding the virtual positioning frame from the center of the whole until it surrounds all participants, and transmit the virtual positioning frame to the target terminal for playback if the first signal is not detected within the set time; as well as, If a first signal is detected within a set time, the device is used to determine a virtual positioning frame of the participant currently speaking based on the first signal, and determine a partial video image of the participant currently speaking in the corresponding video image based on the virtual positioning frame, construct a first video close-up image based on the current posture of the participant currently speaking and the partial video image, and trigger the target terminal to switch the currently playing video image to the first video close-up image, wherein the first video close-up image includes a close-up picture of the participant currently speaking; Among them, determining the local video image of the participant who is currently speaking in the corresponding video image based on the virtual positioning frame includes: if the determined virtual positioning frame overlaps with other virtual positioning frames, generating an overall positioning frame based on other virtual positioning frames that overlap with the virtual positioning frame and the virtual positioning frame, and determining the local video image of the participant who is currently speaking in the corresponding video image based on the video image of the area covered by the overall positioning frame; if the determined virtual positioning frame does not overlap with other virtual positioning frames, determining the local video image of the participant who is currently speaking in the corresponding video image based on the video image of the area covered by the virtual positioning frame.
10. An intelligent video conferencing directing system, comprising: A target terminal located in each of at least two conference venues, and the intelligent video conference directing device according to claim 9; The intelligent directing device is used to execute the intelligent directing method according to any one of claims 1 to 8 based on the acquired video images, so as to intelligently direct the video conference at the target terminal.
11. An electronic device comprising: processor; as well as A memory arranged to store computer-executable instructions, which, when executed, cause the processor to execute the intelligent directing method for video conferencing as described in any one of claims 1 to 8.
12. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device comprising a plurality of application programs, enables the electronic device to execute the intelligent directing method for video conferencing as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Video conferencing system, processing device and video conferencing method
CN105592268A
Wireless video conferencing system
CN109068090A
Speaker positioning method based on sound and image fusion
CN111046850A