Video conference radio switching method and video conference system
By identifying the location and behavioral events of participants in the video conference and automatically adjusting the radio range, the problem that participants need to switch microphone mode by themselves to avoid the meeting being disturbed is solved, and the efficient radio reception and overall efficiency of the video conference are improved.
Patent Information
- Application Number
- CN202311512905.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
During video conferencing, participants need to switch microphone modes themselves to avoid hearing irrelevant comments from remotely other participants, causing the meeting to be disturbed.
By identifying the attendee position and behavioral events in the video signal, the audio range of the radio device is automatically adjusted to achieve intelligent switching between the attendees and the sound filter.
It effectively avoids attendees being radioed without wanting to be heard by others, improving the audio quality of video conferences and the overall efficiency of the conference.
Smart Images

Figure CN120017904A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a video conferencing system and a method for using the same, and more particularly to a method for switching audio reception in a video conferencing system and a video conferencing system. Background Art
[0002] With the development of the Internet, the usage rate of many online conference software has also increased significantly. People can hold video conferences with remote users without moving around. Generally speaking, in video conferences, participants need to switch the microphone mode to mute or receive. If the participant only needs to speak temporarily (such as answering a call or discussing with other participants on the scene) and does not need to be heard by other remote participants, if the participant forgets to switch the microphone to mute mode, other remote participants will hear irrelevant remarks and disturb the meeting. Summary of the invention
[0003] Based on this, it is necessary to provide a method for switching audio reception in a video conference and a video conference system, which can automatically switch between audio reception and audio filtering according to the actions of the participants.
[0004] A method for switching audio reception of a video conference, using a processor to execute the following steps when starting a video conference, including:
[0005] Obtaining a relative position of each of a plurality of participants in a conference space and a behavior event of each of the participants by identifying a video signal;
[0006] Based on the behavior event of each participant, determining whether each participant is in a non-speaking behavior;
[0007] In response to one of the participants being determined to be a non-speaker in the non-speaking behavior, adjusting a sound pickup range of a sound pickup device based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker; and
[0008] In response to one of the participants being determined to be a talker who is not in the non-speaking behavior, the sound pickup range of the sound pickup device is adjusted based on the relative position of the talker in the conference space to pick up the sound of the talker.
[0009] In one embodiment, obtaining a relative position of each of a plurality of participants in a conference space and a behavior event of each of the participants by identifying a video signal includes:
[0010] Identify each of the participants included in the video signal through a portrait recognition module, and obtain the relative position of each of the participants in the conference space; and
[0011] The behavior event of each participant is identified through an image recognition module.
[0012] In one embodiment, it also includes:
[0013] After receiving an audio signal, a human voice is separated from the audio signal through a voiceprint recognition module;
[0014] matching the human voice to one of the corresponding participants; and
[0015] Based on the human voice and the matching behavior event of one of the participants, it is determined whether one of the participants is in the non-speaking behavior.
[0016] In one embodiment, the step of matching the human voice with one of the participants corresponding thereto includes:
[0017] Executing a sound source localization algorithm to determine a source position of the human voice in the conference space; and
[0018] Based on the relative position and the source position, the human voice is matched with one of the participants corresponding thereto.
[0019] In one embodiment, the step of determining whether one of the participants is in the non-speaking behavior based on the human voice and the matching behavior event of one of the participants includes:
[0020] Identifying a sound intensity of the human voice through an audio recognition module;
[0021] Determining whether the sound intensity is less than a default value;
[0022] Determining whether the behavior event of one of the participants matching the human voice conforms to a preset event; and
[0023] When the behavior event conforms to the preset event and the sound intensity is less than the default value, it is determined that one of the participants is in the non-speaking behavior.
[0024] In one embodiment, after initiating the video conference through the processor, the method further includes:
[0025] Capturing images at intervals of a sampling time by an imaging device to obtain the video signal; and
[0026] A sound receiving device is used to receive sound at intervals of the sampling time to obtain the audio signal.
[0027] In one embodiment, it also includes:
[0028] A user interface is displayed on a display device through the processor, and the video segment corresponding to each participant is displayed in the user interface.
[0029] In one embodiment, the user interface includes a plurality of display areas, each of which corresponds to each participant, and each of which is used to display the video clip corresponding to each participant;
[0030] After the user interface is displayed on the display device through the processor, the method further includes:
[0031] Displaying a mute mark in one of the display areas displaying the video segment corresponding to the non-speaker; and
[0032] A sound receiving mark is displayed in another one of the display areas that displays the video segment corresponding to the speaker.
[0033] In one embodiment, the step of filtering the non-speaker includes:
[0034] The processor filters out the human voice corresponding to the relative position of the non-speaker in the conference space from the audio signal.
[0035] A video conferencing system, comprising:
[0036] a storage device storing an application program;
[0037] An imaging device for acquiring a video signal;
[0038] a radio receiving device; and
[0039] A processor is coupled to the storage, the imaging device and the sound receiving device, and is configured to:
[0040] Executing the application to start a video conference, and initiating the video conference, comprising:
[0041] Obtaining a relative position of each of a plurality of participants in a conference space and a behavior event of each of the participants by identifying a video signal;
[0042] Based on the behavior event of each participant, determining whether each participant is in a non-speaking behavior;
[0043] In response to one of the plurality of participants being determined to be a non-speaker in the non-speaking behavior, adjusting a sound pickup range of a sound pickup device based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker; and
[0044] In response to one of the multiple participants being determined to be a talker who is not in the non-speaking behavior, the sound collection range of the sound collection device is adjusted based on the relative position of the talker in the conference space to collect the sound of the talker.
[0045] In one embodiment, the storage includes a portrait recognition module and an image recognition module, and the processor is configured to:
[0046] Identify each participant included in the video signal through the portrait recognition module, and obtain the relative position of each participant in the conference space;
[0047] The behavior event of each participant is identified through the image recognition module.
[0048] In one embodiment, the sound receiving device is used to obtain an audio signal;
[0049] The storage device includes a voiceprint recognition module, and the processor is configured to:
[0050] Separating a human voice from the audio signal through the voiceprint recognition module;
[0051] matching the human voice to one of the corresponding participants; and
[0052] Based on the human voice and the matching behavior event of one of the participants, it is determined whether one of the participants is in the non-speaking behavior.
[0053] In one embodiment, the processor is configured to:
[0054] Executing a sound source localization algorithm to determine a source position of the human voice in the conference space; and
[0055] Based on the relative position and the source position, the human voice is matched with one of the participants corresponding thereto.
[0056] In one embodiment, wherein the storage includes an audio recognition module, the processor is configured to:
[0057] Recognize a sound intensity of the human voice through the audio recognition module;
[0058] Determining whether the sound intensity is less than a default value;
[0059] Determining whether the behavior event of one of the participants matching the human voice conforms to a preset event;
[0060] When the behavior event conforms to the preset event and the sound intensity is less than the default value, it is determined that one of the participants is in the non-speaking behavior.
[0061] In one embodiment, the processor is configured to:
[0062] The image capturing device captures images at intervals of a sampling time to obtain the video signal; and
[0063] The sound receiving device receives sound at intervals of the sampling time to obtain the audio signal.
[0064] In one embodiment, the processor is configured to:
[0065] By identifying the video signal, identifying each of the participants included in the video signal, and obtaining a video segment corresponding to each of the participants from the video signal; and
[0066] A user interface is displayed on a display device, and the video segment corresponding to each participant is displayed in the user interface.
[0067] In one embodiment, the user interface includes a plurality of display areas, each of which corresponds to each participant, and each of which is used to display the video clip corresponding to each participant;
[0068] wherein the processor is configured to:
[0069] Displaying a mute mark in one of the display areas displaying the video segment corresponding to the non-speaker; and
[0070] A sound receiving mark is displayed in another one of the display areas that displays the video segment corresponding to the speaker.
[0071] In one embodiment, the processor is configured to:
[0072] Based on the relative position of the non-speaker in the conference space, the human voice corresponding to the relative position in the audio signal is filtered out.
[0073] The above-mentioned video conference system includes: a storage device storing an application; an imaging device for acquiring a video signal; a sound receiving device; and a processor coupled to the storage device, the imaging device, and the sound receiving device. The processor is configured to: execute the application to start the video conference, and when the video conference is started, the processor includes: obtaining the relative positions of multiple participants in the conference space and the behavior events of each participant by identifying the video signal; judging whether each participant is in a non-speaking behavior based on the behavior events of each participant; in response to one of the participants being determined to be a non-speaker in a non-speaking behavior, adjusting the sound receiving range of the sound receiving device based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker; and in response to one of the participants being determined to be a speaker who is not in a non-speaking behavior, adjusting the sound receiving range of the sound receiving device based on the relative position of the speaker in the conference space to receive the sound of the speaker.
[0074] Based on the above, this method can identify the relative position of each participant in the conference space in the video signal, and decide whether to filter or collect the sound based on whether each participant is in a non-speaking behavior. Accordingly, it can prevent participants from collecting sound when they do not want others to hear them, thereby improving the sound quality of the video conference. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the conventional technology, the drawings required for use in the embodiments or the conventional technology descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0076] Figure 1 This is one of the structural diagrams of a video conferencing system according to an embodiment;
[0077] Figure 2 is a flow chart of a method for switching audio reception of a video conference according to an embodiment;
[0078] Figure 3 This is a second structural diagram of a video conferencing system according to an embodiment;
[0079] Figure 4 is a flow chart of a method for switching audio reception of a video conference according to an embodiment;
[0080] Figure 5A to Figure 5E is a schematic diagram of multiple preset events of an embodiment;
[0081] Figure 6 is a schematic diagram of an application scenario of a video conference according to an embodiment;
[0082] Figure 7 is a schematic diagram of a user interface of an embodiment.
[0083] Description of reference numerals:
[0084] 100: Video Conferencing System
[0085] 110: Processor
[0086] 120: Storage
[0087] 121: Application
[0088] 130: Imaging device
[0089] 140: Radio Device
[0090] 150: Display device
[0091] 210: Audio source identification module
[0092] 211: Voiceprint recognition module
[0093] 213: Portrait recognition module
[0094] 220: Speech behavior recognition module
[0095] 221: Audio recognition module
[0096] 223: Image recognition module
[0097] 230: Radio switching module
[0098] 700: User interface
[0099] 710~740: Display area
[0100] C1~C4: Video clips
[0101] M1~M3: Radio mark
[0102] M4: Mute Marker
[0103] U1~U4: Participants
[0104] V1: Region
[0105] S21~S25: Steps for switching the audio of a video conference
[0106] S404-S419: Steps for switching the audio of a video conference DETAILED DESCRIPTION
[0107] In order to facilitate understanding of the present application, the present application will be described more fully below with reference to the relevant drawings. Embodiments of the present application are provided in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0108] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0109] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first resistor may be referred to as a second resistor, and similarly, a second resistor may be referred to as a first resistor. Both the first resistor and the second resistor are resistors, but they are not the same resistor.
[0110] It can be understood that the “connection” in the following embodiments should be understood as “electrical connection”, “communication connection”, etc. if the connected circuits, modules, units, etc. have electrical signals or data transmission between each other.
[0111] It can be understood that “at least one” means one or more, “plurality” means two or more, and “at least a portion of an element” means a part or all of an element.
[0112] When used herein, the singular forms "a", "an", and "said / the" may also include plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include / comprise" or "have" and the like specify the presence of stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not exclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof. At the same time, the term "and / or" used in this specification includes any and all combinations of the relevant listed items.
[0113] Figure 1 This is one of the structural diagrams of a video conferencing system according to an embodiment. Figure 1 The video conference system 100 includes a processor 110, a storage 120, an imaging device 130, and a sound receiving device 140. The processor 110 is coupled to the storage 120, the imaging device 130, and the sound receiving device 140.
[0114] The processor 110 is, for example, a central processing unit (CPU), a physical processing unit (PPU), a programmable microprocessor, an embedded control chip, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or other similar devices.
[0115] The storage 120 is, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk or other similar device or a combination of these devices. The storage 120 stores one or more code segments, which are executed by the processor 110 after being installed. In this embodiment, the storage 120 includes an application 121 for executing a video conference. When the video conference is started, the processor 110 will execute the following method for switching the audio of the video conference.
[0116] The imaging device 130 may be a video camera or a still camera using a charge coupled device (CCD) lens or a complementary metal oxide semiconductor transistor (CMOS) lens. For example, the imaging device 130 may be a wide-angle camera, a half-spherical camera, a full-spherical camera, or the like.
[0117] The sound receiving device 140 is, for example, a microphone. In one embodiment, only one sound receiving device 140 may be provided. In other embodiments, multiple sound receiving devices 140 may also be provided.
[0118] Figure 2 This is a flow chart of a method for switching audio recording of a video conference according to an embodiment. Figure 1 and Figure 2 In step S21, the processor 110 executes the application 121 to start a video conference.
[0119] When a video conference is started, in step S22, the relative positions of multiple participants in the conference space and the behavioral events of each participant are obtained by identifying the video signal. Assuming that the shooting angle of the imaging device 130 covers all participants in the conference space (such as a conference room, office, lounge, etc.), the video signal captured by the imaging device 130 in the video conference will include the portraits of all participants. Afterwards, the processor 110 executes an image recognition algorithm to find each portrait in the video image, and calculates the relative position of each participant in the conference space through the position of these portraits in each image frame. In addition, the processor 110 executes an image recognition algorithm to identify the behavioral events of each participant. For example, it can be set to detect the gestures corresponding to each portrait, detect the condition of the face of each portrait being covered, detect the distance change between each portrait in multiple image frames, etc., to obtain the behavioral events of each participant.
[0120] Next, in step S23, based on the behavior events of each participant, it is determined whether each participant is in a non-speaking behavior. The non-speaking behavior is defined as: during the video conference, the participant speaks privately and does not want the other party of the video conference to hear, that is, the non-speaking behavior that needs to be filtered. For example, a plurality of preset events belonging to non-speaking behavior can be defined in advance in the storage 120, so that the processor 110 can compare the identified behavior events with the pre-defined preset events to determine whether the participant is in a non-speaking behavior.
[0121] In one embodiment, the preset event may include at least one of a mouth covering event, a palm back facing forward event, and an event of approaching a bystander. In one embodiment, the mouth covering event may be defined as: a portion or all of the mouth area is covered; the palm back facing forward event is defined as: the back of the palm is detected; the event of approaching a bystander is defined as: the distance between the portraits of two participants is less than a default value, which is determined as an event of approaching a bystander. The above embodiments are only examples and are not limited thereto.
[0122] In response to one of the participants being determined to be a non-speaker in a non-speaking behavior, in step S24, the processor 110 adjusts the sound receiving range of the sound receiving device 140 based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker. In one embodiment, the processor 110 filters out the human voice corresponding to the relative position in the audio signal based on the relative position of the non-speaker in the conference space. For example, the corresponding human voice is found according to the relative position of the non-speaker in the conference space, and based on the sound intensity, the sound receiving range of the sound receiving device 140 is adjusted to ensure that the sound of the non-speaker is not recorded.
[0123] There may be only one sound receiving device 140 , or a plurality of sound receiving devices 140 may be provided to form a sound receiving system, which is not limited herein.
[0124] In response to one of the participants being determined as a speaker who is not in a non-speaking behavior, in step S25, the processor 110 adjusts the sound collection range of the sound collection device 140 based on the relative position of the speaker in the conference space to collect the sound of the speaker. In one embodiment, the processor 110 retains the human voice corresponding to the relative position in the audio signal based on the relative position of the speaker in the conference space.
[0125] In addition, in other embodiments, a manual setting function may also be provided in the application 121. Accordingly, the user can manually set the sound collection range of the sound collection device 140 through the manual setting function.
[0126] Figure 3 This is a second structural diagram of a video conferencing system according to an embodiment of the present invention. Figure 3 The embodiment shown is Figure 1 An application example of an embodiment of the present invention. Figure 3 In the embodiment, the video conference system 100 further includes a display device 150. The display device 150 is used to present the video conference screen. For example, the display device 150 may be implemented by a liquid crystal display (LCD), a plasma display, a projection system, etc.
[0127] The storage 120 also includes an audio source identification module 210, a speech behavior recognition module 220, and a sound reception switching module 230. The audio source identification module 210 is used to match the portrait of the participant with the corresponding voice of the participant. The speech behavior recognition module 220 is used to detect whether the participant is in a non-speaking behavior. In one embodiment, the audio source identification module 210 may further include a voiceprint recognition module 211 and a portrait recognition module 213. The speech behavior recognition module 220 includes an audio recognition module 221 and an image recognition module 223.
[0128] The voiceprint recognition module 211 is used to receive audio signals and separate all human voices (one or more human voices) from the audio signals. The portrait recognition module 213 is used to receive video signals and identify each participant included in the video signal, and obtain the relative position of each participant in the conference space. The audio recognition module 221 is used to receive audio signals and identify the sound intensity of each human voice. The image recognition module 223 is used to receive video signals and identify the behavioral events of each participant in the video signal. For example, the face of the portrait corresponding to each participant is detected to be covered, the gestures corresponding to each portrait are detected to determine whether the back of the palm appears, and the distance changes between the portraits in multiple image frames are detected.
[0129] Figure 4This is a flow chart of a method for switching audio recording of a video conference according to an embodiment. Figure 3 and Figure 4 In step S401, the processor 110 uses the image capturing device 130 to capture images at intervals of a sampling time to obtain a video signal. In step S405, the processor 110 uses the sound receiving device 140 to receive sound at intervals of a sampling time to obtain an audio signal. The time for the image capturing device 130 to capture images and the time for the sound receiving device 140 to receive sound are set to be the same, so that a video signal and an audio signal corresponding in time are obtained at intervals of a sampling time.
[0130] After obtaining the video signal and the audio signal, the processor 110 drives the audio source identification module 210 to match each voice in the audio signal with the corresponding portrait of a participant in the video signal. Specifically, in step S403, the processor 110 identifies each participant included in the video signal through the portrait recognition module 213, and obtains the portrait corresponding to each participant, and obtains the relative position of each participant in the conference space. For example, the portrait recognition module 213 can identify each image frame of the video conference and find objects such as furniture or furnishings in the conference space that do not move, thereby determining the relative position of each participant in the conference space. Alternatively, the portrait recognition module 213 can also determine the relative position between movable objects, thereby determining the relative position of each participant in the conference space.
[0131] In step S407, the voices belonging to different people are separated from the audio signal through the voiceprint recognition module 211. In addition, the voiceprint recognition module 211 will further execute the sound source localization algorithm to determine the source position of each voice in the conference space. Afterwards, in step S409, the processor 110 matches each voice with one of the corresponding participants based on each relative position obtained by the portrait recognition module 213 and each source position obtained by the voiceprint recognition module 211 through the audio source recognition module 210. Here, since not every participant will speak during this recording time, the number of voices in the audio signal may be less than the number of participants in the video signal. Accordingly, the processor 110 matches the separated voice with the image of the participant.
[0132] After matching each human voice with the corresponding participant, the processor 110 drives the speech behavior recognition module 220 to determine whether the participant is in a non-speaking behavior based on the video signal and the audio signal. Specifically, in step S411, the behavior event of each participant is identified through the image recognition module 223. After that, it is further determined whether the behavior event conforms to the preset event pre-defined in the storage 120. The preset event may include at least one of the mouth covering event, the palm back facing forward event, and the approaching person event.
[0133] Figure 5A to Figure 5E is a schematic diagram of multiple preset events according to an embodiment of the present invention. Figure 5A The preset events shown include a mouth covering event and a palm back facing forward event. Figure 5B The preset events shown only include the mouth covering event. Figure 5C The preset events shown include a mouth covering event and a palm back facing forward event. Figure 5D The preset events shown include a mouth covering event, a palm back facing forward event, and an approaching person event. Figure 5E The preset events shown include a bystander event. Figure 5A to Figure 5E This is for illustrative purposes only and is not intended to be limiting.
[0134] In step S413 , the sound intensity of each human voice is identified through the audio recognition module 221 . For example, the audio recognition module 221 determines the sound intensity based on the signal corresponding to each human voice separated by the voiceprint recognition module 211 .
[0135] After obtaining the behavior events of each participant and the sound intensity corresponding to each human voice, in step S415, the speaking behavior recognition module 220 determines whether the participant is in a non-speaking behavior. When determining that the participant is in a speaking behavior, the participant is regarded as a speaker. In step S417, the sound reception switching module 230 switches the state corresponding to the speaker to the sound reception state to receive the sound of the speaker. When determining that the participant is in a non-speaking behavior, the participant is regarded as a non-speaker. In step S419, the sound reception switching module 230 switches the state corresponding to the non-speaker to the mute state to filter the sound of the non-speaker.
[0136] In this embodiment, the speaking behavior recognition module 220 may be configured to determine whether each participant is in a non-speaking behavior based on the behavior event of each participant and the sound intensity of the corresponding human voice.
[0137] In one embodiment, the speaking behavior recognition module 220 can be set to determine whether the participant is in a non-speaking behavior based on the matched portraits of each participant and their corresponding voices. For example, after the audio recognition module 221 recognizes the sound intensity of the voice, it further determines whether the sound intensity is less than a default value. And, after the image recognition module 223 obtains the behavioral event of the participant, it determines whether the behavioral event of the participant matched with the voice meets the preset event. When the behavioral event meets the default event and the sound intensity is less than the default value, the participant is determined to be a non-speaker in a non-speaking behavior. Then, in step S419, the non-speaker is filtered.
[0138] On the other hand, if the behavior event matches the preset event and the sound intensity is not less than the default value, the participant is determined to be a speaker in the speaking behavior. Then, in step S417, the speaker's sound is collected.
[0139] In another embodiment, the speaking behavior recognition module 220 may be further configured to determine that the participant is in speaking behavior as long as the detected behavior event does not match the preset event regardless of the sound intensity.
[0140] Figure 6 FIG. 1 is a schematic diagram of an application scenario of a video conference according to an embodiment of the present invention. Figure 6 As shown, multiple participants in a conference room at the local end conduct a video conference with remote participants through the video conference system 100. The video conference system 100 disclosed herein can switch audio for speakers and non-speakers in the same conference space through the above-mentioned embodiments. For example, when one of the participants A in the same conference space temporarily needs to discuss with another participant B, the video conference system 100 can filter the audio of participants A and B through "non-speaking behavior" so that other remote participants will not be affected by the audio of participants A and B at the local end.
[0141] In addition, the processor 110 may further provide a user interface in the display device 150 to display the screen of each participant of the video conference. In one embodiment, the processor 110 identifies each participant included in the video signal by identifying the video signal, obtains the video segment corresponding to each participant from the video signal, and displays the user interface in the display device 150, and displays the video segment corresponding to each participant in the user interface.
[0142] Figure 7 is a schematic diagram of a user interface of an embodiment. Figure 7The user interface 700 includes an area V1 for presenting the local video signal and a plurality of display areas 710-740. The number of display areas 710-740 is based on the number of participants identified from the video signal. In this embodiment, four participants U1-U4 are used as an example for explanation, but the invention is not limited thereto. The display areas 710-740 are each used to display the video segments C1-C4 corresponding to the participants U1-U4. In other embodiments, the representative portraits (static images) corresponding to the participants U1-U4 may also be displayed in the display areas 710-740.
[0143] exist Figure 7 In the illustrated embodiment, it is assumed that participant U4 is determined to be a non-speaker in a non-speaking behavior, and participants U1-U3 are determined to be speakers in a speaking behavior. In the display area 740 that displays the video segment C4 corresponding to participant U4 (non-speaker), a mute mark M4 is displayed. In the display areas 710-730 that display the video segments C1-C3 corresponding to participants C1-C3 (speakers), the sound receiving marks M1-M3 are displayed respectively.
[0144] In addition, a corresponding switching button may be further provided in each of the display areas 710 to 740 for manual switching between receiving and filtering. For example, the receiving marks M1 to M3 and the mute mark M4 have a switching function, and each participant U1 to U4 may be manually controlled to be in a receiving state or a mute state. For example, when the mute mark M4 is enabled, the mute mark M4 is switched to a receiving mark, and the state is switched to receiving the participant U4. For example, when the receiving mark M1 is enabled, the receiving mark M4 is switched to a mute mark, and the state is switched to receiving the participant U1.
[0145] In addition, the speaker who is currently speaking and other speakers who are not currently speaking can be further distinguished through visual display. Figure 7 In the embodiment shown, assuming that the speaker who is speaking is participant U1, the display area 710 is displayed in a bold frame, and the display areas 720-740 are displayed in dotted frames. However, this is only an example and is not limited thereto.
[0146] In one embodiment, it can be further configured that the display device 150 displays the video screen of other remote participants while displaying the user interface 700. Alternatively, the display device 150 displays the video screen of other remote participants while only displaying the display areas 720 to 750. Alternatively, another display device (such as a projector) is configured to display the video screen of other remote participants, and the user interface 700 is displayed on the display device 150.
[0147] In summary, the method can identify the relative position of each participant in the conference space in the video signal, and automatically filter / restore the sound for each participant based on whether each participant is in a non-speaking behavior. In addition, after starting the video conference, the disclosure can further automatically match the participant's portrait with the participant's corresponding voice, and use behavioral events and sound intensity to determine whether the participant is in a non-speaking behavior, which can make the recognition result more accurate. In addition, the disclosure only needs to define the type of preset event (non-speaking behavior) without pre-establishing an image database of portrait behavior, and can realize behavior recognition in real time.
[0148] In the description of this specification, the description with reference to the terms "some embodiments", "other embodiments", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic description of the above terms does not necessarily refer to the same embodiment or example.
[0149] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0150] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for switching audio reception in a video conference, using a processor to execute the following steps when starting a video conference, characterized in that: include: Obtaining a relative position of each of a plurality of participants in a conference space and a behavior event of each of the participants by identifying a video signal; Based on the behavior event of each participant, determining whether each participant is in a non-speaking behavior; In response to one of the participants being determined to be a non-speaker in the non-speaking behavior, adjusting a sound pickup range of a sound pickup device based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker; as well as In response to one of the participants being determined to be a talker who is not in the non-speaking behavior, the sound pickup range of the sound pickup device is adjusted based on the relative position of the talker in the conference space to pick up the sound of the talker.
2. The method for switching video and audio according to claim 1, characterized in that: Wherein, by identifying a video signal, obtaining a relative position of each of the plurality of participants in a conference space and a behavior event of each of the participants includes: Identify each of the participants included in the video signal through a portrait recognition module, and obtain the relative position of each of the participants in the conference space; and The behavior event of each participant is identified through an image recognition module.
3. The method for switching audio reception of a video conference according to claim 1, characterized in that: Also includes: After receiving an audio signal, a human voice is separated from the audio signal through a voiceprint recognition module; matching the human voice to one of the corresponding participants; and Based on the human voice and the matching behavior event of one of the participants, it is determined whether one of the participants is in the non-speaking behavior.
4. The method for switching audio reception of a video conference according to claim 3, characterized in that: The step of matching the human voice with one of the corresponding participants includes: Executing a sound source localization algorithm to determine a source position of the human voice in the conference space; and Based on the relative position and the source position, the human voice is matched with one of the participants corresponding thereto.
5. The method for switching audio reception of a video conference according to claim 3, characterized in that: The step of judging whether one of the participants is in the non-speaking behavior based on the human voice and the matching behavior event of one of the participants includes: Identifying a sound intensity of the human voice through an audio recognition module; Determining whether the sound intensity is less than a default value; Determining whether the behavior event of one of the participants matching the human voice conforms to a preset event; and When the behavior event conforms to the preset event and the sound intensity is less than the default value, it is determined that one of the participants is in the non-speaking behavior.
6. The method for switching audio reception of a video conference according to claim 3, characterized in that: After the video conference is started by the processor, the method further includes: Capturing images at intervals of a sampling time by an imaging device to obtain the video signal; and A sound receiving device is used to receive sound at intervals of the sampling time to obtain the audio signal.
7. The method for switching audio reception of a video conference according to claim 1, characterized in that: Also includes: A user interface is displayed on a display device through the processor, and the video segment corresponding to each participant is displayed in the user interface.
8. The method for switching audio reception of a video conference according to claim 7, characterized in that: The user interface includes a plurality of display areas, each of which corresponds to each of the participants, and each of which is used to display the video clip corresponding to each of the participants; After the user interface is displayed on the display device through the processor, the method further includes: Displaying a mute mark in one of the display areas displaying the video segment corresponding to the non-speaker; as well as A sound receiving mark is displayed in another one of the display areas that displays the video segment corresponding to the speaker.
9. The method for switching audio reception of a video conference according to claim 1, characterized in that: The step of filtering the non-speaker includes: The processor filters out the human voice corresponding to the relative position of the non-speaker in the conference space from the audio signal.
10. A video conferencing system, characterized in that: include: a storage device storing an application program; An imaging device for acquiring a video signal; a radio receiving device; as well as A processor is coupled to the storage, the imaging device and the sound receiving device, and is configured to: Executing the application to start a video conference, and initiating the video conference, comprising: Obtaining a relative position of each of a plurality of participants in a conference space and a behavior event of each of the participants by identifying a video signal; Based on the behavior event of each participant, determining whether each participant is in a non-speaking behavior; In response to one of the plurality of participants being determined to be a non-speaker in the non-speaking behavior, adjusting a sound pickup range of a sound pickup device based on the relative position of the non-speaker in the conference space to filter the sound of the non-speaker; and In response to one of the multiple participants being determined to be a talker who is not in the non-speaking behavior, the sound collection range of the sound collection device is adjusted based on the relative position of the talker in the conference space to collect the sound of the talker.
11. The video conferencing system according to claim 10, characterized in that: The storage device includes a portrait recognition module and an image recognition module, and the processor is configured to: Identify each participant included in the video signal through the portrait recognition module, and obtain the relative position of each participant in the conference space; The behavior event of each participant is identified through the image recognition module.
12. The video conferencing system according to claim 10, characterized in that: The sound receiving device is used to obtain an audio signal; The storage device includes a voiceprint recognition module, and the processor is configured to: Separating a human voice from the audio signal through the voiceprint recognition module; matching the human voice to one of the corresponding participants; and Based on the human voice and the matching behavior event of one of the participants, it is determined whether one of the participants is in the non-speaking behavior.
13. The video conferencing system according to claim 12, characterized in that: wherein the processor is configured to: Executing a sound source localization algorithm to determine a source position of the human voice in the conference space; and Based on the relative position and the source position, the human voice is matched with one of the participants corresponding thereto.
14. The video conferencing system according to claim 12, characterized in that: The storage includes an audio recognition module, and the processor is configured to: Recognize a sound intensity of the human voice through the audio recognition module; Determining whether the sound intensity is less than a default value; Determining whether the behavior event of one of the participants matching the human voice conforms to a preset event; When the behavior event conforms to the preset event and the sound intensity is less than the default value, it is determined that one of the participants is in the non-speaking behavior.
15. The video conferencing system according to claim 12, characterized in that: wherein the processor is configured to: The image capturing device captures images at intervals of a sampling time to obtain the video signal; and The sound receiving device receives sound at intervals of the sampling time to obtain the audio signal.
16. The video conferencing system according to claim 10, characterized in that: wherein the processor is configured to: By identifying the video signal, identifying each of the participants included in the video signal, and obtaining a video segment corresponding to each of the participants from the video signal; and A user interface is displayed on a display device, and the video segment corresponding to each participant is displayed in the user interface.
17. The video conferencing system according to claim 16, characterized in that: The user interface includes a plurality of display areas, each of which corresponds to each of the participants, and each of which is used to display the video clip corresponding to each of the participants; wherein the processor is configured to: Displaying a mute mark in one of the display areas displaying the video segment corresponding to the non-speaker; and A sound receiving mark is displayed in another one of the display areas that displays the video segment corresponding to the speaker.
18. The video conferencing system according to claim 10, characterized in that: wherein the processor is configured to: Based on the relative position of the non-speaker in the conference space, the human voice corresponding to the relative position in the audio signal is filtered out.