Mute control method and device for video conference and computer readable storage medium

By using face detection and lip status analysis in video conferencing, the automatic mute function solves the noise interference problem caused by forgetting to turn off the microphone, ensuring clear transmission of speech content and a good user experience.

CN121750813APending Publication Date: 2026-03-27NANNING FUGUI PRECISION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In video conferences, participants forgetting to turn off their microphones or mute them can cause ambient noise interference, affecting the listening experience of other participants.

Method used

By collecting video data from the meeting space, face detection and lip positioning are performed. The speaking status is determined based on the opening and closing of the lips, and the mute function is automatically turned on or off.

Benefits of technology

It enables intelligent control of the mute function based on the speaking status during video conferences, reducing environmental noise interference, ensuring clear transmission of speaking content, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750813A_ABST
    Figure CN121750813A_ABST
Patent Text Reader

Abstract

A mute control method for a video conference comprises the steps of collecting video data of a conference space and executing face detection, and when a face is detected, positioning the lip of the face. According to the continuous opening and closing state of the lips, the speaking state of a person corresponding to the human face is judged, and meanwhile, the opening or state of a mute function is controlled according to the speaking state of the person. The invention further provides a device for implementing the method and a computer readable storage medium. According to the invention, the speaking state of the conference participant can be judged through the video, and the mute function is automatically controlled to be opened or closed according to the speaking state.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video conference technology, and in particular to a video conference mute control method, device and computer readable storage medium. BACKGROUND

[0002] Currently, users generally use personal computers, notebook computers, tablets, mobile phones and conference room video conference systems to conduct video conferences, and achieve the purpose of communication through video interaction and voice dialogue in the conference process.

[0003] When a conference participant wants to mute, he basically manually closes the voice transmission, such as closing the microphone function or the mute function key in the computer system. However, if the conference participant forgets to close the microphone or press the mute function key, the voice system of the video conference will receive external noise interference, resulting in poor listening effect of other participants in the conference. SUMMARY

[0004] Therefore, the present application aims to provide a video conference mute control method, device and computer readable storage medium, which can ensure that the speech content can be clearly transmitted in the video conference process and reduce environmental background noise.

[0005] An embodiment of the present application provides a video conference mute control method, which comprises: collecting video data of a conference space; performing face detection on an image frame sequence in the video data; judging whether a face is detected according to a result of the face detection; performing lip positioning of the face when it is judged that a face is detected; judging a speech state of a person corresponding to the face according to a continuous opening and closing state of the lips, wherein the speech state comprises a speaking state and a non-speaking state; and controlling opening or closing of a mute function according to the speech state of the person corresponding to the face.

[0006] An embodiment of the present application also provides a video conference mute control device, which is characterized by comprising a processor; and a memory for storing a computer program, which makes the processor realize the video conference mute control method when the computer program is executed by the processor.

[0007] An embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the video conference mute control method.

[0008] Compared with the prior art, the video conference mute control method, device and computer readable storage medium provided by the application can further locate the lips through face detection of the video, judge the speaking state of the person according to the continuous opening and closing state of the lips, and then control the opening or closing of the mute function. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 A block diagram of a video conference mute control device according to an embodiment of the application.

[0010] Figure 2 A flowchart of a video conference mute control method according to an embodiment of the application.

[0011] Figure 3 A schematic diagram of a plurality of feature parameters of lips according to an embodiment of the application.

[0012] Figure 4 A flowchart of a video conference mute control method according to another embodiment of the application.

[0013] Explanation of main element symbols

[0014]

[0015] The following detailed description will further illustrate the application in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0016] In order to facilitate those skilled in the art of the application to understand and implement the application, the following further describes the application in detail in conjunction with the drawings and embodiments. It should be understood that the application provides many applicable inventive concepts, which can be implemented in various specific forms. Those skilled in the art can utilize the details described in these embodiments or other embodiments and other applicable structural, logical and electrical changes to implement the application without departing from the spirit and scope of the application.

[0017] The present application specification provides different embodiments to illustrate the technical features of different embodiments of the application. Among them, the configuration of each component in the embodiment is for illustration, not to limit the application. And part of the repeated figure numbers in the embodiment is to simplify the description, not to mean the relevance between different embodiments. Among them, the same component numbers used in the icons and the specification represent the same or similar components. The drawings of the present application are simplified and not drawn in accurate proportion.

[0018] Furthermore, in describing some embodiments of the present invention, the specification describes the method and / or procedure of the present invention in a specific order of steps. However, since the method and procedure are not necessarily performed according to the specific order of steps described, they are not limited to the specific order of steps. Those skilled in the art will understand that other orders are also possible implementations. Therefore, the specific order of steps described in the specification is not intended to limit the scope of the patent application. Moreover, the scope of the present invention for the method and / or procedure is not limited to the order of execution steps written therein, and those skilled in the art will understand that adjusting the order of execution steps does not depart from the spirit and scope of the present invention.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Some embodiments of the invention are described in detail below with reference to the accompanying drawings.

[0020] Please see Figure 1 The diagram shown is a hardware block diagram of a video conferencing mute control device 100 according to an embodiment of the present invention. The device 100 includes a processor 102, a memory 104, a communication interface 106, an audio / video input module 108, and an audio / video output module 110. Those skilled in the art should understand that... Figure 1 The composition of the device 100 shown does not constitute a limitation of the embodiments of the present invention. Figure 1 The device 100 shown is simplified for ease of description, and in different embodiments it may include fewer or more components than shown.

[0021] In one embodiment, the processor 102 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 102 is the control core of the device 100, connecting various components of the device 100 via various interfaces and lines. It executes computer programs or modules stored in the memory 104 and calls data stored in the memory 104 to perform various functions of the device 100 and process data, such as a mute control method for video conferencing.

[0022] In an embodiment, the memory 104 is configured to store codes of computer programs and various data, such as the mute control method of video conference, and to realize high-speed and automatic access of programs or data during the operation of the apparatus 100. The memory 104 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disc memory, a magnetic tape memory, or any other computer-readable storage medium capable of carrying or storing data.

[0023] In an embodiment, the communication interface 106 is configured by a communication circuit to communicate data or information with other video conference apparatuses, including transmitting and receiving audio and video communications of the apparatus 100.

[0024] In an embodiment, the audio and video input module 108 is, for example, a microphone and a camera, and the audio and video input module 108 is configured to receive sounds or images of a user and output audio signals or video signals accordingly.

[0025] In an embodiment, the audio and video output module 110 is, for example, a headset and a screen, and the audio and video output module 110 is configured to present sounds and images to the user.

[0026] Referring to FIG. 1, a schematic diagram of a video conference apparatus according to an embodiment of the present application is shown. As shown in FIG. 1, the apparatus 100 includes a processor 102, a memory 104, a communication interface 106, an audio and video input module 108, and an audio and video output module 110. Figure 2 In an embodiment, the apparatus 100 is configured to implement the mute control method of video conference. Figure 2 In an embodiment, the apparatus 100 is configured to implement the mute control method of video conference.

[0027] In an embodiment, the apparatus 100 is configured to implement the mute control method of video conference.

[0028] In an embodiment, conference participants in different geographical locations can participate in the same video conference through a multiple control unit (MCU) by using video conference apparatuses. The video conference apparatus, for example, the apparatus 100, is configured to implement the mute control method of video conference. Figure 1The device 100 can be used to collect video data of a conference space.

[0029] In step S202, face detection is performed on the sequence of image frames in the video data.

[0030] Specifically, any face detection algorithm can be performed to detect faces in the sequence of image frames in the video data, thereby obtaining face detection results.

[0031] In step S203, it is determined whether a face is detected.

[0032] According to the face detection results obtained in step S202, it is determined whether a face is detected. When it is determined that no face is detected, step S204 is performed; when it is determined that a face is detected, step S205 is performed.

[0033] In step S204, the mute function is turned on.

[0034] When it is determined in step S203 that no face is detected, it can be inferred that there is no person in the conference space or that all the persons in the conference space are in a non-speaking state, and the mute function is automatically turned on.

[0035] In one example, when the mute function of the video conference device is turned on, a visual prompt such as "mute state" is provided on the display device connected to the video conference device.

[0036] In another example, when the mute function of the video conference device is turned on, the prompt light of the sound pickup device (e.g., microphone) connected to the video conference device is turned on, so that the user can know that the mute function of the local device has been turned on through the prompt light of the sound pickup device.

[0037] In step S205, lip positioning is performed according to the face detection results.

[0038] Specifically, the face detection results in step S202 can obtain multiple face images, and further, the face position and face feature information of each face image can be extracted, and the lip can be positioned through the face feature information.

[0039] In one implementation, the face position and face feature information extracted can also be used to determine the face state. The face state refers to the state of the face of the person relative to the video conference device, which can be one of a front face, a side face, a lowered head, a raised head, or an occluded face. Step S205 is only for positioning the lips of the front face.

[0040] In step S206, the speaking state of the person is determined according to the continuous lip opening and closing state.

[0041] Specifically, a plurality of feature parameters of the lip in each image are extracted, and the opening and closing state of the lip is determined according to the plurality of feature parameters of the lip. The opening and closing state of the lip includes an open state and a closed state, and the speaking state of the person includes a speaking state and a non-speaking state.

[0042] In an embodiment, the plurality of feature parameters of the lip include: an inner lip width ω1, an outer lip width ω0, an upper outer lip height h1, an upper inner lip height h2, a lower inner lip height h3, a lower outer lip height h4, a lip deflection angle θ, a mouth center point coordinate (X c , Y c ), and an offset a off of a quartic curve of the upper outer lip from the coordinate origin, as shown in FIG. 3. Figure 3

[0043] In an embodiment, the inner lip area, the outer lip area, and the mouth shape can be calculated according to the plurality of feature parameters of the lip, and the opening and closing state of the lip can be determined according to the mouth shape.

[0044] In different embodiments, the opening and closing state of the lip can also be determined according to the sum of the upper inner lip height and the lower inner lip height. When the sum of the upper inner lip height and the lower inner lip height is equal to 0, it is the closed state of the lip.

[0045] When the continuous opening and closing state of the lip is the closed state and the closed state lasts for a preset time length, it is determined that the corresponding person is in the non-speaking state; otherwise, it is determined that the corresponding person is in the speaking state.

[0046] In an embodiment, if the current mute function is in the closed state and the face of the previous frame is in the front face state, regardless of the face of the subsequent image frame being in the side face, low head, high head, or occlusion state, it is determined that the speaking state of the corresponding person is the speaking state. However, if the current mute function is in the open state and the face of the previous frame is in the front face state, regardless of the face of the subsequent image frame being in the side face, low head, high head, or occlusion state, it is determined that the speaking state of the corresponding person is the non-speaking state.

[0047] In step S207, the opening or closing of the mute function is controlled according to the speaking state of the person.

[0048] Specifically, when the speaking state of the person is determined to be the non-speaking state, the mute function is controlled to be opened; when the speaking state of the person is determined to be the speaking state, the mute function is controlled to be closed.

[0049] ​In an embodiment, if the speech state of the person is determined to be the non-speaking state and the mute function is currently turned off, a countdown timer for turning on the mute function can be prompted to the user before the mute function is controlled to be turned on. If no instruction for canceling the turning on of the mute function is received during the countdown timer, the mute function is maintained to be turned on. If an instruction for canceling the turning on of the mute function is received during the countdown timer, the mute function is maintained to be turned off.

[0050] In an embodiment, the video conference device includes a plurality of sound pickup devices, for example, a plurality of microphones. According to the number and placement of the plurality of sound pickup devices, the conference space can be divided into a plurality of regions, and the mute control method of the video conference as shown in Figure 2 is performed on each region. Specifically, in the video data collected in step S201, the sound pickup devices can be identified, and the conference space can be divided into a plurality of regions according to the number and positions of the identified sound pickup devices. That is, each region corresponds to a sound pickup device, and the mute control method of the video conference as shown in Figure 2 is performed on the sound pickup device corresponding to the region.

[0051] Please refer to Figure 4 , which is a flowchart of the mute control method of the video conference according to another embodiment of the present application. After the mute function is turned on, a method for identifying the speaking intention of the user and controlling the mute function to be turned off is needed to avoid disturbing the conference participants. As shown in Figure 4 , the method includes:

[0052] Step S401, collecting video data of the conference space.

[0053] Step S402, performing human body detection on the image frame sequence in the video data to identify the conference participants.

[0054] The human body detection on the image frame sequence in the video data can detect one or more human bodies in the conference space, wherein each detected human body corresponds to one conference participant participating in the video conference.

[0055] In an embodiment, a human body detection model can be pre-trained, the input of the human body detection model is each image frame, and the output is the bounding box of the human body contained in the image frame. The human body detection model can be trained by using pre-input training data through deep learning or machine learning method. In the execution of the method flow, the feedback information of the user can be continuously collected for self-learning and optimization to improve the accuracy of the human body detection model in identifying the conference participants.

[0056] Step S403, identifying the limb movement and gaze direction of the conference participant.

[0057] After the human body detection identifies the participant in step S402, the relative position of the hand and the body center of gravity and the relative position of the hand and the head of the participant can be further extracted as characteristic parameters, which are input into the limb action detection model to obtain the recognition result of the limb action.

[0058] According to the human body detection result, it can be further determined whether the human face is included. After it is determined that the human face is included, the facial feature points and the eye feature points can be extracted to calculate the orientation of the face and the line of sight.

[0059] In step S404, the speaking intention of the participant is analyzed according to the limb action and the line of sight of the participant.

[0060] The limb action and the line of sight can be used to convey the information of the psychological activity of the human body, as a clue to analyze the intention. For example, when a person speaks, there can be limb actions such as raising hands and waving hands, or subtle eye movements such as staring at the microphone or the display screen. Collecting data such as the limb action and the line of sight of the participant can be used to analyze the speaking habit of the participant, and the collected data can be further trained and optimized to model, so as to analyze the speaking intention of the participant according to the limb action and the line of sight of the participant.

[0061] In step S405, the opening or closing of the mute function is controlled according to whether the participant has the speaking intention.

[0062] When it is determined that the participant has the speaking intention, the mute function is controlled to be closed; when it is determined that the participant does not have the speaking intention, the mute function is controlled to be opened.

[0063] In an embodiment, when it is determined that the participant has the speaking intention and the mute function is controlled to be closed, if no voice data is collected for a period of time, it is determined that the participant does not have the speaking intention, and the mute function is controlled to be opened.

[0064] In the process of executing the method of Figure 4 , the feedback of the user can be received at any time, and the related models used are updated and corrected to improve the accuracy of the related recognition.

[0065] In an embodiment, the video conference device includes a plurality of sound pickup devices, for example, a plurality of microphones. According to the number and placement position of the plurality of sound pickup devices, the conference space can be divided into a plurality of regions, and the mute control method of the video conference as shown in Figure 4 is performed on each region. Specifically, in the video data collected in step S401, the identification of the sound pickup device can be performed, and the conference space is divided into a plurality of regions according to the number and position of the identified sound pickup devices. That is, each region corresponds to a sound pickup device, and the mute control method of the video conference as shown in Figure 4The mute control method of the video conference shown can control the mute function of the sound pickup device corresponding to the area.

[0066] In summary, the mute control method of the video conference, the device and the computer readable storage medium can intelligently control the opening or closing of the mute function according to the opening and closing state of the lips, the body movement and the line of sight of the video conference participants, so that the participants have a better use experience.

[0067] It is worth noting that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for controlling mute in a video conference, characterized in that, The method includes: Collect video data from the meeting space; Perform face detection on the image frame sequence in the video data; Based on the face detection results, determine whether a face has been detected; When a face is detected, the lips of the face are located. Based on the continuous opening and closing of the lips, the speech state of the person corresponding to the face is determined, wherein the speech state includes a speaking state and a non-speaking state; and The mute function is controlled to be turned on or off based on the speaking state of the person corresponding to the face.

2. The video conferencing mute control method as described in claim 1, characterized in that, The method further includes: When it is determined that no face is detected, the mute function is activated.

3. The video conferencing mute control method as described in claim 1, characterized in that, The method further includes: when the mute function is activated, illuminating the indicator light on the microphone.

4. The video conferencing mute control method as described in claim 1, characterized in that, The opening and closing states of the lips include an open state and a closed state.

5. The video conferencing mute control method as described in claim 4, characterized in that, The step of determining the speaking state of the person corresponding to the face based on the continuous opening and closing state of the lips further includes: Extract multiple feature parameters of the lips, including inner lip width, outer lip width, upper outer lip height, upper inner lip height, lower inner lip height, lower outer lip height, lip deflection angle, lip center point coordinates, and the offset of the upper outer lip's quadratic curve from the origin. Calculate the area of ​​the inner lip and the area of ​​the outer lip based on the aforementioned multiple feature parameters; The mouth shape is determined based on the area of ​​the inner lip and the area of ​​the outer lip; and The current opening and closing state of the lips is determined based on the mouth shape.

6. The video conferencing mute control method as described in claim 1, characterized in that, The method further includes: When the lips are in a closed state for a preset duration during continuous opening and closing, the speech state of the person corresponding to the face is determined to be a non-speaking state. as well as When it is determined that the person corresponding to the face is not speaking, the mute function is activated.

7. The video conferencing mute control method as described in claim 6, characterized in that, The method further includes: If it is determined that the person is not speaking and the mute function is currently off, then before controlling the mute function to turn on, the user is prompted with a countdown timer for turning on the mute function; If no cancellation command is received from the user during the countdown period, the mute function remains enabled. as well as When a cancellation command is received from the user during the countdown period, the mute function is deactivated.

8. The video conferencing mute control method as described in claim 1, characterized in that, The method further includes: Human detection is performed on the image frame sequence in the video data to identify attendees; Identify the body movements and gaze direction of the attendees; Analyze whether the participants intend to speak based on their body language and eye contact. as well as The mute function is turned on or off depending on whether the attendees intend to speak.

9. A mute control device for video conferencing, characterized in that, include Processor; and A memory for storing a computer program that, when executed by the processor, causes the processor to implement the mute control method for video conferencing as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the mute control method for video conferencing as described in any one of claims 1 to 8.