Video conference control method, device and equipment and computer readable storage medium
By identifying the current speaker and generating a speech interruption strategy based on speaking information and facial information, the problem of meeting chaos caused by multiple people speaking at the same time is solved, and meeting efficiency is improved.
Patent Information
- Application Number
- CN202310431579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-04-20
AI Technical Summary
Existing video conferencing systems are prone to causing chaos and affecting meeting efficiency when multiple people are participating, as multiple people speak at the same time.
By acquiring information about meeting participants and the current meeting content, the current speaker is identified, and the speaking status is judged based on speaking information and facial information. Speaking interruption strategies are generated to optimize the speaking order and reduce confusion.
This improved the efficiency of the meeting, reduced the possibility of multiple people speaking at the same time and disrupting the meeting, and ensured that the meeting proceeded normally.
Smart Images

Figure CN116436715B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of conferencing systems, and in particular to a video conferencing control method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] With the continuous development of internet technology, people have increasingly higher requirements for the timeliness and urgency of problem-solving, and video conferencing has been widely used in daily life.
[0003] Current video conferencing only offers two modes: speaking with an open microphone and stopping speaking with an open microphone. When there are many participants, it is easy for multiple people to speak simultaneously, affecting the quality of the meeting. Furthermore, when other interfering factors are present, the meeting will descend into chaos, hindering its normal progress and thus impacting its efficiency. Summary of the Invention
[0004] To improve the efficiency of meetings, this application provides a video conferencing control method, apparatus, device, and computer-readable storage medium.
[0005] Firstly, this application provides a video conferencing control method, which adopts the following technical solution:
[0006] A video conferencing control method, comprising:
[0007] Obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content;
[0008] In response to a speech interruption request, obtain the current speaker's speech information and facial information;
[0009] The current speaker's speaking status is determined based on the spoken information and the facial information;
[0010] Based on the speaking status, determine whether the current speaker can be interrupted immediately;
[0011] If the current speaker can be interrupted immediately, the information of the person who requested the interruption request is obtained;
[0012] A speech interruption notification is generated based on the requester information, and the speech interruption notification is sent to the meeting interface of the current speaker.
[0013] If the current speaker cannot be interrupted immediately, the interruption information of the person who requested the interruption request will be collected.
[0014] A speaking strategy is generated based on the interruption information and the speaking status.
[0015] By adopting the above technical solution, the current speaker is determined based on the current meeting content. After the current speaker is determined, the speaker's speaking information and facial information are collected. When a speaking interruption request is received, the current speaker's speaking status is determined based on the speaking information and facial information. If the speaking status allows interruption, the current speaker's speech is interrupted and a prompt is given. If not, a speaking strategy is generated to allow the requester to speak at an appropriate time. This reduces the possibility of multiple people speaking at the same time and disrupting the normal progress of the meeting, thereby improving the efficiency of the meeting.
[0016] Optionally, determining the current speaker based on the personnel information and the current meeting content includes:
[0017] Obtain the first voice information of the meeting participants and the second voice information of the sender who sent the current meeting content;
[0018] Find the sound information in the first sound information that is the same as the second sound information;
[0019] The current speaker is determined based on the voice information and the personnel information.
[0020] Optionally, determining the current speaker's speaking status based on the speaking information and the facial information includes:
[0021] The speaking speed and volume of the current speaker are determined based on the speaking information.
[0022] The urgency level of the current speaker is determined based on the speaking speed, the speaking volume, and preset level rules.
[0023] The current speaker's emotional state is determined based on the facial information and preset emotion rules;
[0024] The current speaker's speaking status is determined based on the urgency level of the speech and the emotional state.
[0025] Optionally, generating a speech interruption notification based on the requester information includes:
[0026] Obtain the speech content of the current speaker, analyze the speech content, and generate a preset interruption reason;
[0027] Obtain the request time and number of times the speech interruption request was made;
[0028] Calculate the request frequency based on the request time and the number of requests;
[0029] Obtain preset request level rules, and determine the request level based on the preset level rules and the request frequency;
[0030] A speech interruption notification is generated based on the interruption reason and the request level.
[0031] Optionally, the interruption information includes the interruption content; the generation of a speaking strategy based on the interruption information and the speaking status includes:
[0032] Calculate the correlation between the interrupted speech and the current meeting content;
[0033] Determine whether the correlation is not less than a preset correlation threshold;
[0034] If the correlation is not less than a preset correlation threshold, the interrupted speech content will be played directly when the current speaker pauses.
[0035] If the relevance is less than a preset relevance threshold, a speaking prompt is generated based on the interrupted speech content, and the speaking prompt is sent to the current speaker's meeting interface.
[0036] Optionally, generating a speaking prompt based on the interrupted speech content includes:
[0037] Keyword extraction is performed on the speech content, and a speech summary is generated based on the extracted keywords;
[0038] Obtain the name of the requester, and generate a message prompt based on the message summary and the name.
[0039] Optional, also includes:
[0040] Obtain the content and duration of all statements made by the participants in the meeting;
[0041] A meeting report is generated based on the content of the statements made at the meeting and the time when those statements were made.
[0042] Secondly, this application provides a video conferencing control device, which adopts the following technical solution:
[0043] A video conferencing control device, comprising:
[0044] The current speaker determination module is used to obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content;
[0045] The speech information acquisition module is used to acquire the speech information and facial information of the current speaker in response to a speech interruption request;
[0046] The speaking status confirmation module is used to determine the speaking status of the current speaker based on the speaking information and the facial information;
[0047] The speech interruption judgment module is used to determine whether the current speaker can be interrupted immediately based on the speech status.
[0048] The interruption request acquisition module is used to acquire the requester information of the speech interruption request;
[0049] An interruption notification generation module is used to generate a speech interruption notification based on the requester information and send the speech interruption notification to the meeting interface of the current speaker.
[0050] The interruption information acquisition module is used to interrupt and collect the interruption speech information of the person requesting the interruption request.
[0051] The speaking strategy generation module is used to generate a speaking strategy based on the interruption information and the speaking status.
[0052] By adopting the above technical solution, the current speaker is determined based on the current meeting content. After the current speaker is determined, the speaker's speaking information and facial information are collected. When a speaking interruption request is received, the current speaker's speaking status is determined based on the speaking information and facial information. If the speaking status allows interruption, the current speaker's speech is interrupted and a prompt is given. If not, a speaking strategy is generated to allow the requester to speak at an appropriate time. This reduces the possibility of multiple people speaking at the same time and disrupting the normal progress of the meeting, thereby improving the efficiency of the meeting.
[0053] Thirdly, this application provides an electronic device that adopts the following technical solution:
[0054] An electronic device includes a processor coupled to a memory;
[0055] The processor is configured to execute a computer program stored in the memory, such that the electronic device executes the computer program of the video conferencing control method according to any one of the first aspects.
[0056] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:
[0057] A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the video conferencing control method according to any one of the first aspects. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating a video conferencing control method provided in an embodiment of this application.
[0059] Figure 2This is a structural block diagram of a video conferencing control device provided in an embodiment of this application.
[0060] Figure 3 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0061] The present application will be further described in detail below with reference to the accompanying drawings.
[0062] This application provides a video conferencing control method, which can be executed by an electronic device. The electronic device can be a server or a terminal device. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet computer, desktop computer, etc., but is not limited to these.
[0063] Figure 1 This is a flowchart illustrating a video conferencing control method provided in an embodiment of this application.
[0064] like Figure 1 As shown, the main process of this method is described below (steps S101 to S108):
[0065] Step S101: Obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content.
[0066] For step S101, obtain the first voice information of the meeting participants and the second voice information of the sender who sent the current meeting content; find the voice information in the first voice information that is the same as the second voice information; determine the current speaker based on the voice information and personnel information.
[0067] In this embodiment, the personnel information of the meeting participants includes the personnel's name, position, voice information, and facial information. The personnel's voice information is the first voice information, which includes the personnel's timbre, pronunciation characteristics, and pronunciation frequency. The personnel's facial information includes facial features and habitual expressions.
[0068] During a meeting, there's a possibility that multiple participants may have their microphones on simultaneously, but only one participant may be speaking. In such cases, it's insufficient to determine the current speaker solely based on microphone availability. Therefore, it's necessary to identify the actual speaker by analyzing the audio information of the person sending the meeting content. This involves collecting the audio from the sender of the meeting content to obtain second audio information. This second audio information is then compared with the first audio information in the participant information. The person whose first audio information matches the second is the current speaker. Once the current speaker is identified, only their microphone remains on, while the microphones of other participants who have their microphones on but are not speaking are muted.
[0069] Step S102: In response to the speech interruption request, obtain the current speaker's speech information and facial information.
[0070] In this embodiment, the microphone status of each meeting participant is detected in real time. When a participant turns on their microphone, a speaking query is sent to the participant's meeting interface. If the query is answered yes, the microphone-on operation is considered a speaking interruption request. When a speaking interruption request is generated, the speaker's speaking information and facial information are collected in detail.
[0071] Step S103: Determine the current speaker's speaking status based on the speaking information and facial information.
[0072] For step S103, the current speaker's speaking speed and volume are determined based on the speaking information; the current speaker's speaking urgency level is determined based on the speaking speed, speaking volume, and preset level rules; the current speaker's emotional state is determined based on facial information and preset emotion rules; and the current speaker's speaking state is determined based on the speaking urgency level and emotional state.
[0073] In this embodiment, the speech information includes the current speaker's speaking speed and volume, and the facial information includes the current speaker's facial expressions. The preset level rules include speech speed level rules and volume level rules. The speech speed level is determined based on the speech speed level rules and the speaking speed, and the volume level is determined based on the speaking volume and the volume level rules. The speech speed level and volume level are added together to obtain the speaking urgency level. The speech speed level rules set a range of speaking word counts, with each range corresponding to a specific speech speed level. The larger the value of the speaking word count range, the higher the corresponding speech speed level. The number of speaking words within a unit of time is collected, and the speaking word count is matched with the speaking word count range. The speech speed level of the successfully matched speaking word count range is the current speaker's speech speed level. The range values of the speaking word count ranges can be the same or decreasing sequentially, and must be set according to the normal and extreme values of human speaking speed, not deviating from the actual range of human speaking speed. Furthermore, the unit of time must be set to an integer number of seconds within one minute, such as 10 seconds, 15 seconds, or 20 seconds, depending on actual needs.
[0074] The volume level rule determines the volume level based on the average decibel value per unit time. Each decibel corresponds to one volume level. The average decibel value is compared with the volume decibel value set in the volume level rule. The volume level corresponding to the volume decibel value set in the volume level rule that is the same as the average decibel value is taken as the volume level of the current speaker.
[0075] After determining the urgency level of speaking, the speaker's emotional state is determined based on their facial expressions. Different facial expressions correspond to different emotional states. The collected facial expressions are matched against preset emotion rules, which specify the emotional state and level for each expression. The matching result reflects the speaker's current emotional state and level. For example, a frowning expression with tense and focused facial muscles corresponds to anger, with an anger level of 5. The speaking state is the sum of the urgency level, emotional state, and emotional state level. It should be noted that the specific emotional state and level need to be set based on actual human responses; no specific limitations are made here.
[0076] Step S104: Determine whether the current speaker can be interrupted immediately based on the speaking status.
[0077] In addition to the speaking status, the network environment is also considered. If either of these conditions is not met, it is determined that the current speaker cannot be interrupted immediately. That is, only when both conditions are met can the current speaker be interrupted immediately. The preset conditions are meeting conditions set according to the current network environment and the ability to communicate normally. They need to be adjusted according to the actual situation.
[0078] Step S105: If the current speaker can be interrupted immediately, obtain the information of the person who requested the interruption request.
[0079] In this embodiment, the requester information includes the requester's name, gender, voice information, and facial information.
[0080] Step S106: Generate a speech interruption notification based on the requester's information and send the speech interruption notification to the current speaker's meeting interface.
[0081] For step S106, obtain the speech content of the current speaker, analyze the speech content, and generate a preset interruption reason; obtain the request time and number of requests for speech interruption; calculate the request frequency based on the request time and number of requests; obtain the preset request level rules, determine the request level based on the preset level rules and request frequency; and generate a speech interruption notification based on the interruption reason and request level.
[0082] In this embodiment, after determining that the current speaker can be interrupted immediately, to ensure that the current speaker's emotions are not affected, the requester is prompted to speak normally. This speech is not a complete speech, but a brief summary of the reason for speaking. The requester's speech is recorded and stored, and semantic analysis is performed on the speech content to determine the key content. At the same time, the aforementioned method for determining the urgency level of the speech is used to determine the urgency level of the speech. A preset interruption reason is generated based on the key content and the urgency level of the speech. Furthermore, the request time and the number of times the speech interruption request is collected and recorded, and the ratio of the request time to the number of requests is calculated. This ratio is used as the request frequency. A preset request level rule is set so that each request frequency corresponds to a request level, and different request frequencies can correspond to different request levels. The preset interruption reason and the request level are combined to generate a speech interruption notification, so that the current speaker can be informed of the reason for being interrupted in a timely manner.
[0083] Step S107: If the current speaker cannot be interrupted immediately, collect the interruption information of the person who requested the interruption request.
[0084] Step S108: Generate a speaking strategy based on interruption information and speaking status.
[0085] For step S108, calculate the relevance between the interrupted speech content and the current meeting content; determine whether the relevance is not less than a preset relevance threshold; if the relevance is not less than the preset relevance threshold, then play the interrupted speech content directly when the current speaker pauses; if the relevance is less than the preset relevance threshold, then generate a speech prompt based on the interrupted speech content and send the speech prompt to the current speaker's meeting interface.
[0086] Furthermore, keywords are extracted from the speech content, and a speech summary is generated based on the extracted keywords; the name of the requester is obtained, and a speech prompt is generated based on the speech summary and the name.
[0087] In this embodiment, when it is determined that the current speaker cannot be interrupted immediately, to ensure that the requester's current speech is not affected by time, the requester needs to be prompted to immediately express the interruption information. All interruption information is stored, and keywords are extracted from it. The extracted keywords are compared with the keywords in the current meeting content to determine relevance. When keywords are the same, the relevance is set to the highest value. When keywords are different but belong to the same type or context, the relevance is determined based on a preset difference relevance value. When the relevance is greater than or equal to a preset relevance threshold, the interrupted speech content will be played immediately whenever the current speaker pauses. A pause is defined as the current speaker not speaking for a certain period of time; for example, if no speaking action is detected within 5 seconds after the current speaker begins speaking, it can be considered a pause. When the relevance is less than the preset relevance threshold, a speech summary is generated based on the keywords extracted from the interruption information. Based on the requester's name, a speech prompt is generated according to the speech summary and name, and sent to the current speaker's meeting interface. The interrupted speech content is played when the current speaker allows it.
[0088] In this embodiment, the content and timing of all meeting participants' remarks are obtained; a meeting report is generated based on the content and timing of their remarks.
[0089] Most video conferences take notes, but due to interruptions in recording, storage, and playback timing, the information may be disjointed when reviewing the notes. Therefore, a meeting report is generated by arranging the remarks of each participant in chronological order. This allows for quick reference to the time and person who made what remarks.
[0090] Figure 2 This is a structural block diagram of a video conferencing control device 200 provided in the application embodiment.
[0091] like Figure 2 As shown, the video conferencing control device 200 mainly includes:
[0092] The current speaker determination module 201 is used to obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content;
[0093] The speech information acquisition module 202 is used to acquire the speech information and facial information of the current speaker in response to a speech interruption request;
[0094] The speaking status confirmation module 203 is used to determine the current speaker's speaking status based on speaking information and facial information;
[0095] The speech interruption judgment module 204 is used to determine whether the current speaker can be interrupted immediately based on the speech status.
[0096] The interruption request acquisition module 205 is used to acquire the information of the person who requested the interruption request.
[0097] Interruption notification generation module 206 is used to generate a speech interruption notification based on the requester information and send the speech interruption notification to the current speaker's meeting interface;
[0098] The interruption information acquisition module 207 is used to interrupt and collect the interruption information of the person requesting the interruption request.
[0099] The speaking strategy generation module 208 is used to generate speaking strategies based on interruption information and speaking status.
[0100] As an optional implementation of this embodiment, the current speaker determination module 201 is specifically used to obtain the first voice information of the meeting participants and the second voice information of the sender of the current meeting content; search for the voice information in the first voice information that is the same as the second voice information; and determine the current speaker based on the voice information and the personnel information.
[0101] As an optional implementation of this embodiment, the speaking status confirmation module 203 is specifically used to determine the current speaker's speaking speed and volume based on speaking information; determine the current speaker's speaking urgency level based on speaking speed, speaking volume and preset level rules; determine the current speaker's emotional state based on facial information and preset emotion rules; and determine the current speaker's speaking status based on speaking urgency level and emotional state.
[0102] As an optional implementation of this embodiment, the interruption notification generation module 206 is specifically used to obtain the speech content of the current speaker, analyze the speech content, generate a preset interruption reason; obtain the request time and number of requests for the speech interruption request; calculate the request frequency based on the request time and number of requests; obtain a preset request level rule, determine the request level based on the preset level rule and the request frequency; and generate a speech interruption notification based on the interruption reason and the request level.
[0103] As an optional implementation of this embodiment, the speaking strategy generation module 208 includes:
[0104] The relevant calculation module is used to calculate the relevance of the interrupted speech to the current meeting content;
[0105] The correlation judgment module is used to determine whether the correlation is not less than a preset correlation threshold;
[0106] The speech playback module is used to directly play the interrupted speech when the current speaker pauses.
[0107] The prompt generation module is used to generate a prompt based on the interrupted speech content and send the prompt to the current speaker's meeting interface.
[0108] In this optional embodiment, the prompt generation module is specifically used to extract keywords from the speech content, generate a speech summary based on the extracted keywords, obtain the name of the requester, and generate a speech prompt based on the speech summary and name.
[0109] As an optional implementation of this embodiment, the video conferencing control device 200 further includes:
[0110] The meeting speech acquisition module is used to acquire the content and duration of all meeting speeches by all participants.
[0111] The meeting report generation module is used to generate meeting reports based on the content and timing of the statements made at the meeting.
[0112] In one example, the module in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0113] For example, when modules in a device can be implemented via a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these modules can be integrated together as a system-on-a-chip (SOC).
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0115] Figure 3 This is a structural block diagram of the electronic device 300 provided in an embodiment of this application.
[0116] like Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.
[0117] The processor 301 controls the overall operation of the electronic device 300 to complete all or part of the steps of the video conferencing control method described above. The memory 302 stores various types of data to support the operation of the electronic device 300. This data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as one or more of Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0118] I / O interface 303 provides an interface between processor 301 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 304 is used for wired or wireless communication between electronic device 300 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 104 may include a Wi-Fi component, a Bluetooth component, and an NFC component.
[0119] The electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the video conferencing control method given in the above embodiments.
[0120] The communication bus 305 may include a path for transmitting information between the aforementioned components. The communication bus 305 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 305 may be divided into an address bus, a data bus, a control bus, etc.
[0121] Electronic device 300 may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers, and may also be servers.
[0122] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the video conferencing control method described above.
[0123] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0124] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0125] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.
Claims
1. A video conferencing control method, characterized in that, include: Obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content; In response to a speech interruption request, obtain the current speaker's speech information and facial information; The current speaker's speaking status is determined based on the spoken information and the facial information; Based on the speaking status, determine whether the current speaker can be interrupted immediately; If the current speaker can be interrupted immediately, the information of the person who requested the interruption request is obtained; A speech interruption notification is generated based on the requester information, and the speech interruption notification is sent to the meeting interface of the current speaker. If the current speaker cannot be interrupted immediately, the interruption information of the person who requested the interruption request will be collected. A speaking strategy is generated based on the interruption information and the speaking status.
2. The method according to claim 1, characterized in that, The process of determining the current speaker based on the personnel information and the current meeting content includes: Obtain the first voice information of the meeting participants and the second voice information of the sender who sent the current meeting content; Find the sound information in the first sound information that is the same as the second sound information; The current speaker is determined based on the voice information and the personnel information.
3. The method according to claim 1, characterized in that, Determining the current speaker's speaking status based on the speaking information and the facial information includes: The speaking speed and volume of the current speaker are determined based on the speaking information. The urgency level of the current speaker is determined based on the speaking speed, the speaking volume, and preset level rules. The current speaker's emotional state is determined based on the facial information and preset emotion rules; The current speaker's speaking status is determined based on the urgency level of the speech and the emotional state.
4. The method according to claim 1, characterized in that, The generation of a speech interruption notification based on the requester information includes: Obtain the speech content of the current speaker, analyze the speech content, and generate a preset interruption reason; Obtain the request time and number of times the speech interruption request was made; Calculate the request frequency based on the request time and the number of requests; Obtain preset request level rules, and determine the request level based on the preset request level rules and the request frequency; A speech interruption notification is generated based on the interruption reason and the request level.
5. The method according to claim 1, characterized in that, The interruption information includes the interruption content; the speech generation strategy based on the interruption information and the speech status includes: Calculate the correlation between the interrupted speech and the current meeting content; Determine whether the correlation is not less than a preset correlation threshold; If the correlation is not less than a preset correlation threshold, the interrupted speech content will be played directly when the current speaker pauses. If the relevance is less than a preset relevance threshold, a speaking prompt is generated based on the interrupted speech content, and the speaking prompt is sent to the current speaker's meeting interface.
6. The method according to claim 5, characterized in that, The generation of a speaking prompt based on the interrupted speech content includes: Keyword extraction is performed on the speech content, and a speech summary is generated based on the extracted keywords; Obtain the name of the requester, and generate a message prompt based on the message summary and the name.
7. The method according to claim 1, characterized in that, Also includes: Obtain the content and duration of all statements made by the participants in the meeting; A meeting report is generated based on the content of the statements made at the meeting and the time when those statements were made.
8. A video conferencing control device, characterized in that, include: The current speaker determination module is used to obtain the personnel information of the meeting participants and the current meeting content, and determine the current speaker based on the personnel information and the current meeting content; The speech information acquisition module is used to acquire the speech information and facial information of the current speaker in response to a speech interruption request; The speaking status confirmation module is used to determine the speaking status of the current speaker based on the speaking information and the facial information; The speech interruption judgment module is used to determine whether the current speaker can be interrupted immediately based on the speech status. The interruption request acquisition module is used to acquire the requester information of the speech interruption request; An interruption notification generation module is used to generate a speech interruption notification based on the requester information and send the speech interruption notification to the meeting interface of the current speaker. The interruption information acquisition module is used to interrupt and collect the interruption speech information of the person requesting the interruption request. The speaking strategy generation module is used to generate a speaking strategy based on the interruption information and the speaking status.
9. An electronic device, characterized in that, Includes a processor, which is coupled to a memory; The processor is configured to execute a computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It includes a computer program or instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Terminal role switching method and device, terminal device and storage medium
CN111193896A
Duplex communication for improving session AI by dynamically responding to interrupted content
CN115700878A