Video conference control method, system and equipment based on remote controller and medium
By automatically determining the speech conditions of audio data in the remote control meeting, determining and transmitting the audio data of the speech terminal, the problem of strong dependence on manual operations in the prior art is solved, and the meeting efficiency and participation experience are improved.
Patent Information
- Application Number
- CN202510508903.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-15
AI Technical Summary
In existing remote control meetings, speaker switching relies on manual operations, resulting in inefficient meetings and poor participation experience, and the main control is concentrated on the host terminal that is not flexible enough.
By obtaining the audio data of the conference terminal in the remote control meeting, it is automatically judged whether the preset speech conditions are met, and the terminal that meets the conditions is determined as a speech terminal, and its audio data is transmitted to other conference terminals, including a judgment method based on the gain value and text content.
The speech switching process is simplified, the meeting efficiency is improved, the response time of speech terminal switching is shortened, the problem of strong dependence on manual operations is solved, and the independent determination and direct participation of speech terminals is realized.
Smart Images

Figure CN120499333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video conferencing, and in particular to a video conferencing control method based on a remote controller, a video conferencing control system based on a remote controller, an electronic device, and a computer-readable storage medium. Background Art
[0002] In existing remote control conferencing technology, speaker switching usually relies on manual operation. For example, the host terminal's remote control must be used to select a terminal as the speaking terminal in the participant list so that other terminals can receive the audio and video data of the terminal.
[0003] In the current technical solution, the main control is concentrated in the host's venue, and the operation is not flexible enough. When other terminals want to speak, they must notify the control personnel of the host's terminal to operate, resulting in low meeting efficiency and poor meeting experience. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention are proposed to provide a remote control-based video conferencing control method, a remote control-based video conferencing control system, an electronic device, and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.
[0005] In order to solve the above problems, an embodiment of the present invention discloses a video conference control method based on a remote controller, the method comprising:
[0006] In a remote control conference, obtaining audio data of at least one first conference terminal;
[0007] Determining whether at least one of the audio data meets a preset speaking condition;
[0008] The first conference terminal corresponding to the audio data meeting the speaking condition is determined as a speaking terminal, and the audio data of the speaking terminal is transmitted to other conference terminals in the remote control conference.
[0009] Optionally, the determining whether at least one audio data meets a preset speaking condition includes:
[0010] detecting a gain value of at least one of the audio data;
[0011] comparing at least one of the gain values with a preset gain threshold;
[0012] If there is at least one gain value greater than or equal to the gain threshold, screening out the largest gain value from the at least one gain value greater than or equal to the gain threshold;
[0013] The audio data corresponding to the screened gain value is used as the audio data that meets the speaking condition.
[0014] Optionally, the determining whether at least one audio data meets a preset speaking condition includes:
[0015] identifying text content of at least one of the audio data;
[0016] Detecting whether at least one of the text contents contains a preset speech keyword;
[0017] The audio data corresponding to the text content containing the speech keyword is regarded as the audio data meeting the speech condition.
[0018] Optionally, transmitting the audio data of the speaking terminal to other conference terminals in the remote control conference includes:
[0019] generating a switching instruction for the speaking terminal, wherein the switching instruction includes terminal identification information of the speaking terminal;
[0020] The audio data of the speaking terminal is transmitted to the other conference terminals according to the switching instruction.
[0021] Optionally, the determining whether at least one audio data meets a preset speaking condition includes:
[0022] transmitting at least one of the audio data to a second terminal;
[0023] The second terminal is used to determine whether at least one of the audio data meets the speaking condition.
[0024] Optionally, after transmitting the at least one audio data to the second terminal, the method further includes:
[0025] detecting that the second terminal is in an offline state;
[0026] Filtering a third terminal from the remote control conference, and transmitting at least one audio data to the third terminal;
[0027] The third terminal is used to determine whether at least one audio data meets the speaking condition.
[0028] Optionally, screening out a third terminal from the remote control conference includes:
[0029] Acquiring current status information or preset priority information of other conference terminals in the remote control conference except the second terminal;
[0030] The third terminal is screened out according to the current state information or the priority information.
[0031] The embodiment of the present invention further discloses a video conference control system based on a remote controller, the system comprising:
[0032] An audio data acquisition module, configured to acquire audio data of at least one first conference terminal in a remote control conference;
[0033] A speaking condition judging module, configured to judge whether at least one of the audio data meets a preset speaking condition;
[0034] The speaking data transmission module is configured to determine the first conference terminal corresponding to the audio data meeting the speaking condition as a speaking terminal, and transmit the audio data of the speaking terminal to other conference terminals in the remote control conference.
[0035] Optionally, the speaking condition judgment module includes:
[0036] a gain detection module, configured to detect a gain value of at least one of the audio data;
[0037] a gain comparison module, configured to compare at least one of the gain values with a preset gain threshold;
[0038] a gain screening module, configured to, if there is at least one gain value greater than or equal to the gain threshold, screen out the maximum gain value from the at least one gain value greater than or equal to the gain threshold;
[0039] The first audio data determining module is configured to use the audio data corresponding to the screened gain value as the audio data meeting the speaking condition.
[0040] Optionally, the speaking condition judgment module includes:
[0041] a text recognition module, configured to recognize text content of at least one of the audio data;
[0042] A text detection module, configured to detect whether at least one of the text contents contains a preset speech keyword;
[0043] The second audio data determination module is configured to take the audio data corresponding to the text content containing the speech keyword as the audio data meeting the speech condition.
[0044] Optionally, the speech data transmission module includes:
[0045] An instruction generating module, configured to generate a switching instruction for the speaking terminal, wherein the switching instruction includes terminal identification information of the speaking terminal;
[0046] An audio transmission module is used to transmit the audio data of the speaking terminal to the other conference terminals according to the switching instruction.
[0047] Optionally, the speaking condition judgment module includes:
[0048] A data transmission module, configured to transmit at least one of the audio data to a second terminal;
[0049] The data judgment module is used to use the second terminal to judge whether at least one of the audio data meets the speaking condition.
[0050] Optionally, the system further comprises:
[0051] A state detection module, configured to detect that the second terminal is in an offline state after the data transmission module transmits at least one audio data to the second terminal;
[0052] a terminal screening module, configured to screen out a third terminal from the remote control conference and transmit at least one audio data to the third terminal;
[0053] The condition judgment module is used to use the third terminal to judge whether at least one audio data meets the speaking condition.
[0054] Optionally, the terminal screening module includes:
[0055] A status priority acquisition module, configured to acquire current status information or preset priority information of other conference terminals in the remote control conference except the second terminal;
[0056] The third terminal screening module is configured to screen out the third terminal according to the current state information or the priority information.
[0057] An embodiment of the present invention also discloses an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enables the electronic device to execute the remote control-based video conferencing control method described above.
[0058] An embodiment of the present invention further discloses a computer-readable storage medium, which stores a computer program that enables a processor to execute the video conference control method based on a remote controller as described above.
[0059] The embodiments of the present invention include the following advantages:
[0060] The remote control-based video conferencing control solution provided by an embodiment of the present invention obtains audio data of at least one first conference terminal in a remote control conference; determines whether at least one audio data meets preset speaking conditions; determines the first conference terminal corresponding to the audio data that meets the speaking conditions as the speaking terminal, and transmits the audio data of the speaking terminal to other conference terminals in the remote control conference.
[0061] Compared with the background technology, the embodiments of the present invention have the following beneficial effects:
[0062] The embodiment of the present invention automatically determines whether the audio data meets the speaking conditions, eliminates the tedious steps of manual operation, simplifies the speaking switching process, and solves the problem of strong dependence on manual operation in the background technology. Automatically determining the speaking terminal and transmitting audio data shortens the response time of speaking terminal switching, avoids the operation delay of notifying the host terminal, thereby significantly improving the efficiency of the meeting and solving the problem of low meeting efficiency in the background technology. The determination of the speaking terminal no longer relies on the centralized control of the host terminal, but is completed autonomously based on preset speaking conditions, allowing other terminals to participate in the speech more directly, solving the problem of centralized control in the background technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a flowchart of the steps of a remote control-based video conferencing control method according to an embodiment of the present invention;
[0064] Figure 2 This is a schematic diagram of the principle of a conference speech management solution based on a visual network according to an embodiment of the present invention;
[0065] Figure 3 This is a structural block diagram of a video conferencing control system based on a remote controller according to an embodiment of the present invention. DETAILED DESCRIPTION
[0066] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0067] An embodiment of the present invention provides a remote control-based video conferencing control method, one of the purposes of which is to achieve automatic switching of speaking terminals in a remote control conference. Specifically, the method obtains audio data from at least one first conference terminal, determines whether it meets preset speaking conditions, and determines the terminal that meets the conditions as the speaking terminal. Its audio data is transmitted to other conference terminals, thereby improving conference efficiency and meeting experience. For example, in a visual network environment, the audio data of other participating terminals except the host terminal is transmitted to the host terminal. The host terminal detects the audio gain value and determines the participating terminal with the loudest sound as the speaking terminal. Then, its audio and video data are distributed to other terminals, achieving an autonomous switching effect based on sound, overcoming the limitations of traditional manual operation.
[0068] Reference Figure 1 , shows a flowchart of the steps of a remote control-based video conferencing control method according to an embodiment of the present invention. The remote control-based video conferencing control method can be applied to systems such as conference systems, video conferencing systems, and conference management systems (hereinafter referred to as systems). The remote control-based video conferencing control method may specifically include the following steps:
[0069] Step 101: In a remote control conference, obtain audio data of at least one first conference terminal.
[0070] A remote control conference is a type of conference controlled by a terminal remote control. It typically supports multiple participating terminals, with a maximum limit of, for example, no more than 10. In this step, the system first identifies all participating terminals. These terminals can be any participating device in the conference, such as video conferencing terminals installed in the conference room or portable mobile devices. The system then acquires audio data from each terminal via a network connection. This audio data is typically real-time voice input from a participant's microphone on the terminal device. To ensure the integrity and accuracy of the audio data, the system configures the appropriate number of audio channels, for example, setting the number of audio channels to support simultaneous audio data transmission from multiple terminals, ensuring that audio data from each terminal is captured in real time. Furthermore, the system can append identification information for the first conference terminal, such as the terminal number or device name, to the audio data stream, enabling accurate differentiation of audio sources from different terminals in subsequent steps. When acquiring audio data, the system must also ensure the continuity of data collection to avoid interruptions due to network fluctuations or device failures, which could affect the accuracy of subsequent judgments.
[0071] Step 102: Determine whether at least one audio data meets a preset speaking condition.
[0072] The preset speaking conditions are a set of rules or thresholds predefined by the system that are used to filter audio data suitable for speaking terminals. A common judgment method is based on the size of the audio gain value. The system detects the gain value of the audio data of each first conference terminal and compares these gain values with the preset gain threshold. If the gain value of an audio data item is greater than or equal to the preset gain threshold, it is considered to potentially meet the speaking conditions. Furthermore, if the gain values of multiple audio data items exceed the gain threshold, the system selects the one with the largest gain value and identifies the corresponding audio data as meeting the speaking conditions. This approach ensures that the first conference terminal with the loudest voice is preferentially identified as a potential speaking terminal. In addition, the system can also use other auxiliary judgment conditions, such as identifying the text content in the audio data and analyzing it through speech-to-text technology to see if it contains preset speaking keywords (such as "Attention" or "Start speaking"). If so, the audio data is considered to meet the speaking conditions. To ensure real-time judgment, the system periodically performs audio gain value detection and condition judgment at preset intervals (e.g., once per second) to ensure timely response to changes in sound or speaking needs during the meeting. For example, in a remote control conference, assuming there are five first conference terminals participating, the system detects the audio gain value once a second. If it is found that the gain value of one terminal is significantly higher than that of other terminals and exceeds the preset gain threshold, the audio data of the terminal will be marked as meeting the speaking conditions.
[0073] Step 103: Determine the first conference terminal corresponding to the audio data meeting the speaking condition as the speaking terminal, and transmit the audio data of the speaking terminal to other conference terminals in the remote control conference.
[0074] Based on the results of the previous step, the system first identifies audio data that meets the preset speaking conditions. Using the identifier attached to the audio data stream, the system accurately locates the first conference terminal corresponding to the audio data and sets it as the speaking terminal. The speaking terminal role means that the terminal's audio data will be preferentially transmitted and displayed to other participants in the remote conference, enabling dynamic switching of speakers within the conference. After determining the speaking terminal, the system retrieves the speaking terminal's audio data stream and distributes it via a network channel to the other conference terminals in the remote conference, ensuring that all participants can hear the speaking terminal's voice content in real time. Furthermore, during the transmission process, the system generates a switching instruction containing the speaking terminal's identifier to adjust the audio data stream's transmission path, ensuring that the audio data is accurately delivered to the other conference terminals without miscommunication or confusion. For example, in a remote conference, if the system determines that the audio data from the first conference terminal, numbered "Terminal A," meets the speaking conditions, the system will set "Terminal A" as the speaking terminal, retrieve its audio data stream, and distribute it to the other four conference terminals, ensuring that participants at the other terminals can hear "Terminal A's" speech in real time. The system also needs to ensure that audio data transmission is not interrupted or significantly delayed when the speaking terminal switches, maintaining the continuity and smoothness of the meeting. If a network anomaly or device failure is detected during transmission, the system will take measures such as retransmission or switching to an alternative channel to ensure reliable delivery of audio data.
[0075] The remote control-based video conferencing control solution provided by an embodiment of the present invention obtains audio data of at least one first conference terminal in a remote control conference; determines whether at least one audio data meets preset speaking conditions; determines the first conference terminal corresponding to the audio data that meets the speaking conditions as the speaking terminal, and transmits the audio data of the speaking terminal to other conference terminals in the remote control conference.
[0076] Compared with the background technology, the embodiments of the present invention have the following beneficial effects:
[0077] The embodiment of the present invention automatically determines whether the audio data meets the speaking conditions, eliminates the tedious steps of manual operation, simplifies the speaking switching process, and solves the problem of strong dependence on manual operation in the background technology. Automatically determining the speaking terminal and transmitting audio data shortens the response time of speaking terminal switching, avoids the operation delay of notifying the host terminal, thereby significantly improving the efficiency of the meeting and solving the problem of low meeting efficiency in the background technology. The determination of the speaking terminal no longer relies on the centralized control of the host terminal, but is completed autonomously based on preset speaking conditions, allowing other terminals to participate in the speech more directly, solving the problem of centralized control in the background technology.
[0078] In an exemplary embodiment of the present invention, an implementation method for determining whether at least one audio data meets a preset speaking condition is: detecting a gain value of at least one audio data; comparing at least one gain value with a preset gain threshold; if there is at least one gain value greater than or equal to the gain threshold, filtering out the maximum gain value from at least one gain value greater than or equal to the gain threshold; and using the audio data corresponding to the filtered gain value as audio data that meets the speaking condition.
[0079] The system performs real-time detection on the audio data of at least one first conference terminal obtained from the remote control conference and extracts the gain value of each piece of audio data. This process usually relies on audio signal processing technology and can accurately reflect the volume performance of each first conference terminal. The system compares the at least one extracted gain value with the preset gain threshold one by one. The gain threshold is a standard value pre-set by the system and is used to filter out audio data with too low volume or background noise, ensuring that only audio data with a certain volume level is considered. Then, if the detection result shows that there are one or more gain values greater than or equal to the preset gain threshold, the system will further filter out the maximum gain value from these qualified gain values to ensure that the audio data finally selected corresponds to the first conference terminal with the strongest voice. Finally, the system determines the audio data corresponding to the filtered maximum gain value as audio data that meets the preset speaking conditions, laying the foundation for subsequently setting the first conference terminal corresponding to the audio data as the speaking terminal.
[0080] For example, in a remote control conference, suppose there are five first conference terminals participating, and the system detects that their audio data gain values are 30, 45, 20, 60 and 25 (in decibels), respectively, and the preset gain threshold is 40 decibels. The system will first filter out two audio data with gain values of 45 and 60, and then select the audio data with a gain value of 60 as the audio data that meets the speaking conditions, indicating that the first conference terminal corresponding to the audio data has the strongest sound and is most suitable as a speaking terminal.
[0081] This implementation method uses gain value-based detection and screening, and the system can automatically identify the first conference terminal with the loudest voice and determine its corresponding audio data as audio data that meets the speaking conditions. This process does not require human intervention, significantly improving the efficiency and accuracy of determining the speaking terminal, while avoiding the tedious manual operation and delay problems in traditional remote control meetings, thereby optimizing the smoothness of the meeting and the meeting experience.
[0082] In an exemplary embodiment of the present invention, an implementation method for determining whether at least one audio data meets the preset speaking conditions is: identifying the text content of at least one audio data; detecting whether at least one text content contains preset speaking keywords; and taking the audio data corresponding to the text content containing the speaking keywords as the audio data that meets the speaking conditions.
[0083] The system processes the audio data of at least one first conference terminal obtained from the remote control conference and converts the audio data into corresponding text content using voice recognition technology. This process usually relies on advanced speech-to-text technology (Speech-to-Text, abbreviated as STT), which can convert the participant's voice into text information for analysis in real time, ensuring the accuracy and real-time nature of the conversion. The system detects the at least one extracted text content one by one and analyzes whether it contains preset speech keywords. These speech keywords are a set of specific words or phrases pre-defined by the system, such as "start speaking", "please pay attention" or "I have an opinion", which are used to indicate that the participant intends to become a speaking terminal. Finally, if it is detected that a certain text content contains preset speech keywords, the system will determine the audio data corresponding to the text content as audio data that meets the preset speaking conditions, thereby providing a basis for subsequently setting the first conference terminal corresponding to the audio data as a speaking terminal.
[0084] For example, in a remote control conference, suppose there are four participating first-party conference terminals. The system uses speech recognition technology to convert their audio data into text content, including "The weather is nice today," "Please note, I have an important report," "Hello everyone," and "Wait a moment." If the preset speech keywords include "Please note," the system will detect that the second text content contains this keyword and mark the corresponding audio data as meeting the speech conditions, indicating that the participant at the first-party conference terminal clearly expressed their intention to speak. When performing recognition and detection, the system must also consider the accuracy of speech recognition. For example, by optimizing algorithms or introducing noise filtering mechanisms, it can avoid text content recognition errors caused by background noise or accent differences, ensuring that each piece of audio data is fairly and accurately evaluated.
[0085] This implementation method uses text content recognition and speech keyword detection to enable the system to intelligently judge the speaking intention of the first conference terminal and determine the audio data that meets the speaking conditions. This process is not only highly automated, but also can accurately capture the participants' clear expressions, significantly improving the targeted selection of speaking terminals and the efficiency of conference interaction, thereby optimizing the overall experience of remote control conferences.
[0086] In an exemplary embodiment of the present invention, an implementation method for transmitting the audio data of the speaking terminal to other conference terminals in a remote control conference is: generating a switching instruction for the speaking terminal, the switching instruction including the terminal identification information of the speaking terminal; and transmitting the audio data of the speaking terminal to other conference terminals according to the switching instruction.
[0087] After the system determines that a first conference terminal is the speaking terminal, it generates a dedicated switching instruction. This switching instruction contains the terminal identification information of the speaking terminal, such as the terminal number, device name, or other unique identifier. It is used to clearly indicate which terminal's audio data needs to be transmitted and displayed to other participants. The design of the switching instruction ensures that the system can accurately identify the target terminal, avoiding confusion or errors during the transmission process. Based on the generated switching instruction, the system adjusts the audio data transmission path, distributing the audio data stream of the speaking terminal via the network channel to other conference terminals in the remote control conference, ensuring that all participants at other conference terminals can receive the voice content of the speaking terminal in real time.
[0088] For example, in a remote control conference, assuming that the system has determined that the first conference terminal numbered "Terminal B" is the speaking terminal, the system will generate a switching instruction containing the terminal identification information of "Terminal B", and then distribute the audio data stream of "Terminal B" to the other five conference terminals according to the instruction, so that the participants at other terminals can clearly hear the speech content of the participant at "Terminal B". In addition, when executing the transmission, the system must also consider the issue of network stability. For example, by setting up a backup transmission channel or a data retransmission mechanism, it can ensure that in the event of network fluctuations or equipment failures, the audio data of the speaking terminal can still be smoothly delivered to other conference terminals to avoid conference interruptions or audio loss. The system can also track the status of each speaking terminal switch and audio data transmission through the logging function of the switching instruction, which is convenient for subsequent analysis and optimization, and ensures the traceability and reliability of the entire transmission process.
[0089] This implementation generates a switching instruction containing terminal identification information and transmits the audio data of the speaking terminal accordingly. The system can accurately and efficiently distribute the audio data to other conference terminals. It not only realizes the automation and real-time switching of the speaking terminal, but also avoids the delay and error risk of traditional manual operation, significantly improving the fluency of remote control conferences and the meeting experience.
[0090] In an exemplary embodiment of the present invention, one implementation method for determining whether at least one audio data meets a preset speaking condition is: transmitting at least one audio data to a second terminal; and using the second terminal to determine whether the at least one audio data meets the speaking condition.
[0091] After the system obtains audio data from at least one first conference terminal in a remote control conference, it transmits the audio data to the second terminal through a network channel. The second terminal is usually a device with strong computing power or special configuration, such as the host terminal, which is responsible for centrally processing the analysis task of the audio data. In order to ensure the reliability and real-time performance of the transmission, the system will configure a sufficient number of audio channels to support the simultaneous transmission of audio data from multiple first conference terminals, and attach the identification information of the first conference terminal (such as the terminal number or device name) to the audio data stream so that the second terminal can accurately distinguish audio data from different sources. After receiving at least one audio data, the second terminal uses its built-in analysis mechanism or algorithm to determine whether each piece of audio data meets the preset speaking conditions. These speaking conditions may include whether the audio gain value exceeds a preset threshold, whether the audio content contains specific keywords, and other standards. The second terminal will evaluate each piece of audio data one by one and feedback the judgment result to the system for subsequent determination of the speaking terminal.
[0092] For example, in a remote control conference, suppose there are six participating first-party terminals. The system transmits their audio data to a pre-defined second-party terminal. This second-party terminal, which may be a dedicated conference control device, is responsible for detecting the gain value of each audio data and determining which audio data meets the speaking conditions. For example, if the audio gain value of a first-party terminal is the highest and exceeds a preset threshold, it will be marked as meeting the conditions. Furthermore, the system must ensure the integrity of the audio data during transmission to avoid misjudgment by the second-party terminal due to network fluctuations or data loss. Furthermore, the second-party terminal must have a fault-tolerant mechanism, such as multiple detection or a comprehensive judgment based on multiple conditions, to ensure the accuracy and reliability of the results.
[0093] This implementation method realizes distributed processing by transmitting audio data to the second terminal and using its specialized processing capabilities to determine whether the speaking conditions are met. This not only reduces the computing burden of other devices, but also uses the optimization algorithm of the second terminal to improve the accuracy and efficiency of the judgment, thereby ensuring the rationality of the speaking terminal selection and optimizing the overall control effect of the remote control conference.
[0094] In an exemplary embodiment of the present invention, after transmitting at least one audio data to the second terminal, an implementation method is: detecting that the second terminal is in an offline state; screening out a third terminal from the remote control conference, transmitting at least one audio data to the third terminal; and using the third terminal to determine whether the at least one audio data meets the speaking conditions.
[0095] After the system transmits the audio data of at least one first conference terminal to the second terminal, it will continuously detect the online status of the second terminal to determine whether it is offline. The offline status may be caused by network interruption, equipment failure or power problem. The system uses heartbeat signal detection or status feedback mechanism to confirm whether the second terminal can normally receive and process audio data. If it is detected that the second terminal is offline, the system will immediately select a new terminal from other terminals in the remote control conference as the third terminal. The screening process can be based on preset priority rules (such as the computing power of the terminal or network stability) or current status information (such as online time or device load) to ensure that the third terminal has sufficient processing power and reliability to take over the task of the second terminal. After the screening is completed, the system will retransmit the audio data of at least one first conference terminal to the third terminal, and append the identification information of the first conference terminal (such as terminal number or device name) to the audio data stream so that the third terminal can accurately distinguish audio data from different sources. After receiving the audio data, the third terminal uses its built-in analysis mechanism or algorithm to determine whether each piece of audio data meets the preset speaking conditions. These speaking conditions may include whether the audio gain value exceeds the preset threshold or whether the audio content contains specific keywords. The third terminal will evaluate each piece of audio data one by one and feedback the judgment results to the system for subsequent determination of the speaking terminal.
[0096] For example, in a remote control conference, suppose the system initially transmits audio data from five primary conference endpoints to a secondary endpoint for evaluation. However, it detects that the secondary endpoint is offline due to a network issue. The system then selects a device numbered "Terminal C" from the other endpoints based on a preset priority level as the third endpoint and retransmits the audio data to "Terminal C." Terminal C then checks the gain of each audio data stream to determine which audio data meets the speaking criteria. Furthermore, the system must ensure the integrity of the audio data during transmission to avoid data loss due to multiple transmissions. Furthermore, the third endpoint must have a fault-tolerant mechanism to ensure the accuracy and reliability of the evaluation results.
[0097] This implementation method detects the offline status of the second terminal and promptly selects the third terminal to take over its task. The system can ensure the continuity of the audio data judgment process and avoid the interruption of conference control due to the offline status of the second terminal. This mechanism significantly improves the stability and robustness of the remote control conference, thereby ensuring the smooth progress of the conference process.
[0098] In an exemplary embodiment of the present invention, an implementation method of screening out a third terminal from a remote control conference is: obtaining current status information or preset priority information of other conference terminals in the remote control conference except the second terminal; and screening out the third terminal based on the current status information or priority information.
[0099] The system obtains the current status information or preset priority information of all conference terminals in the remote control conference, excluding the second terminal. This current status information typically includes real-time data such as the online status, network connection quality, device load, and battery level of other conference terminals. This information can reflect the terminal's current ability to operate stably. The preset priority information may be a set of predefined rules or levels, such as ranking the importance of other conference terminals based on their hardware performance, historical usage history, or user-specified priority. The system collects this information through database queries or real-time detection mechanisms to ensure comprehensive and accurate data. Based on the current status information or priority information obtained, the system conducts a comprehensive evaluation and screening of other conference terminals, ultimately determining the most suitable device for the third terminal. The screening process may prioritize the terminal with the most stable network connection and the lowest device load as shown in the current status information, or select the terminal with the highest ranking in the priority information. If there is a conflict between the current status information and the priority information, the system may use a weighted algorithm or preset rules to balance the two to ensure the rationality of the screening results.
[0100] For example, in a remote control conference, if the second terminal goes offline due to a malfunction, the system needs to select a third terminal from the other five conference terminals. The system first obtains the current status information of these terminals and discovers that one terminal has the lowest network latency and a device load of only 20%. It also ranks second in preset priority, just behind another terminal with slightly worse network status. After a comprehensive evaluation, the system selects the terminal with the best network status as the third terminal. Furthermore, the system must consider real-time information updates during the screening process. For example, by periodically refreshing the status data of other conference terminals, it can avoid screening errors caused by information lags. The system can also implement a backup mechanism for screening failures. If the currently selected third terminal also experiences problems within a short period of time, the screening process will be re-executed to ensure continuity of conference control.
[0101] This implementation method obtains the current status information or priority information of other conference terminals and filters the third terminal accordingly. The system can ensure that the selected third terminal has the best operating conditions or meets the preset importance ranking. This mechanism significantly improves the adaptability and stability of remote control conferences in the event of equipment failure or offline, thereby ensuring the smooth progress of the conference process.
[0102] Based on the above description of an embodiment of a video conference control method based on a remote control, the following introduces a conference speech management solution based on a visual network, which aims to realize automatic switching of speaking terminals in remote control conferences, and improve meeting efficiency and meeting experience. This solution involves the following key roles and equipment: a specific terminal (host terminal A), which is responsible for receiving and processing audio data from other first conference terminals, and determining the speaking terminal. Ordinary terminals (ordinary terminals B, C, D, etc.), which receive audio and video data from speaking terminals as participants. Monitoring equipment, used to monitor the status of the meeting. The conference management system is responsible for the control and management of the conference and the switching scheduling of audio and video streams.
[0103] Reference Figure 2 , showing a schematic diagram of the principles of a conference speech management solution based on visual networking in an embodiment of the present invention. Figure 2 The figure shows the audio data transmission relationship between specific terminals and ordinary terminals, that is, the specific terminal receives audio data from multiple ordinary terminals (for example, the audio channel of each terminal is configured to support multi-channel transmission), and distributes the audio and video data of the determined speaking terminal to other terminals through the conference management system.
[0104] After the remote control conference is started, the conference management system will transmit the audio data of other first conference terminals except the specific terminal to the specific terminal. To support the concurrent transmission of audio data from multiple terminals, the system configures the number of audio channels to a sufficient number (such as adjusting from 4 to 9) to ensure that the specific terminal can receive audio data from multiple first conference terminals at the same time. In addition, the system appends terminal identification information (such as terminal number) to the audio data stream through the visual networking protocol so that the specific terminal can accurately distinguish audio data from different sources.
[0105] The specific terminal checks the gain value of the audio data received from each first conference terminal every second to determine whether it meets the preset speaking conditions. Specifically, the gain value of each audio data point is compared with a preset gain threshold. If a gain value greater than or equal to the gain threshold exists, the maximum gain value is selected and the audio data corresponding to that gain value is determined to meet the speaking conditions. The specific terminal encapsulates the screening results into a switching instruction containing the terminal identification information of the speaking terminal and transmits it to the conference management system via transparent transmission.
[0106] Upon receiving the switch command, the conference management system sets the corresponding first conference terminal as the speaking terminal and distributes its audio and video data streams via the visual network to the other conference terminals in the remote conference, ensuring that all participants can listen to and view the content from the speaking terminal in real time. During the transmission process, the system utilizes the high-speed transmission characteristics of the visual network to ensure low latency and high definition of audio and video data.
[0107] If the system detects that a specific terminal is offline, the conference management system automatically selects a new terminal from the other conference terminals as the designated terminal. This selection process involves obtaining the current status information (such as network connection quality and device load) or preset priority information of the other conference terminals, and based on this information, selecting the most suitable terminal as the new designated terminal. The system then transmits the audio data from the other first conference terminals to the new designated terminal, continuing the audio data processing and speech condition determination tasks.
[0108] For example, in a remote control conference based on a visual network, there are a total of 6 first conference terminals participating, one of which is designated as a specific terminal. The conference management system transmits the audio data of the other 5 terminals to the specific terminal. The specific terminal detects the audio gain value every second and finds that the audio data gain value of "Terminal C" is the highest and exceeds the preset threshold. It then generates a switching instruction containing the terminal identification information of "Terminal C" and sends it to the conference management system. The conference management system sets "Terminal C" as the speaking terminal and distributes its audio and video data to other terminals. If the specific terminal goes offline due to network problems, the system selects "Terminal D" with the most stable network as the new specific terminal based on the current status information and continues the audio data processing process.
[0109] This embodiment, through automatic audio data processing and dynamic switching of speaking terminals in a visual networking environment, enables voice intensity-based speech management in remote-controlled conferences, overcoming the tediousness and delays of traditional manual operations and significantly improving meeting efficiency and the participant experience. Furthermore, a specific terminal offline processing mechanism ensures the continuity and stability of conference control, enhancing the robustness of the system.
[0110] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0111] Reference Figure 3 , shows a structural block diagram of a video conference control system based on a remote controller according to an embodiment of the present invention. The video conference control system based on a remote controller may specifically include the following modules.
[0112] The audio data acquisition module 31 is used to acquire audio data of at least one first conference terminal in a remote control conference;
[0113] A speaking condition judging module 32 is configured to judge whether at least one of the audio data meets a preset speaking condition;
[0114] The speaking data transmission module 33 is configured to determine the first conference terminal corresponding to the audio data meeting the speaking condition as a speaking terminal, and transmit the audio data of the speaking terminal to other conference terminals in the remote control conference.
[0115] In an exemplary embodiment of the present invention, the speaking condition determination module 32 includes:
[0116] a gain detection module, configured to detect a gain value of at least one of the audio data;
[0117] a gain comparison module, configured to compare at least one of the gain values with a preset gain threshold;
[0118] a gain screening module, configured to, if there is at least one gain value greater than or equal to the gain threshold, screen out the maximum gain value from the at least one gain value greater than or equal to the gain threshold;
[0119] The first audio data determining module is configured to use the audio data corresponding to the screened gain value as the audio data meeting the speaking condition.
[0120] In an exemplary embodiment of the present invention, the speaking condition determination module 32 includes:
[0121] a text recognition module, configured to recognize text content of at least one of the audio data;
[0122] A text detection module, configured to detect whether at least one of the text contents contains a preset speech keyword;
[0123] The second audio data determination module is configured to take the audio data corresponding to the text content containing the speech keyword as the audio data meeting the speech condition.
[0124] In an exemplary embodiment of the present invention, the speech data transmission module 33 includes:
[0125] An instruction generating module, configured to generate a switching instruction for the speaking terminal, wherein the switching instruction includes terminal identification information of the speaking terminal;
[0126] An audio transmission module is used to transmit the audio data of the speaking terminal to the other conference terminals according to the switching instruction.
[0127] In an exemplary embodiment of the present invention, the speaking condition determination module 32 includes:
[0128] A data transmission module, configured to transmit at least one of the audio data to a second terminal;
[0129] The data judgment module is used to use the second terminal to judge whether at least one of the audio data meets the speaking condition.
[0130] In an exemplary embodiment of the present invention, the system further comprises:
[0131] A state detection module, configured to detect that the second terminal is in an offline state after the data transmission module transmits at least one audio data to the second terminal;
[0132] a terminal screening module, configured to screen out a third terminal from the remote control conference and transmit at least one audio data to the third terminal;
[0133] The condition judgment module is used to use the third terminal to judge whether at least one audio data meets the speaking condition.
[0134] In an exemplary embodiment of the present invention, the terminal screening module includes:
[0135] A status priority acquisition module, configured to acquire current status information or preset priority information of other conference terminals in the remote control conference except the second terminal;
[0136] The third terminal screening module is configured to screen out the third terminal according to the current state information or the priority information.
[0137] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0138] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0139] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0141] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0143] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0144] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0145] The above is a detailed introduction to a remote control-based video conferencing control method and a remote control-based video conferencing control system provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A video conference control method based on a remote controller, characterized in that: The method comprises: In a remote control conference, obtaining audio data of at least one first conference terminal; Determining whether at least one of the audio data meets a preset speaking condition; The first conference terminal corresponding to the audio data meeting the speaking condition is determined as a speaking terminal, and the audio data of the speaking terminal is transmitted to other conference terminals in the remote control conference.
2. The method according to claim 1, characterized in that The determining whether at least one of the audio data meets the preset speaking condition includes: detecting a gain value of at least one of the audio data; comparing at least one of the gain values with a preset gain threshold; If there is at least one gain value greater than or equal to the gain threshold, screening out the largest gain value from the at least one gain value greater than or equal to the gain threshold; The audio data corresponding to the screened gain value is used as the audio data that meets the speaking condition.
3. The method according to claim 1 or 2, characterized in that The determining whether at least one of the audio data meets the preset speaking condition includes: identifying text content of at least one of the audio data; Detecting whether at least one of the text contents contains a preset speech keyword; The audio data corresponding to the text content containing the speech keyword is regarded as the audio data meeting the speech condition.
4. The method according to claim 1, wherein The transmitting the audio data of the speaking terminal to other conference terminals in the remote control conference includes: generating a switching instruction for the speaking terminal, wherein the switching instruction includes terminal identification information of the speaking terminal; The audio data of the speaking terminal is transmitted to the other conference terminals according to the switching instruction.
5. The method according to claim 1, characterized in that The determining whether at least one of the audio data meets a preset speaking condition includes: transmitting at least one of the audio data to a second terminal; The second terminal is used to determine whether at least one of the audio data meets the speaking condition.
6. The method according to claim 1, wherein After transmitting at least one of the audio data to the second terminal, the method further includes: detecting that the second terminal is in an offline state; Filtering a third terminal from the remote control conference, and transmitting at least one audio data to the third terminal; The third terminal is used to determine whether at least one audio data meets the speaking condition.
7. The method according to claim 6, characterized in that The step of screening out a third terminal from the remote control conference includes: Acquiring current status information or preset priority information of other conference terminals in the remote control conference except the second terminal; The third terminal is screened out according to the current state information or the priority information.
8. A video conference control system based on a remote controller, characterized in that: The system comprises: An audio data acquisition module, configured to acquire audio data of at least one first conference terminal in a remote control conference; A speaking condition judging module, configured to judge whether at least one of the audio data meets a preset speaking condition; The speaking data transmission module is configured to determine the first conference terminal corresponding to the audio data meeting the speaking condition as a speaking terminal, and transmit the audio data of the speaking terminal to other conference terminals in the remote control conference.
9. An electronic device, characterized in that: include: one or more processors; and One or more machine-readable media having instructions stored thereon, when executed by the one or more processors, enable the electronic device to execute the video conference control method based on the remote control according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer program stored therein enables the processor to execute the video conference control method based on the remote controller as described in any one of claims 1 to 7.