Control method and system of online conference system, electronic equipment and storage medium
Through the online conference system control method based on multimodal perception, users' voice and gesture information are collected in real time, and equipment parameters and system resources are dynamically adjusted, which solves the problems of insufficient equipment status perception, low energy consumption utilization, lack of user behavior response and insufficient energy use efficiency in existing online conference systems, achieving high-quality user experience and efficient resource utilization.
Patent Information
- Application Number
- CN202510505100.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing online conference systems have challenges in insufficient device status perception, low energy utilization rate of equipment, lack of user behavior response, and insufficient energy use efficiency.
Using a control method based on multimodal perception, users' voice information and gesture information from the participating end are collected in real time, user status is determined based on these information, and equipment parameters and system resources are dynamically adjusted to realize intelligent resource scheduling.
By dynamically adjusting equipment parameters and system resources, optimizing user experience and system performance, improving energy usage efficiency, and achieving a balance between high-quality performance and resource utilization.
Smart Images

Figure CN120066797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of online meetings. Specifically, it relates to a control method and system for an online meeting system, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of Internet technology and communication technology, application scenarios such as real-time audio and video communication, online education, and remote meetings have become increasingly popular.
[0003] However, with the continuous improvement of users' requirements for high-definition picture quality, ultra-low latency, and immersive interaction, existing online meeting systems face great challenges in aspects such as audio and video meeting optimization, device control algorithms, dynamic scheduling of data resources, and user experience.
[0004] For example, at the level of audio and video meeting optimization, existing technologies have problems such as insufficient device status perception, low utilization rate of device energy consumption, and lack of user behavior response; for another example, at the levels of device control algorithms and dynamic scheduling of data resources, existing technologies also have problems such as inability to predict device failures and dynamically adjust the cooling strategy, resulting in insufficient energy use efficiency.
[0005] The content in the background art section is only the technology known to the applicant and does not necessarily represent the prior art in this field. Summary of the Invention
[0006] The present invention provides a control method and system for an online meeting system, an electronic device, and a storage medium, aiming to solve the problems of insufficient device status perception, low utilization rate of device energy consumption, lack of user behavior response, and insufficient energy use efficiency existing in the prior art.
[0007] According to one aspect of the present invention, there is provided a control method for an online meeting system based on multimodal perception, including: real-time collecting voice information and gesture information of users at the participating end; determining the user status of the participating end according to the voice information and gesture information, so as to dynamically adjust the parameter configuration information and the system resources occupied by the participating end according to the user status; determining the current meeting status of the online meeting system according to the number of currently active online meetings, so as to dynamically adjust the resource scheduling of the online meeting system.
[0008] According to some embodiments of the present invention, determining the user status of the participating end based on voice information and gesture information includes: when the voice information meets the first preset condition and the gesture information meets the second preset condition, determining that the user status is an active user status; when the voice information does not meet the first preset condition and / or the gesture information does not meet the second preset condition, determining that the user status is an inactive status; wherein, the first preset condition is that the voice activity of the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a gesture for indication.
[0009] According to some embodiments of the present invention, dynamically adjusting the resource scheduling of the online meeting system according to the current meeting status includes: when the current meeting status is a peak period meeting status, scheduling the system resources of the host end from the data center batch task pool to the audio-video codec calculation module; when the current meeting status is a non-peak period meeting status, releasing the idle resources of the host end to the data center batch task pool.
[0010] According to some embodiments of the present invention, the control method further includes: collecting the user expression information of the participating end in real time; determining the user emotion state according to the user expression information and the voice information; and sending a prompt message to the host end when the user emotion state is a fatigue state.
[0011] According to another aspect of the present invention, the present invention provides a control system for an online meeting system based on multi-modal perception, including an information collection module and a resource processing module. The information collection module collects the voice information and gesture information of the user at the participating end in real time; the resource processing module determines the user status of the participating end according to the voice information and the gesture information, so as to dynamically adjust the parameter configuration information and the occupied system resources of the participating end according to the user status, and determine the current meeting status of the online meeting system according to the number of currently active online meetings, so as to dynamically adjust the resource scheduling of the online meeting system according to the current meeting status.
[0012] According to some embodiments of the present invention, the resource processing module determines that the user status is an active user status when the voice information meets the first preset condition and the gesture information meets the second preset condition; the resource processing module determines that the user status is an inactive status when the voice information does not meet the first preset condition and / or the gesture information does not meet the second preset condition; wherein, the first preset condition is that the voice activity of the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a gesture for indication.
[0013] According to some embodiments of the present invention, when the current meeting state is a peak meeting state, the resource processing module schedules the system resources of the main speaking end from the data center batch task pool to the audio and video codec computing module; when the current meeting state is a non-peak meeting state, the resource processing module releases the idle resources of the main speaking end to the data center batch task pool.
[0014] According to some embodiments of the present invention, the information collection module real-time collects the user expression information of the participating end; the resource processing module determines the user emotion state according to the user expression information and the voice information; when the user emotion state is a fatigue state, the resource processing module sends a prompt message to the main speaking end.
[0015] According to another aspect of the present invention, the present invention also provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the control method as described above.
[0016] According to another aspect of the present invention, the present invention also provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the control method as described above.
[0017] According to another aspect of the present invention, the present invention also provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, when the program instructions are executed by a computer, enabling the computer to execute the control method as described above.
[0018] Beneficial effects The present invention can dynamically adjust device parameters according to the user's behavior information, so as to optimize the system resources according to the user's real-time interaction state. And the present invention can adopt an intelligent resource scheduling algorithm according to the current meeting state, transforming the resource management of the original meeting system from static allocation to dynamic scheduling, and can dynamically allocate computing resources (especially GPU resources) according to the real-time meeting load situation, so as to achieve the balance between high-quality performance and resource utilization. Brief description of the drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0020] Figure 1A flowchart showing a control method of an online conferencing system according to an embodiment of the present invention; Figure 2 Another flowchart showing a control method of an online conferencing system according to an embodiment of the present invention; Figure 3 A structural diagram showing a control system of an online conferencing system according to an embodiment of the present invention.
[0021] Explanation of reference numerals: Control system 1; Information acquisition module 10; Resource processing module 20. Detailed implementation manners
[0022] The following will clearly and completely describe the technical solutions of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The current online conferencing system at least still has the following problems: 1. Insufficient device status perception: The current online conferencing system mainly relies on basic audio and video codec technologies to achieve the transmission of audio and video data, but lacks deep integration with hardware devices (such as microphone arrays, cameras, etc.), resulting in a lack of real-time perception and dynamic adjustment of device status (such as microphone array power consumption, camera angle, etc.); 2. Low device energy consumption utilization rate: The current online conferencing system cannot dynamically optimize resource allocation according to the behaviors of participants (such as speaking frequency, gestures), but operates with fixed parameters all the time. Even when there is no one speaking, it still works with high power consumption, resulting in high energy consumption of terminal devices; 3. Lack of response to user behaviors: The current online conferencing system cannot make dynamic adjustments according to user behaviors, and there is a problem of lack of emotional interaction blank (such as being unable to adjust the meeting rhythm according to the user's emotional state), resulting in low meeting efficiency; 4. Insufficient energy use efficiency: For example, the current power environment monitoring system still has problems such as a single monitoring dimension (such as only collecting basic environmental data), using static threshold alarms and lacking dynamic adaptability, and lacking a prediction model.
[0024] According to one aspect of the present invention, the present invention provides a control method for an online conferencing system based on multi-modal perception.
[0025] Figure 1 A flowchart showing a control method of an online conferencing system according to an embodiment of the present invention. As Figure 1As shown, the control method may include steps S100 - S300. Exemplarily, the control method may be executed by a control system with computing capabilities (such as a server or a host, etc.).
[0026] An online meeting is a virtual meeting form that connects participants located in different geographical locations through Internet technology, enabling functions such as real - time audio - video communication, screen sharing, and file transfer.
[0027] According to an exemplary embodiment, an online meeting may include at least one host end and multiple participant ends. For example, the host end may be the terminal device where the host of the meeting is located; the participant ends may be the terminal devices where other participants of the meeting are located. In the present invention, the terminal device includes, but is not limited to, devices such as computers, mobile phones, and tablets, and the present invention does not limit this.
[0028] Exemplarily, the application scenario of the host end in the present invention may be a physical meeting scenario, and corresponding hardware devices (such as a microphone array, a camera device, and a computer room dynamic environment monitoring system, etc.) are also included at the host end. Below, taking the application scenario of the host end as a physical meeting scenario as an example, the present invention will be described and introduced.
[0029] According to an exemplary embodiment, in step S100, the control system collects the voice information and gesture information of the users at the participant ends in real - time.
[0030] For example, the participant end may collect the voice information of the user through a microphone array (such as a microphone) provided at the participant end; and the participant end may also collect the gesture information of the user (such as capturing the gesture trajectory of the user through a thermal imaging sensor, etc.) through a camera device (such as a camera) provided at the participant end.
[0031] The control system may communicate with multiple participant ends through a preset network interface to receive the voice information and gesture information sent by the multiple participant ends.
[0032] In step S200, the control system determines the user status of the participants at the participant ends based on the voice information and gesture information, so as to dynamically adjust the parameter configuration information and the system resources occupied by the participant ends. The user status includes an active user status and an inactive user status.
[0033] For example, after receiving the voice information and gesture information of multiple participant ends, the control system determines whether the current user status of the participant ends is an active user status or an inactive user status based on the voice information and gesture information.
[0034] Exemplarily, the active user state can be that the current user is speaking or actively interacting (such as raising a hand to ask a question); the inactive user state can be that the current user has not participated in the interaction for a long time (such as being silent, having no gesture, etc.).
[0035] According to the exemplary embodiment, after determining the user state of the current participating end, the control system adjusts the parameter configuration information and the occupied system resources of the participating end according to the user state.
[0036] For example, when the user state is the active user state, the control system can dynamically increase the parameter configuration information of the participating end, such as increasing the resolution of the camera device of the participating end to a first preset threshold (such as increasing from 720p to 1080p), turning on the noise reduction algorithm of the camera device, turning on the infrared fill light of the camera device, and increasing the system resources occupied by the participating end (such as GPU resources: Graphics Processing Unit, graphics processor), to ensure the picture clarity of the participating end in the active user state.
[0037] Again, for example, when the user state is the inactive user state, the control system can dynamically reduce the parameter configuration information of the participating end, such as reducing the resolution of the camera device of the participating end to a second preset threshold (such as reducing from 1080p to 720p), turning off the noise reduction algorithm of the camera device, turning off the infrared fill light of the camera device, and releasing the system resources (such as GPU resources) occupied by the participating end to the host end, to reduce the system resources occupied by the participating end in the inactive user state, and can improve the performance of the host end (such as increasing the screen sharing frame rate of the host end, etc.).
[0038] As an embodiment, in an online meeting, the online meeting system can include a host end and 50 participating ends. The control system determines that the user states of 3 participating ends are inactive user states according to the voice information and gesture information of the users (such as the voice activity levels of the 3 participating ends are all less than the preset threshold, and the gesture information is all non-signaling gestures), then the control system reduces the resolution of the camera devices of the 3 participating ends from 1080p to 720p, and turns off the infrared fill light and the noise reduction algorithm of their camera devices, etc., so that the device power consumption of the 3 participating ends can be reduced from 7.5w to 4w. The control system also releases the system resources of the 3 participating ends to the host end, so that the screen sharing frame rate of the host end can be increased to 60fps. Exemplarily, in this way, the overall system resource utilization rate of the online meeting can be reduced from 90% to about 65%, and the total power consumption of the online meeting can be reduced by 18%.
[0039] It can be understood here that the control system can monitor the user status of the participating end in real time. Even if the user status of the current participating end is an inactive user status, as long as it is detected that the user becomes active again, the control system can restore the participating end to the high-performance mode within 1 s, thereby ensuring the user experience effect.
[0040] Through the above embodiments, the present invention can dynamically adjust device parameters according to the user's behavior information, so as to optimize system resources according to the user's real-time interaction status.
[0041] Optionally, in step S200, when the voice information meets the first preset condition and the gesture information meets the second preset condition, the control system determines that the user status is an active user status. And when the voice information does not meet the first preset condition and / or the gesture information does not meet the second preset condition, the control system determines that the user status is an inactive status.
[0042] The first preset condition is that the voice activity in the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a signaling gesture.
[0043] For example, the voice information may include the sound pressure level (dB) and the voice duration. The control system can calculate the current voice activity of the participating end according to the average sound pressure level per minute and the voice duration ratio. The gesture information may include signaling gestures (such as raising a hand or waving) and non-signaling gestures (such as both hands being still). The control system can identify the gesture type according to the MediaPipe model (an open-source cross-platform framework for building real-time multimedia machine learning applications).
[0044] When the voice activity of the current participating end is greater than the preset threshold and the gesture information is a signaling gesture, the control system determines that the user status of the participating end is an active user status. When the voice activity of the current participating end is less than the preset threshold and / or the gesture information is a non-signaling gesture, the control system determines that the user status of the participating end is an inactive user status.
[0045] In step S300, the control system determines the current conference status of the online conference system according to the number of currently active online conferences. The current conference status includes a peak period status and a non-peak period status.
[0046] For example, the online conference system can load multiple online conferences at the same time. The control system can count the number of currently active online conferences. When the number of currently active online conferences is greater than a third preset threshold (such as 10 online conferences), the control system determines that the current conference status is a peak period conference status. Otherwise, the control system determines that the current conference status is a non-peak period status.
[0047] According to an exemplary embodiment, the control system dynamically adjusts the resource scheduling of the online meeting system according to the current meeting state. For example, the control system can schedule the system resources at the host end in real time according to the current meeting state to maximize the resource utilization rate.
[0048] Optionally, in step S300, when the current meeting state is a peak period meeting state, the control system schedules the system resources of the host end from the data center batch task pool to the audio-video codec calculation module. And when the current meeting state is a non-peak period meeting state, the control system releases the idle resources of the host end to the data center batch task pool.
[0049] For example, the system resources of the host end can be GPU resources. In the case of a peak period meeting state, the control system reclaims GPU resources from the data center batch task pool (such as terminating batch data analysis tasks), and centrally allocates the GPU computing power to calculation modules such as audio-video codec and audio noise reduction. In the case of a non-peak period meeting state, the control system can schedule the idle GPU resources to the data center batch task pool to complete calculations such as batch meeting video recording transcribing, meeting summary generation, and operation and maintenance log analysis.
[0050] Through the above embodiments, the present invention can adopt an intelligent resource scheduling algorithm according to the current meeting state, change the resource management of the original meeting system from static allocation to dynamic scheduling, and can dynamically allocate computing resources (especially GPU resources) according to the real-time meeting load situation, so as to achieve a balance between high-quality performance and resource utilization.
[0051] Optionally, the control system can also predict the meeting load of the online meeting system in a preset future time period (such as the next 1 hour) based on a prediction model (such as an ARIMA model: AutoRegressive Integrated Moving Average), so that the control system can dynamically adjust the resource scheduling strategy of the online meeting system according to the predicted meeting load.
[0052] For example, the control system can collect historical meeting data (such as the number of meetings, the number of participants, the audio-video bit rate, the resource occupancy rate, etc.) and external variable data (such as time period and holiday marks, etc.) in real time, and train a prediction model (such as an ARIMA model or an LSTM model) according to the historical meeting data.
[0053] The ARIMA model can capture the periodicity (such as daily peak hours), trend (such as the difference between weekdays and weekends), and random fluctuations of the meeting load based on time series analysis. It can also eliminate non-stationarity through differencing and combine autoregressive and moving average terms to predict the meeting load in a preset future time period. The LSTM (Long Short-Term Memory) model can process multi-dimensional time series data (such as CPU / GPU utilization and network bandwidth, etc.) using long short-term memory networks, and can predict load peaks and resource requirements by learning the non-linear relationships of the load.
[0054] Exemplarily, the control system can schedule resources in advance (such as reserving GPU computing power) or trigger an expansion mechanism (such as starting a standby service, etc.) according to the predicted meeting load to avoid insufficient resources during peak periods.
[0055] Optionally, the control system can also monitor the real-time utilization rate of system resources (such as GPU resources) in real time, and trigger an expansion mechanism when the real-time utilization rate is greater than the fourth preset threshold.
[0056] Exemplarily, the expansion mechanism includes but is not limited to starting a standby server or invoking an idle server, etc., and the present invention does not limit this.
[0057] Optionally, the control system can also use diversified real-time parameter data such as device temperature, vibration frequency, current load, network traffic, etc. as inputs, input them into a prediction model (such as the LSTM model) to predict the health status of hardware devices (such as microphones, camera devices, etc.) at the speaker end, and issue a warning through a dynamic threshold method.
[0058] For example, the control system can collect parameter information such as device temperature, vibration frequency, current load, network traffic, etc. in real time, then extract the key features of the parameter information, and train a prediction model (such as the LSTM model) based on the key features of the parameter information.
[0059] By outputting the key features of the parameter information (i.e., multi-dimensional time series data), the LSTM model can learn the device operation status mode, and thus can predict future failure risks (such as the probability of fan failure, etc.).
[0060] The control system can dynamically adjust the alarm threshold according to the prediction result of the LSTM model to avoid false alarms or missed alarms caused by static thresholds. Exemplarily, the temperature alarm threshold can be adaptively adjusted as the load increases.
[0061] The control system can also generate device health assessment information (such as 0 - 100 points) based on the prediction results of various parameter information. When the device health assessment information is lower than the preset threshold, the control system triggers an alarm, such as marking potential faulty components and suggesting maintenance measures (such as cleaning the fan, replacing the power supply), to remind the user to perform device maintenance or repair, etc.
[0062] As an embodiment, the acquisition methods of parameter information can include: 1. Real - time monitoring of the device temperature of key components (such as GPU, power module) through the built - in thermal sensor or infrared thermal imaging module of the device.
[0063] 2. By deploying a micro - accelerometer on the device housing, abnormal vibrations (such as vibrations generated by fan imbalance or hard disk jitter) can be captured.
[0064] 3. Recording current fluctuations through the power management unit to identify current load or short - circuit risks.
[0065] 4. Statistically analyzing real - time traffic through the network interface card to analyze the correlation between burst traffic and device performance.
[0066] As an embodiment, the data processing process of parameter information can include: 1. Data normalization processing: Standardize parameters with different dimensions (such as converting temperature to a percentage value for representation).
[0067] 2. Time - series alignment processing: Align multi - source data through timestamps to ensure data synchronization.
[0068] 3. Abnormal filtering processing: Adopt a sliding window algorithm to eliminate instantaneous noise (such as filtering short - term current spikes, etc.).
[0069] Optionally, the control system can also perform dynamic refrigeration.
[0070] For example, the control system can trigger the start of the liquid - cooling system when the temperature of key components exceeds the dynamic threshold (dynamically adjusted according to the load).
[0071] Exemplarily, the control system can adjust the coolant flow rate and pump power according to the temperature range (such as low flow rate in the low - temperature section and full - speed operation in the high - temperature section). The control system can also directionally cool high - heat - generating areas (such as GPU) through a micro - channel liquid - cooling plate to reduce overall energy consumption. And the control system can also dynamically adjust the cooling intensity in combination with the ambient temperature (room temperature and humidity sensor) to avoid over - refrigeration.
[0072] The control system can also set a flow - monitoring module in the liquid - cooling system. When there are problems such as blockage or leakage in the pipeline, it triggers the emergency start of the standby air - cooling module.
[0073] Optionally, the control system can also collect the device parameter information of the hardware devices at the host end in real time, such as microphone arrays, imaging devices, screens, power modules, etc.
[0074] For example, the parameter information of the microphone array can include power consumption information, pickup sensitivity, array angle, etc. The parameter information of the imaging device can include resolution configuration, focal length parameter, fill light status, and lens cleanliness, etc. The parameter information of the screen can include brightness value, refresh rate, and panel temperature, etc. The parameter information of the power module can include input voltage, output current, and efficiency percentage, etc.
[0075] The collection methods of the device parameter information can include: 1. Hardware interface collection: Read sensor data through I2C, SPI buses, etc.
[0076] 2. Driver layer interface collection: Call the device driver API to obtain the device parameter information.
[0077] 3. Image analysis collection: Analyze the lens stains or physical damages through the self-taken pictures of the imaging device.
[0078] The control system can receive the device parameter information in real time through the MQTT protocol (encapsulated in JSON format). The control system can monitor the device status of the hardware devices at the host end in real time through the device parameter information of the hardware devices at the host end.
[0079] Optionally, Figure 2 Another flowchart showing the control method of the online conference system according to an embodiment of the present invention is shown.
[0080] As Figure 2 shown, the control method can include steps S400 - S600.
[0081] In step S400, the control system collects the user expression information at the participant end in real time.
[0082] For example, the participant end can capture the facial micro-expressions of the user (such as the upward curvature of the mouth corner, the frequency of frowning, etc., which are expressions used to express the user's emotions) based on the imaging device. The control system can obtain the user expression information at the participant end in real time.
[0083] In step S500, the control system determines the user emotional state according to the user expression information and the voice information.
[0084] For example, the control system can extract information such as the standard deviation of the fundamental frequency, speech rate, and energy distribution in the voice information, and can judge the emotional intensity of the user, etc. Then, the control system can determine the user emotional state of the current user according to the determined user expression information combined with the emotional intensity of the user.
[0085] Exemplarily, the user emotional state may include a fatigue state and a non-fatigue state.
[0086] In step S600, when the user emotional state of the control system is in a fatigue state, a prompt message is sent to the host end.
[0087] For example, the control system may send the user emotional state of the participating end to the host end by means of a prompt message, so that the host of the host end can control the meeting process according to the prompt message.
[0088] Exemplarily, when the number of participating ends with a fatigue user emotional state exceeds a certain number, the host of the host end can take an appropriate mid-meeting break, so as to ensure the user experience and the meeting effect.
[0089] Through the above embodiments, the present invention provides a control method for an online meeting system based on multi-modal perception, which can realize device management and resource scheduling of an online meeting through multi-angle and multi-faceted perceptions such as user behavior perception, meeting state perception, device state perception, and user emotion perception.
[0090] According to another aspect of the present invention, the present invention also provides a control system for an online meeting system based on multi-modal perception. Figure 3 The structural schematic diagram of the control system of the online meeting system according to the embodiment of the present invention is shown. As Figure 3 shown, the control system 1 includes an information collection module 10 and a resource processing module 20.
[0091] The information collection module 10 collects the voice information and gesture information of the users at the participating ends in real time.
[0092] For example, the participating end can collect the voice information of the user through a microphone array (such as a microphone) arranged at the participating end; and the participating end can also collect the gesture information of the user through a camera device (such as a camera) arranged at the participating end (such as capturing the gesture trajectory of the user through a thermal imaging sensor, etc.).
[0093] The information collection module 10 can be communicatively connected to multiple participating ends through a preset network interface to receive the voice information and gesture information sent by the multiple participating ends.
[0094] The resource processing module 20 determines the user state of the participating end according to the voice information and gesture information, so as to dynamically adjust the parameter configuration information and the system resources occupied by the participating end. The user state includes an active user state and an inactive user state.
[0095] For example, after receiving voice information and gesture information from multiple conference participants, the resource processing module 20 determines whether the user status of the current conference participant is an active user status or an inactive user status according to the voice information and gesture information.
[0096] Exemplarily, the active user status may be that the current user is speaking or actively interacting (such as raising hand to ask a question); the inactive user status may be that the current user has not participated in the interaction for a long time (such as silence, no gesture, etc.).
[0097] According to an exemplary embodiment, after determining the user status of the current participant terminal, the resource processing module 20 adjusts the parameter configuration information and the occupied system resources of the participant terminal according to the user status.
[0098] For example, when the user status is an active user status, the resource processing module 20 can dynamically increase the parameter configuration information of the participating end, such as increasing the resolution of the camera device of the participating end to a first preset threshold (such as from 720p to 1080p), turning on the camera device noise reduction algorithm, turning on the camera device infrared fill light, and increasing the system resources occupied by the participating end (such as GPU resources: Graphics Processing Unit, graphics processor) to ensure the picture clarity of the participating end in the active user status.
[0099] For another example, when the user status is an inactive user status, the resource processing module 20 can dynamically reduce the parameter configuration information of the participating end, such as reducing the resolution of the camera device of the participating end to a second preset threshold (such as from 1080p to 720p), turning off the camera device noise reduction algorithm, turning off the camera device infrared fill light, and releasing the system resources occupied by the participating end (such as GPU resources) to the main speaker end, so as to reduce the system resources occupied by the participating end in the inactive user status, and improve the performance of the main speaker end (such as increasing the screen sharing frame rate of the main speaker end).
[0100] It can be understood here that the resource processing module 20 can monitor the user status of the participating end in real time. Even if the current user status of the participating end is an inactive user status, as long as the resource processing module 20 detects that the user is active again, the participating end can be restored to high-performance mode within 1 second, thereby ensuring the user experience.
[0101] Through the above embodiments, the present invention can dynamically adjust device parameters according to user behavior information, thereby optimizing system resources according to the real-time interactive status of the user.
[0102] Optionally, when the voice information meets the first preset condition and the gesture information meets the second preset condition, the resource processing module 20 determines that the user status is the active user status. And when the voice information does not meet the first preset condition and / or the gesture information does not meet the second preset condition, the resource processing module 20 determines that the user status is the inactive status.
[0103] The first preset condition is that the voice activity in the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a signaling gesture.
[0104] For example, the voice information may include the sound pressure level (dB) and the voice duration. The resource processing module 20 may calculate the voice activity of the current participating end according to the average sound pressure level per minute and the proportion of the voice duration. The gesture information may include signaling gestures (such as raising a hand or waving) and non-signaling gestures (such as both hands being stationary). The resource processing module 20 may identify the gesture type according to the MediaPipe model (an open-source cross-platform framework for building real-time multimedia machine learning applications).
[0105] When the voice activity of the current participating end is greater than the preset threshold and the gesture information is a signaling gesture, the resource processing module 20 determines that the user status of this participating end is the active user status. When the voice activity of the current participating end is less than the preset threshold and / or the gesture information is a non-signaling gesture, the resource processing module 20 determines that the user status of this participating end is the inactive user status.
[0106] According to the exemplary embodiment, the resource processing module 20 determines the current meeting status of the online meeting system according to the number of currently active online meetings. The current meeting status includes the peak period status and the non-peak period status.
[0107] For example, the online meeting system can load multiple online meetings simultaneously. The resource processing module 20 can count the number of currently active online meetings. When the number of currently active online meetings is greater than a third preset threshold (such as 10 online meetings), the resource processing module 20 determines that the current meeting status is the peak period meeting status. Otherwise, the resource processing module 20 determines that the current meeting status is the non-peak period status.
[0108] According to the exemplary embodiment, the resource processing module 20 dynamically adjusts the resource scheduling of the online meeting system according to the current meeting status. For example, the resource processing module 20 can schedule the system resources at the host end in real time according to the current meeting status to maximize the resource utilization rate.
[0109] Optionally, when the current meeting state is the peak meeting state, the resource processing module 20 schedules the system resources of the host end from the data center batch task pool to the audio and video codec calculation module. And when the current meeting state is the non-peak meeting state, the resource processing module 20 releases the idle resources of the host end to the data center batch task pool.
[0110] For example, the system resources of the host end can be GPU resources. In the case of the peak meeting state, the resource processing module 20 reclaims GPU resources from the data center batch task pool (such as terminating batch data analysis tasks), and centrally allocates the GPU computing power to calculation modules such as audio and video codec and audio noise reduction. In the case of the non-peak meeting state, the resource processing module 20 can schedule the idle GPU resources to the data center batch task pool to complete calculations such as batch meeting video recording transcription, meeting summary generation, and operation and maintenance log analysis.
[0111] Through the above embodiments, the present invention can adopt an intelligent resource scheduling algorithm according to the current meeting state, change the resource management of the original meeting system from static allocation to dynamic scheduling, and can dynamically allocate computing resources (especially GPU resources) according to the real-time meeting load situation, so as to achieve a balance between high-quality performance and resource utilization.
[0112] Optionally, the resource processing module 20 can also predict the meeting load of the online meeting system in a preset future time period (such as the next 1 hour) based on a prediction model (such as the ARIMA model: AutoRegressive Integrated Moving Average), so that the resource processing module 20 can dynamically adjust the resource scheduling strategy of the online meeting system according to the predicted meeting load.
[0113] For example, the resource processing module 20 can collect historical meeting data (such as the number of meetings, the number of participants, audio and video bitrates, resource occupancy rates, etc.) and external variable data (such as time period and holiday marks, etc.) in real time, and train a prediction model (such as the ARIMA model or the LSTM model) based on the historical meeting data.
[0114] The ARIMA model can capture the periodicity (such as daily peak hours), trend (such as the difference between weekdays and weekends), and random fluctuations of the meeting load based on time series analysis. And it can eliminate non-stationarity through differencing processing, and combine autoregressive and moving average terms to predict the meeting load in a preset future time period. The LSTM model can use long short-term memory networks to process multi-dimensional time series data (such as CPU / GPU utilization and network bandwidth, etc.), and can predict load peaks and resource requirements by learning the non-linear relationship of the load.
[0115] Exemplarily, the resource processing module 20 may schedule resources in advance according to the predicted conference load (such as reserving GPU computing power) or trigger the expansion of machines (such as starting standby services, etc.) to avoid insufficient resources during peak periods.
[0116] Optionally, the resource processing module 20 may also monitor the real-time utilization rate of system resources (such as GPU resources) in real time, and trigger an expansion mechanism when the real-time utilization rate is greater than the fourth preset threshold.
[0117] Exemplarily, the expansion mechanism includes but is not limited to starting a standby server or invoking an idle server, etc., and the present invention does not limit this.
[0118] Optionally, the resource processing module 20 may also use diversified real-time parameter data such as device temperature, vibration frequency, current load, network traffic, etc. as inputs, input them into a prediction model (such as an LSTM model) to predict the health status of hardware devices (such as microphones, camera devices, etc.) at the speaker end, and issue a warning through a dynamic threshold method.
[0119] For example, the resource processing module 20 may collect parameter information such as device temperature, vibration frequency, current load, network traffic, etc. in real time, then extract the key features of the parameter information, and train a prediction model (such as an LSTM model) based on the key features of the parameter information.
[0120] By outputting the key features of the parameter information (i.e., multi-dimensional time series data), the LSTM model can learn the device operation state mode, and thus can predict future failure risks (such as the probability of fan failure, etc.).
[0121] The resource processing module 20 may dynamically adjust the alarm threshold according to the prediction result of the LSTM model to avoid false alarms or missed alarms caused by static thresholds. Exemplarily, the temperature alarm threshold may be adaptively adjusted as the load increases.
[0122] The resource processing module 20 may also generate device health assessment information (such as 0 - 100 points) according to the prediction results of various parameter information. When the device health assessment information is lower than the preset threshold, the control system triggers an alarm, such as marking potential faulty components and suggesting maintenance measures (such as cleaning the fan, replacing the power supply), to remind the user to perform device maintenance or repair, etc.
[0123] As an embodiment, the acquisition method of parameter information may include: 1. Real-time monitoring of the device temperature of key components (such as GPU, power module) through a built-in thermal sensor or an infrared thermal imaging module of the device.
[0124] 2. By deploying a micro accelerometer on the device shell, abnormal vibrations (such as vibrations generated by fan imbalance or hard disk jitter) can be captured.
[0125] 3. Record current fluctuations through the power management unit to identify current load or short - circuit risks.
[0126] 4. Count real - time traffic through the network interface card to analyze the correlation between burst traffic and device performance.
[0127] As an embodiment, the data processing process of parameter information may include: 1. Data normalization processing: Standardize parameters with different dimensions (such as converting temperature to a percentage value).
[0128] 2. Time - series alignment processing: Align multi - source data through timestamps to ensure data synchronization.
[0129] 3. Abnormal filtering processing: Use a sliding - window algorithm to eliminate instantaneous noise (such as filtering short - term current spikes, etc.).
[0130] Optionally, the resource processing module 20 can also perform dynamic refrigeration.
[0131] For example, the resource processing module 20 can trigger the start of the liquid - cooling system when the temperature of key components exceeds a dynamic threshold (dynamically adjusted according to the load).
[0132] Exemplarily, the resource processing module 20 can adjust the coolant flow rate and pump power according to the temperature range (such as low flow rate in the low - temperature section and full - speed operation in the high - temperature section). The resource processing module 20 can also directionally cool high - heat - generating areas (such as the GPU) through a micro - channel liquid - cooling plate to reduce overall energy consumption. And the resource processing module 20 can also dynamically adjust the cooling intensity in combination with the ambient temperature (room temperature and humidity sensor) to avoid over - refrigeration.
[0133] The resource processing module 20 can also set a flow - monitoring module in the liquid - cooling system to trigger the emergency start of the standby air - cooling module in case of blockage or leakage problems in the pipeline.
[0134] Optionally, the resource processing module 20 can also collect device parameter information of the main - speaking end's hardware devices (such as microphone array, imaging device, screen, and power module, etc.).
[0135] For example, the parameter information of the microphone array can include power consumption information, pickup sensitivity, array angle, etc. The parameter information of the imaging device can include resolution configuration, focal length parameter, fill - light status, and lens cleanliness, etc. The parameter information of the screen can include brightness value, refresh rate, and panel temperature, etc. The parameter information of the power module can include input voltage, output current, and efficiency percentage, etc.
[0136] The acquisition methods of device parameter information can include: 1. Hardware interface acquisition: Read sensor data through I2C, SPI buses, etc.
[0137] 2. Driver layer interface acquisition: Call device driver APIs to obtain device parameter information.
[0138] 3. Image analysis acquisition: Analyze the lens stains or physical damages through the self-taken pictures of the imaging device.
[0139] The resource processing module 20 can receive the device parameter information of this device in real time through the MQTT protocol (encapsulated in JSON format). The resource processing module 20 can monitor the device status of the hardware device at the main speaking end in real time through the device parameter information of the hardware device at the main speaking end.
[0140] Optionally, the information acquisition module 10 acquires the user expression information of the participating end in real time.
[0141] For example, the participating end can capture the facial micro-expressions of the user based on the imaging device (such as the expressions used to express the user's emotions like the upward curl of the mouth, the frequency of frowning, etc.). The information acquisition module 10 can obtain the user expression information of this participating end in real time.
[0142] The resource processing module 20 determines the user emotion state according to the user expression information and the voice information.
[0143] For example, the resource processing module 20 can extract information such as the standard deviation of the fundamental frequency, the speech rate, and the energy distribution in the voice information, and can judge the emotion intensity of the user, etc. Then, the resource processing module 20 combines the determined user expression information with the emotion intensity of the user to determine the user emotion state of the current user.
[0144] Exemplarily, the user emotion state can include a fatigue state and a non-fatigue state.
[0145] When the user emotion state is the fatigue state, the resource processing module 20 sends a prompt message to the main speaking end.
[0146] For example, the resource processing module 20 can send the user emotion state of the participating end to the main speaking end in the form of a prompt message, so that the main speaker at the main speaking end can control the meeting process according to this prompt message.
[0147] Exemplarily, when the number of participating ends with the user emotion state being the fatigue state exceeds a certain amount, the main speaker at the main speaking end can take an appropriate mid-meeting break, thereby ensuring the user experience and the meeting development effect.
[0148] Through the above embodiments, the present invention provides a control method for an online meeting system based on multi-modal perception, which can achieve device management and resource scheduling for online meetings through multi-faceted perceptions such as user behavior perception, meeting state perception, device state perception, and user emotion perception.
[0149] According to another aspect of the present invention, the present invention also provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the control method as described above.
[0150] According to another aspect of the present invention, the present invention also provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the control method as described above.
[0151] According to another aspect of the present invention, the present invention also provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the control method as described above.
[0152] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions of the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A control method for an online conference system, characterized in that: The online conference system includes a speaker terminal and a participant terminal, and the control method includes: Collecting voice information and gesture information of the user at the conference participant in real time; Determining the user status of the conference participant terminal according to the voice information and the gesture information, so as to dynamically adjust the parameter configuration information and the occupied system resources of the conference participant terminal according to the user status; The current conference state of the online conference system is determined according to the number of currently active online conferences, so as to dynamically adjust resource scheduling of the online conference system according to the current conference state.
2. The control method according to claim 1, characterized in that: The determining the user status of the conference participant terminal according to the voice information and the gesture information includes: When the voice information satisfies a first preset condition and the gesture information satisfies a second preset condition, determining that the user status is an active user status; When the voice information does not satisfy the first preset condition and / or the gesture information does not satisfy the second preset condition, determining that the user status is an inactive status; Among them, the first preset condition is that the voice activity of the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a gestural gesture.
3. The control method according to claim 1, characterized in that: The dynamically adjusting resource scheduling of the online conference system according to the current conference status includes: When the current conference state is a peak conference state, the system resources of the main speaker terminal are dispatched from the data center batch task pool to the audio and video codec calculation module; When the current conference state is a non-peak conference state, the idle resources of the main speaker terminal are released to the data center batch task pool.
4. The control method according to claim 1, characterized in that: The control method further comprises: Collecting the user expression information of the participant terminal in real time; Determining the user's emotional state based on the user's expression information and the voice information; When the user's emotional state is a fatigue state, a prompt message is sent to the main speaker terminal.
5. A control system for an online conference system, characterized in that: The online conference system includes a speaker terminal and a participant terminal, and the control system includes: An information collection module collects voice information and gesture information of the user at the conference terminal in real time; A resource processing module determines the user status of the participant end according to the voice information and the gesture information, so as to dynamically adjust the parameter configuration information and the occupied system resources of the participant end according to the user status, and determines the current conference status of the online conference system according to the number of currently active online conferences, so as to dynamically adjust the resource scheduling of the online conference system according to the current conference status.
6. The control system according to claim 5, characterized in that: The resource processing module determines that the user status is an active user status when the voice information meets a first preset condition and the gesture information meets a second preset condition; The resource processing module determines that the user status is an inactive status when the voice information does not meet the first preset condition and / or the gesture information does not meet the second preset condition; Among them, the first preset condition is that the voice activity of the voice information is greater than or equal to a preset threshold, and the second preset condition is that the gesture information is a gestural gesture.
7. The control system according to claim 5, characterized in that: The resource processing module dispatches the system resources of the main speaker terminal from the data center batch task pool to the audio and video codec calculation module when the current conference state is a peak conference state; When the current conference state is a non-peak conference state, the resource processing module releases the idle resources of the main speaker terminal to the data center batch task pool.
8. The control system according to claim 5, characterized in that: The information collection module collects the user expression information of the participating terminal in real time; The resource processing module determines the user's emotional state based on the user's expression information and the voice information; When the user's emotional state is fatigue, the resource processing module sends a prompt message to the main speaker.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the control method as described in any one of claims 1-4.
10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the control method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Multiparty communication control system, multiparty communication system and multiparty communication processing method
CN102523422A
Conference audio control method, system and device and computer readable storage medium
CN110300001A
Video conference data transmission method and video conference system
CN115941881A
Resource scheduling method and device, storage medium and product
CN116506570A
Video conference resource scheduling method and device, equipment and storage medium
CN117478822A