Video conference remote supervision system and method based on conference environment information

By constructing multi-dimensional environmental data and attention assessment, and automatically executing dynamic resource allocation strategies, the problem of interruption of the core participants' experience during video conferencing systems when the network fluctuates has been solved, thus optimizing resource utilization and meeting quality.

CN121567833APending Publication Date: 2026-02-24INFORMATION & COMMUNICATION BRANCH STATE GRID JIBEI ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511508925.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing video conferencing systems cannot adaptively adjust to fluctuations in network bandwidth, leading to interruptions in the experience of key participants, affecting the meeting process, and neglecting the impact of the physical meeting environment and the behavior of participants on meeting quality.

Method used

By collecting data on network performance, audio quality, video quality, and physical environment of participants, multi-dimensional environmental data is constructed. A weighted fusion algorithm is used to analyze the real-time meeting quality index, and dynamic resource allocation strategies are automatically executed based on the focus status to prioritize the experience of core participants.

Benefits of technology

When network congestion occurs, priority should be given to ensuring the audio stream of core participants, while reducing the resolution of the video stream of non-core participants to ensure the basic progress of the meeting and improve resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567833A_ABST
    Figure CN121567833A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video conference remote supervision, in particular to a video conference remote supervision system and method based on meeting place environment information, and the method comprises the steps: collecting the network performance data of a corresponding server of a video conference in real time, and when the transmission delay in the network performance data of the corresponding server is monitored to be greater than or equal to a first preset threshold value, sending the video conference to a server; based on the real-time concentration state of the conference participants, a dynamic resource allocation strategy is automatically executed, and video stream resource occupation of non-core conference participants is reduced. According to the method, conference participant concentration degree evaluation based on visual analysis is introduced to serve as a key execution basis of a dynamic resource allocation strategy, so that when a server network is congested, a video conference can preferentially guarantee experience of key persons and concentrators, resources of non-core conference participants are limited to a certain extent, and the experience of the video conference is guaranteed. Therefore, the core conference experience is maximized under limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video conferencing remote monitoring technology, specifically to a video conferencing remote monitoring system and method based on meeting room environment information. Background Technology

[0002] With the widespread adoption of remote work and global collaboration, video conferencing has become a core tool for daily business operations. However, existing video conferencing systems and their monitoring methods have significant limitations. Current remote monitoring systems primarily focus on technical indicators at the network transmission layer, such as bandwidth, packet loss rate, latency, and basic audio / video connectivity. However, they severely neglect the actual impact of the physical meeting environment and participant behavior on meeting quality and efficiency. Furthermore, when network bandwidth fluctuates drastically, existing remote monitoring systems often resort to indiscriminate global quality degradation (e.g., reducing video resolution for all participants). This approach cannot adaptively adjust to the actual dynamics of the meeting, leading to interruptions for key participants (such as the speaker) and severely impacting the meeting's progress. Summary of the Invention

[0003] The purpose of this invention is to provide a video conferencing remote monitoring system and method based on meeting room environment information, so as to solve the problems mentioned in the background art.

[0004] To address the aforementioned technical problems, this invention provides the following technical solution: a video conferencing remote monitoring method based on meeting room environment information, comprising: Each video conference corresponds to multiple endpoints and one server. Each endpoint is bound to each participant in the corresponding video conference, and different endpoints exchange information through the server. Network performance data, audio quality data, video data, and physical environment data of each participant in the video conference are collected separately to construct multi-dimensional environmental data of the corresponding participant's venue, and then transmitted to the server through the corresponding edge device of the participant. Based on the multi-dimensional environmental data of the participants in the meeting room, the server uses a weighted fusion algorithm to analyze and obtain the real-time meeting quality index of each participant in the corresponding video conference; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The system collects network performance data of the corresponding servers in real time. When the transmission delay in the network performance data of the corresponding server is greater than or equal to the first preset threshold, it automatically executes a dynamic resource allocation strategy based on the real-time focus status of the participants to reduce the video stream resource consumption of non-core participants.

[0005] Furthermore, the network performance data includes one or more of network jitter, latency, and bandwidth; the audio quality data includes the signal-to-noise ratio of the audio signal received by the corresponding edge from the server; the video image data includes the smoothness of the video image received by the corresponding edge from the server; and the physical environment data includes one or more of air quality, temperature, and humidity, as well as the noise level of the venue where the participants are located.

[0006] This invention unifies the collection of information that was originally scattered across network devices, audio and video devices, and IoT sensors, laying a data foundation for comprehensive intelligent analysis of the real-time meeting quality index corresponding to the participants.

[0007] Furthermore, the analysis yields a real-time meeting quality index for each participant in the corresponding video conference, calculated using the following formula: Among them, MQI i This indicates the real-time meeting quality index of the corresponding video conference based on the i-th participant; Sn i Sa represents the network dimension quality sub-score corresponding to the most recently acquired network performance data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; i This represents the audio dimension quality sub-score corresponding to the most recently acquired audio quality data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Sv i This represents the video dimension quality sub-score corresponding to the most recently acquired video frame data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Se i This represents the environmental dimension quality sub-score corresponding to the most recently acquired physical environment data of the i-th participant in the corresponding video conference within the database's preset form; Wn, Wa, Wv, and We represent the preset weight coefficients for the network dimension, audio dimension, video dimension, and environmental dimension, respectively.

[0008] This invention normalizes and comprehensively analyzes complex, multi-dimensional environmental data of meeting venues to a certain extent, greatly reducing the cognitive burden on administrators.

[0009] Furthermore, the specific implementation steps for obtaining the real-time attention status of participants are as follows: Step 1: Capture the facial images of the attendees using a camera, and extract the temporal features of the head pose angle, eye gaze direction vector, and mouth aspect ratio from the facial images of the attendees; the head pose angle represents the angle between the facial plane and the camera plane in the facial key point detection results obtained through face recognition. Step 2: Based on the obtained time-series characteristics, calculate the real-time attention level assessment value of the participants. The calculation formula is as follows: Among them, AS i FF represents the real-time attention level assessment value of the i-th participant; i FG represents the cosine of the maximum head posture angle of the i-th participant within the most recent preset unit time; i FY represents the ratio of the duration during which the eye gaze direction vector of the i-th participant points to the corresponding edge of the screen area in the video conference within the most recent preset unit of time; i denoted as the number of yawning actions identified in the temporal features of the mouth aspect ratio of the i-th participant within the most recent preset unit time; e represents the natural constant; λ1, λ2, and λ3 represent the weight coefficients of the first feature, the second feature, and the third feature, respectively. In the process of obtaining the number of yawning actions identified from the temporal features of the mouth's aspect ratio, each yawning action corresponds to a continuous temporal segment in the temporal features of the mouth's aspect ratio. The yawning action is identified by comparing the temporal change relationship of the mouth's aspect ratio within the corresponding continuous temporal segment with the temporal change relationship of the mouth's aspect ratio corresponding to the yawning action in the database's preset form.

[0010] This invention integrates multiple features such as head posture, eye gaze, and yawning frequency to avoid misjudgment based on a single feature (such as relying solely on face detection), making the results of attention state assessment more reliable.

[0011] Furthermore, the automatic execution of the dynamic resource allocation strategy includes: Based on the real-time focus status and real-time meeting quality index of the participants, calculate the resource allocation priority evaluation value for each participant in the corresponding video conference. Based on the priority evaluation value of the corresponding resource allocation, the participants in the corresponding video conference, excluding the current speaker, are sorted in descending order to construct a candidate sequence of participants for optimization in the corresponding video conference. The difference between the transmission delay and the first preset threshold in the network performance data of the corresponding server at the current time is used to determine the number of attendees to be optimized in the database, which is recorded as the reference value for the number of optimized objects. The set of attendees in the candidate sequence of attendees to be optimized, counting the reference value of the number of optimized objects from the end to the beginning, is recorded as the set of attendees to be optimized. All attendees in the set of attendees to be optimized are considered as non-core attendees at the current time. The dynamic resource allocation strategy for non-core attendees is obtained. The dynamic resource allocation strategy includes reducing the resolution of the video stream subsequently received by the attendees, and the reduced resolution of the video stream received by the attendees. The reduced resolution of the video stream received by the attendees is obtained by querying the video stream resolution bound in the database based on the resource allocation priority evaluation value of the attendees at the current time. When the transmission latency in the network performance data of the corresponding server is less than the first preset threshold, there is no need to automatically execute the resource allocation strategy.

[0012] In this invention, when the server network is congested, the audio stream of the speaker is prioritized, ensuring the effective transmission of core meeting information and effectively preventing meeting interruptions caused by video buffering, thus guaranteeing the basic progress of the meeting. At the same time, by introducing a resource allocation priority evaluation value, limited bandwidth resources are allocated to those who need them most (more focused participants with better meeting quality), while resources are restricted for participants who are distracted or have poor meeting quality. This approach optimizes the return on investment of resources to a certain extent, minimizing the overall negative impact on the user experience while ensuring the core experience.

[0013] Furthermore, the resource allocation priority evaluation value for the i-th participant in the corresponding video conference is the product of the i-th participant's real-time focus status and the i-th participant's real-time meeting quality index.

[0014] A video conferencing remote monitoring system based on meeting venue environment information includes: The multi-dimensional data acquisition module for meeting room environment information is used to collect network performance data, audio quality data, video image data and physical environment data of each participant in the video conference, construct the multi-dimensional environment data of the meeting room for the corresponding participant, and transmit it to the server through the corresponding edge device of the participant. The real-time meeting quality analysis module controls the server to analyze the multi-dimensional environmental data of the meeting room based on the participants, and obtain the real-time meeting quality index of each participant in the corresponding video conference through a weighted fusion algorithm; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The resource allocation dynamic management module is used to collect network performance data of the corresponding servers in real time. When the transmission delay in the network performance data of the corresponding server is greater than or equal to the first preset threshold, the dynamic resource allocation strategy is automatically executed based on the real-time focus status of the participants to reduce the video stream resource consumption of non-core participants.

[0015] Furthermore, the resource allocation dynamic management module includes a server network performance comparison unit, a focus status analysis unit, and a dynamic resource allocation strategy generation unit; The server network performance comparison unit is used to collect network performance data of the corresponding server in real time and compare the transmission delay in the network performance data of the corresponding server with a first preset threshold. The focus state analysis unit is used to obtain the real-time focus state of each participant. The dynamic resource allocation strategy generation unit automatically executes a dynamic resource allocation strategy based on the real-time focus status of each participant obtained by the focus status analysis unit, thereby reducing the video stream resource consumption of non-core participants.

[0016] Compared with the prior art, the beneficial effects achieved by the present invention are: (1) This invention integrates network, audio, video and physical environment data to construct a unified meeting quality evaluation system covering the three dimensions of technology, environment and participants, enabling administrators to accurately and effectively grasp the overall health status of the meeting; (2) This invention introduces visual analysis-based participant attention assessment as the key basis for dynamic resource allocation strategy, thereby enabling video conferencing to prioritize the experience of key personnel and focused participants when the server network is congested, while the resources of non-core participants are limited to a certain extent, thus maximizing the core meeting experience with limited resources. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of the video conferencing remote monitoring system based on meeting room environment information according to the present invention; Figure 2 This is a flowchart illustrating the video conferencing remote monitoring method based on meeting room environment information according to the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please see Figures 1-2 The present invention provides a technical solution: such as Figure 1 As shown, this embodiment provides a video conferencing remote monitoring system based on meeting room environment information, including: The multi-dimensional data acquisition module for meeting room environment information is used to collect network performance data, audio quality data, video image data and physical environment data of each participant in the video conference, construct the multi-dimensional environment data of the meeting room for the corresponding participant, and transmit it to the server through the corresponding edge device of the participant. The real-time meeting quality analysis module controls the server to analyze the multi-dimensional environmental data of the meeting room based on the participants, and obtain the real-time meeting quality index of each participant in the corresponding video conference through a weighted fusion algorithm; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The resource allocation dynamic management module includes a server network performance comparison unit, a focus status analysis unit, and a dynamic resource allocation strategy generation unit. The server network performance comparison unit is used to collect network performance data of the corresponding server in real time and compare the transmission delay in the network performance data of the corresponding server with a first preset threshold. The focus state analysis unit is used to obtain the real-time focus state of each participant. The dynamic resource allocation strategy generation unit automatically executes a dynamic resource allocation strategy based on the real-time focus status of each participant obtained by the focus status analysis unit, thereby reducing the video stream resource consumption of non-core participants.

[0020] like Figure 2 As shown, this embodiment provides a method for remote monitoring of video conferencing based on meeting room environment information, including: Each video conference corresponds to multiple endpoints and one server. Each endpoint is bound to each participant in the corresponding video conference, and different endpoints exchange information through the server. Network performance data, audio quality data, video data, and physical environment data of each participant in the video conference are collected separately to construct multi-dimensional environmental data of the corresponding participant's venue, and then transmitted to the server through the corresponding edge device of the participant. The network performance data includes one or more of network jitter, latency, and bandwidth; the audio quality data includes the signal-to-noise ratio of the audio signal received by the corresponding edge from the server; the video image data includes the smoothness of the video image received by the corresponding edge from the server; the physical environment data includes one or more of air quality, temperature, and humidity, as well as the noise level of the venue where the participants are located.

[0021] Based on the multi-dimensional environmental data of the participants in the meeting room, the server uses a weighted fusion algorithm to analyze and obtain the real-time meeting quality index of each participant in the corresponding video conference; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The analysis yields a real-time meeting quality index for each participant in the corresponding video conference, calculated using the following formula: Among them, MQI i This indicates the real-time meeting quality index of the corresponding video conference based on the i-th participant; Sn i Sa represents the network dimension quality sub-score corresponding to the most recently acquired network performance data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; i This represents the audio dimension quality sub-score corresponding to the most recently acquired audio quality data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Sv i This represents the video dimension quality sub-score corresponding to the most recently acquired video frame data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Se i This represents the environmental dimension quality sub-score corresponding to the most recently acquired physical environment data of the i-th participant in the corresponding video conference within the database's preset form; Wn, Wa, Wv, and We represent the preset weight coefficients for the network dimension, audio dimension, video dimension, and environmental dimension, respectively.

[0022] The system collects network performance data of the corresponding servers in real time. When the transmission delay in the network performance data of the corresponding server is greater than or equal to the first preset threshold, it automatically executes a dynamic resource allocation strategy based on the real-time focus status of the participants to reduce the video stream resource consumption of non-core participants.

[0023] The specific steps for obtaining the real-time focus status of participants are as follows: Step 1: Capture the facial images of the attendees using a camera, and extract the temporal features of the head pose angle, eye gaze direction vector, and mouth aspect ratio from the facial images of the attendees; the head pose angle represents the angle between the facial plane and the camera plane in the facial key point detection results obtained through face recognition. Step 2: Based on the obtained time-series characteristics, calculate the real-time attention level assessment value of the participants. The calculation formula is as follows: Among them, AS i FF represents the real-time attention level assessment value of the i-th participant; iFG represents the cosine of the maximum head posture angle of the i-th participant within the most recent preset unit time; i FY represents the ratio of the duration during which the eye gaze direction vector of the i-th participant points to the corresponding edge of the screen area in the video conference within the most recent preset unit of time; i denoted as the number of yawning actions identified in the temporal features of the mouth aspect ratio of the i-th participant within the most recent preset unit time; e represents the natural constant; λ1, λ2, and λ3 represent the weight coefficients of the first feature, the second feature, and the third feature, respectively. In the process of obtaining the number of yawning actions identified from the temporal features of the mouth aspect ratio (in this embodiment, the mouth aspect ratio refers to the quotient of mouth width divided by mouth height), each yawning action corresponds to a continuous temporal segment in the temporal features of the mouth aspect ratio. The yawning action is identified by comparing the temporal change relationship of the mouth aspect ratio within the corresponding continuous temporal segment with the temporal change relationship of the mouth aspect ratio corresponding to the yawning action in the preset form of the database.

[0024] The automatic execution of the dynamic resource allocation strategy includes: Based on the real-time focus status and real-time meeting quality index of the participants, calculate the resource allocation priority evaluation value for each participant in the corresponding video conference; the resource allocation priority evaluation value for the i-th participant in the corresponding video conference is the product of the i-th participant's real-time focus status and the i-th participant's real-time meeting quality index. Based on the priority evaluation value of the corresponding resource allocation, the participants in the corresponding video conference, excluding the current speaker, are sorted in descending order to construct a candidate sequence of participants for optimization in the corresponding video conference. The difference between the transmission delay and the first preset threshold in the network performance data of the corresponding server at the current time is used to determine the number of attendees to be optimized in the database, which is recorded as the reference value for the number of optimized objects. The set of attendees in the candidate sequence of attendees to be optimized, counting the reference value of the number of optimized objects from the end to the beginning, is recorded as the set of attendees to be optimized. All attendees in the set of attendees to be optimized are considered as non-core attendees at the current time. The dynamic resource allocation strategy for non-core attendees is obtained. The dynamic resource allocation strategy includes reducing the resolution of the video stream subsequently received by the attendees, and the reduced resolution of the video stream received by the attendees. The reduced resolution of the video stream received by the attendees is obtained by querying the video stream resolution bound in the database based on the resource allocation priority evaluation value of the attendees at the current time. When the transmission latency in the network performance data of the corresponding server is less than the first preset threshold, there is no need to automatically execute the resource allocation strategy.

[0025] In this embodiment, when automatically executing the dynamic resource allocation strategy, the identity of the speaker and attendee at the corresponding time point in the meeting will be automatically identified; the audio stream of the speaker and attendee will be marked as the highest priority core media stream, and guaranteed network bandwidth will be allocated to it. At the same time, the video stream resolution of one or more non-speaker attendees (non-core attendees) is automatically reduced to save overall network bandwidth; After the dynamic resource allocation strategy is executed automatically, the dynamic resource allocation management module will record the triggering conditions, execution actions, and changes in the meeting quality index after execution of the corresponding dynamic resource allocation strategy, forming an optimization log; This facilitates administrators in updating the weighting coefficients in the weighted fusion algorithm and the trigger thresholds (first preset thresholds) for dynamic resource allocation strategies.

[0026] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0027] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for remote monitoring of video conferencing based on meeting room environment information, characterized in that: include: Each video conference corresponds to multiple endpoints and one server. Each endpoint is bound to each participant in the corresponding video conference, and different endpoints exchange information through the server. Network performance data, audio quality data, video data, and physical environment data of each participant in the video conference are collected separately to construct multi-dimensional environmental data of the corresponding participant's venue, and then transmitted to the server through the corresponding edge device of the participant. Based on the multi-dimensional environmental data of the participants in the meeting room, the server uses a weighted fusion algorithm to analyze and obtain the real-time meeting quality index of each participant in the corresponding video conference; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The system collects network performance data of the corresponding servers in real time. When the transmission delay in the network performance data of the corresponding server is greater than or equal to the first preset threshold, it automatically executes a dynamic resource allocation strategy based on the real-time focus status of the participants to reduce the video stream resource consumption of non-core participants.

2. The video conferencing remote monitoring method based on meeting room environment information according to claim 1, characterized in that: The network performance data includes one or more of network jitter, latency, and bandwidth; the audio quality data includes the signal-to-noise ratio of the audio signal received by the corresponding edge from the server; the video image data includes the smoothness of the video image received by the corresponding edge from the server; the physical environment data includes one or more of air quality, temperature, and humidity, as well as the noise level of the venue where the participants are located.

3. The video conferencing remote monitoring method based on meeting room environment information according to claim 1, characterized in that: The analysis yields a real-time meeting quality index for each participant in the corresponding video conference, calculated using the following formula: Among them, MQI i This indicates the real-time meeting quality index of the corresponding video conference based on the i-th participant; Sn i Sa represents the network dimension quality sub-score corresponding to the most recently acquired network performance data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; i This represents the audio dimension quality sub-score corresponding to the most recently acquired audio quality data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Sv i This represents the video dimension quality sub-score corresponding to the most recently acquired video frame data of the i-th participant in the corresponding video conference, within a pre-defined form in the database; Se i This represents the environmental dimension quality sub-score corresponding to the most recently acquired physical environment data of the i-th participant in the corresponding video conference within the database's preset form; Wn, Wa, Wv, and We represent the preset weight coefficients for the network dimension, audio dimension, video dimension, and environmental dimension, respectively.

4. The video conferencing remote monitoring method based on meeting room environment information according to claim 1, characterized in that: The specific steps for obtaining the real-time focus status of participants are as follows: Step 1: Capture the facial images of the attendees using a camera, and extract the temporal features of the head pose angle, eye gaze direction vector, and mouth aspect ratio from the facial images of the attendees; the head pose angle represents the angle between the facial plane and the camera plane in the facial key point detection results obtained through face recognition. Step 2: Based on the obtained time-series characteristics, calculate the real-time attention level assessment value of the participants. The calculation formula is as follows: Among them, AS i FF represents the real-time attention level assessment value of the i-th participant; i FG represents the cosine of the maximum head posture angle of the i-th participant within the most recent preset unit time; i FY represents the ratio of the duration during which the eye gaze direction vector of the i-th participant points to the corresponding edge of the screen area in the video conference within the most recent preset unit of time; i denoted as the number of yawning actions identified in the temporal features of the mouth aspect ratio of the i-th participant within the most recent preset unit time; e represents the natural constant; λ1, λ2, and λ3 represent the weight coefficients of the first feature, the second feature, and the third feature, respectively. In the process of obtaining the number of yawning actions identified from the temporal features of the mouth's aspect ratio, each yawning action corresponds to a continuous temporal segment in the temporal features of the mouth's aspect ratio. The yawning action is identified by comparing the temporal change relationship of the mouth's aspect ratio within the corresponding continuous temporal segment with the temporal change relationship of the mouth's aspect ratio corresponding to the yawning action in the database's preset form.

5. The video conferencing remote monitoring method based on meeting room environment information according to claim 1, characterized in that: The automatic execution of the dynamic resource allocation strategy includes: Based on the real-time focus status and real-time meeting quality index of the participants, calculate the resource allocation priority evaluation value for each participant in the corresponding video conference. Based on the priority evaluation value of the corresponding resource allocation, the participants in the corresponding video conference, excluding the current speaker, are sorted in descending order to construct a candidate sequence of participants for optimization in the corresponding video conference. The difference between the transmission delay and the first preset threshold in the network performance data of the corresponding server at the current time is used to determine the number of attendees to be optimized in the database, which is recorded as the reference value for the number of optimized objects. The set of attendees in the candidate sequence of attendees to be optimized, counting the reference value of the number of optimized objects from the end to the beginning, is recorded as the set of attendees to be optimized. All attendees in the set of attendees to be optimized are considered as non-core attendees at the current time. The dynamic resource allocation strategy for non-core attendees is obtained. The dynamic resource allocation strategy includes reducing the resolution of the video stream subsequently received by the attendees, and the reduced resolution of the video stream received by the attendees. The reduced resolution of the video stream received by the attendees is obtained by querying the video stream resolution bound in the database based on the resource allocation priority evaluation value of the attendees at the current time. When the transmission latency in the network performance data of the corresponding server is less than the first preset threshold, there is no need to automatically execute the resource allocation strategy.

6. The video conferencing remote monitoring method based on meeting room environment information according to claim 5, characterized in that: The resource allocation priority evaluation value for the i-th participant in the corresponding video conference is the product of the i-th participant's real-time focus status and the i-th participant's real-time conference quality index.

7. A video conferencing remote monitoring system based on meeting room environment information, employing the video conferencing remote monitoring method based on meeting room environment information as described in any one of claims 1-6, characterized in that, include: The multi-dimensional data acquisition module for meeting room environment information is used to collect network performance data, audio quality data, video image data and physical environment data of each participant in the video conference, construct the multi-dimensional environment data of the meeting room for the corresponding participant, and transmit it to the server through the corresponding edge device of the participant. The real-time meeting quality analysis module controls the server to analyze the multi-dimensional environmental data of the meeting room based on the participants, and obtain the real-time meeting quality index of each participant in the corresponding video conference through a weighted fusion algorithm; and presents the real-time meeting quality index of all participants in the corresponding video conference to the remote administrator through a graphical user interface. The resource allocation dynamic management module is used to collect network performance data of the corresponding servers in real time. When the transmission delay in the network performance data of the corresponding server is greater than or equal to the first preset threshold, the dynamic resource allocation strategy is automatically executed based on the real-time focus status of the participants to reduce the video stream resource consumption of non-core participants.

8. The video conferencing remote monitoring system based on meeting room environment information according to claim 7, characterized in that: The resource allocation dynamic management module includes a server network performance comparison unit, a focus status analysis unit, and a dynamic resource allocation strategy generation unit. The server network performance comparison unit is used to collect network performance data of the corresponding server in real time and compare the transmission delay in the network performance data of the corresponding server with a first preset threshold. The focus state analysis unit is used to obtain the real-time focus state of each participant. The dynamic resource allocation strategy generation unit automatically executes a dynamic resource allocation strategy based on the real-time focus status of each participant obtained by the focus status analysis unit, thereby reducing the video stream resource consumption of non-core participants.