Cross-domain cooperative monitoring method constructed based on endogenous characteristics of enterprise-level video conference system
By building a cross-domain collaborative monitoring method in an enterprise-level video conferencing system, audio and video devices are monitored in real time and alarm information is pushed, which solves the problem of difficult fault location in cross-domain collaboration, realizes efficient fault handling and global visualization analysis without adding new hardware, and improves the reliability and operation and maintenance efficiency of the system.
Patent Information
- Application Number
- CN202511286419.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-11-28
AI Technical Summary
Enterprise-level video conferencing systems suffer from information silos due to fault propagation in cross-domain collaboration, making fault localization difficult and requiring on-site support from professional personnel, which makes it difficult to meet the timeliness requirements of business operations.
Based on the inherent characteristics of enterprise-level video conferencing systems, a cross-domain collaborative monitoring method is constructed. This method monitors the health status of audio and video devices in real time at the local meeting room, pushes alarm information to remote meeting rooms, and centrally reports it to a multi-point control unit, generating a global visual alarm status interface to achieve cross-domain collaborative processing.
It enables real-time monitoring of the status of all nodes' devices and cross-domain collaborative processing without the need for additional hardware, breaks down the silos of fault information transmission, provides comprehensive system health assessment and operation and maintenance decision support, and improves the reliability and operation and maintenance efficiency of video conferencing systems.
Smart Images

Figure CN121037526A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of enterprise video conferencing, in particular to a cross-domain collaborative monitoring method based on the endogenous characteristics of an enterprise video conferencing system. BACKGROUND
[0002] Enterprise video conferencing is usually in the form of multipoint video conferencing. The system is based on MCU architecture and is composed of a central service subsystem and multiple video conference site subsystems. It is deeply integrated with the internal network system of an enterprise by relying on private network hardware. Its high security and high concurrency performance advantages make it a core carrier for cross-domain collaboration of modern organizations. In large meeting rooms commonly used in enterprise video conferences, the video conference site system is deployed in a split interconnection mode. The device networking is complex, the devices are heterogeneous, and there are high-density cable connections.
[0003] Unlike internet-based video conference software, the reliability maintenance of the MCU video conference system requires professional personnel to be on site. For meetings involving sensitive enterprise content, technical personnel who are not participants are prohibited from entering the meeting room, and the speed of emergency troubleshooting is difficult to meet the timeliness of video conference services. Secondly, in the multi-domain collaborative business mode of enterprise video conferencing across organizations and regions, cross-domain fault transmission often results in information silos. Specifically, local meeting rooms cannot sensitively perceive abnormalities, and continuous transmission leads to asynchronous disconnection of audio and video signals at remote sites, severely damaging the continuity of meeting interactions.
[0004] In summary, the MCU video conference system provides strong information technology support for cross-domain collaboration of enterprise organizations due to its information interaction, spatial span, scene realism, and high security and high concurrency that internet video conference software does not have. However, its hardware heterogeneity and complex networking require high reliability guarantee costs. The present application proposes a high-reliability monitoring method for MCU conference systems based on the endogenous characteristics of the system, which does not require additional hardware for multi-domain meeting room collaboration. It has a remote-local-cross-domain three-level linkage mechanism for reliability guarantee, which has important engineering value for improving the operation efficiency of video conference systems and ensuring the continuity of key meetings. SUMMARY
[0005] The present application aims to provide a cross-domain collaborative monitoring method based on the endogenous characteristics of an enterprise video conferencing system to solve the problems raised in the background.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0007] A cross-domain collaborative monitoring method based on the endogenous characteristics of an enterprise video conferencing system, comprising the following steps:
[0008] Step S1, the local conference hall monitors the audio and video equipment health state in real time, and outputs the alarm information through the video display device;
[0009] Step S2, the local alarm information is pushed to other conference halls of the same conference, and is displayed in real time at the remote conference hall;
[0010] Step S3, the whole network alarm information is centrally reported to the multipoint control unit, and after aggregation and deduplication, is pushed to the conference management platform to generate a global visual alarm situation interface.
[0011] In the application, in the step S1, the local real-time monitoring specifically comprises:
[0012] Step S101, a state detection instruction is sent to the audio subsystem through the centralized control system to obtain the audio equipment health state;
[0013] Step S102, a state detection instruction is sent to the video subsystem through the centralized control system to obtain the video equipment health state;
[0014] Step S103, the detection result is sent to the local video conference terminal through the video conference control protocol.
[0015] In the application, in the step S2, the alarm information display specifically comprises:
[0016] Step S201, the video conference terminal sends the detection result to the video splicing processing unit in the form of streaming media data;
[0017] Step S202, the video splicing processing unit sends the video stream containing the alarm text picture to the local conference hall video display device;
[0018] Step S203, the alarm information is output and displayed in real time in the form of a video picture at the local conference hall.
[0019] In the application, in the step S2, the local alarm pushing specifically comprises:
[0020] Step S211, the local video conference terminal encodes and packages the video data containing the alarm content into a real-time transport protocol packet;
[0021] Step S212, the real-time transport protocol packet is sent to the multipoint control unit through a user datagram protocol stream;
[0022] Step S213, the multipoint control unit generates a single video stream after mixing and sends it to the remote video conference terminal of the same conference in real time.
[0023] In the application, in the step S3, the whole network alarm information centralized reporting specifically comprises:
[0024] Step S301, each video conference terminal carries local alarm information with a timestamp to the multipoint control unit;
[0025] Step S302, the multipoint control unit aggregates and removes the alarm based on the entry dimension and the time window;
[0026] Step S303, the multipoint control unit reports the processed alarm information to the conference management platform.
[0027] In the present application, the global visual alarm situation interface generated in step S3 specifically includes:
[0028] Step S311, the conference management platform receives the global alarm event stream in real time through a distributed message queue;
[0029] Step S312, the alarm data is cleaned, normalized and subjected to multi-dimensional correlation analysis;
[0030] Step S313, a dynamic visualization interface is constructed based on a vector topological layer, realizing the visualization mapping of alarm hot area distribution and time sequence evolution trend.
[0031] In the present application, the state detection instruction includes:
[0032] The channel state detection instruction of the audio subsystem;
[0033] The signal stability detection instruction of the video subsystem;
[0034] The device connection state detection instruction.
[0035] In the present application, the alarm text picture includes:
[0036] The conference name;
[0037] The device name;
[0038] The device fault type identifier;
[0039] The fault occurrence time;
[0040] The recommended treatment measures.
[0041] In the present application, the real-time transport protocol packet encapsulation includes:
[0042] Alarm content coding;
[0043] Local conference identifier information;
[0044] Timestamp information.
[0045] In the present application, the aggregation and deduplication processing includes:
[0046] Alarm classification based on terminal ID and conference ID;
[0047] Merging of repeated alarms within a time window;
[0048] Alarm priority determination.
[0049] Compared with the prior art, the beneficial effects of the present application are:
[0050] 1. The present application realizes real-time monitoring and cross-domain collaborative processing of full-node device states without adding information devices by multiplexing the existing signaling channel and video stream transmission mechanism of MCU architecture. The system completely relies on the endogenous functions of existing hardware such as video conference terminals and MCUs, avoiding the network complexity problem caused by external device incremental deployment;
[0051] 2. The present application innovatively uses video code stream embedding technology to present multi-source alarm information of different conference sites in the same conference on local display devices in real time. The alarm information is superimposed on the video picture through OSD subtitle, and the alarm level is distinguished by hierarchical color identification (red / yellow / green). The display content includes fault type, occurrence time and processing suggestion. This mechanism effectively breaks the fault information transmission island in traditional systems;
[0052] 3. The three-state analysis system constructed based on space-time dimensions provides comprehensive system health evaluation. The distribution and influence range of current alarms are visualized, the alarm evolution trend is displayed on the time axis, and the fault conduction analysis between conference sites based on geographic topology. This model supports statistical analysis according to conference sessions, conference domains, device types and other dimensions, providing data support for operation and maintenance decisions. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 The enterprise-level video conference system multi-domain business model architecture of the present application;
[0054] Figure 2 The local conference monitoring alarm time sequence flowchart of the present application;
[0055] Figure 3 The video conference cross-domain conference alarm pushing schematic diagram of the present application;
[0056] Figure 4 The global alarm hub collection schematic diagram of the present application. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0058] Embodiment one
[0059] The present application is directed to the problem of cross-domain fault conduction leading to information silos and difficulty in fault location in enterprise-level video conference systems. Through the use of a cross-domain collaborative monitoring system constructed using intrinsic characteristics of the system, precise positioning and agile handling of video conference system faults are achieved. Due to the involvement of distributed audio and video device state monitoring and cross-domain collaborative processing, the following technical objectives are achieved relying on the MCU hardware architecture:
[0060] 1. Real-time automated monitoring of local conference audio and video device status (detection delay ≤ 500 ms);
[0061] 2. Establishing a real-time sharing mechanism for alarm information across conference domains (end-to-end transmission delay ≤ 3 s);
[0062] 3. Building a globally visualized alarm situation awareness platform (data refresh period ≤ 1 s);
[0063] 4. Ensuring system reliability in the presence of network jitter (packet loss rate ≤ 20%) and device offline (reconnection timeout ≤ 30 s).
[0064] Based on the above technical requirements, the present embodiment adopts the "multi-domain collaborative video conference monitoring system" architecture as shown in Figure 1
[0065] Centralized control subsystem: deployed in each local conference, including:
[0066] Central control host: dual-core processor, connected to audio and video devices through RS-485 bus
[0067] Diagnosis module: integrated audio analysis unit (sampling rate ≥ 48 kHz) and video analysis unit (supporting 4K resolution detection)
[0068] Alarm generator: fault level evaluation algorithm is implemented as follows:
[0069] L = α·S + β·D + γ·F
[0070] Where L represents the alarm level, taking values in the range 1-5 (1 being the lowest and 5 being the highest), S represents signal distortion, taking values in the range 0-1 (0 indicating no distortion and 1 indicating complete distortion), D is the duration, in minutes, and F is the influence range coefficient, taking values 1-3 (1 indicating a single device, 2 indicating a subsystem, and 3 indicating the entire field domain), α, β, γ are weighting coefficients, with default values of 0.6, 0.3, and 0.1, respectively.
[0071] Multipoint control unit: deployed in the data center, containing:
[0072] Mixed flow processor: support H.265 encoding, maximum concurrent processing 64 video streams;
[0073] Alarm aggregation engine: time window-based alarm deduplication algorithm window size Δt = 5min;
[0074] Topology management module: maintain dynamic conference connection matrix M n×n ;
[0075] Elastic layer: based on Kubernetes to realize automatic scaling, load balancing algorithm:
[0076]
[0077] Where, W i represents the weight value of the i-th node, C i represents the CPU core number of the i-th node, Q i represents the current queue depth of the i-th node, Q max represents the maximum queue depth allowed by the system.
[0078] Conference management platform: based on microservice architecture, including:
[0079] Message queue service: Kafka cluster, throughput ≥100,000 / s
[0080] Analysis engine: sliding window model is used for time series analysis, window function:
[0081]
[0082] Where, W(t) represents the weighted alarm value at time point t, A i represents the original alarm value at the i-th time point, N represents the sliding window size, the default value is 30 (unit: time point), and λ represents the decay coefficient, the default value is 0.1.
[0083] Security protection module: using national encryption SM4 and mTLS authentication.
[0084] As Figures 2-4 shown, the specific process of the cross-domain cooperative monitoring method executed on the architecture includes:
[0085] Step S1: The local control processing subsystem initiatively issues monitoring control instructions to the conference host and the splicing processor, and then initiates diagnostic instructions to the audio subsystem and the video subsystem through the conference host and the splicing processor respectively. The conference host and the splicing processor respectively return the diagnostic results of the audio subsystem and the video subsystem to the central control in the form of alarms, and the central control sends the alarm information of different types of local fault sources to the video conference terminal. The local video conference terminal loads the alarm information in the form of text pictures into the video stream and displays the alarm text pictures on the main screen in the meeting room through the splicing processor control, so as to intuitively remind the user.
[0086] Dynamic visualization interface generation sub-process:
[0087] The data preprocessing adopts sliding window normalization:
[0088]
[0089] wherein, represents the normalized alarm value at time point t, A i represents the original alarm value at the i-th time point, w represents the size of the sliding window, the default value is 15, ω represents the decay factor, the default value is 0.9.
[0090] The topology mapping uses three-dimensional coordinate conversion:
[0091]
[0092] wherein, (x', y', z') represents the converted three-dimensional coordinates, (x, y, z) represents the original physical coordinates, (x0, y0, z0) represents the multi-point control unit reference coordinates, k x , k y represents the x / y axis scaling factor, θ represents the rotation angle, and z0 represents the height reference value.
[0093] The heat map generation is based on kernel density estimation:
[0094]
[0095] wherein, H(x, y, t) represents the heat value at coordinate (x, y,) at time t, L i represents the grade value of the i-th alarm, h represents the broadband parameter, the default value is 0.8, (x i , y i ) represents the coordinate position of the i-th alarm.
[0096] Step S2: The local control system in the local conference (A1) sends the audio and video alarm to the local video conference terminal, the video conference terminal reports the alarm information to the backend video conference multipoint control system, the multipoint control system distributes the alarm text information of the reserved entry conference topology conference A1 to the video conference terminals of the same conference entry conference B1 and C1 in the form of video stream, and the video conference terminals of each entry conference send the alarm text picture of conference A1 to the splicing processor, and the splicing processor controls the alarm text picture to be displayed on the conference display, thereby providing the participants of the same conference across the domain with collaborative basic information.
[0097] The intelligent analysis engine processing sub-process is as follows:
[0098] The improved DBSCAN algorithm is used for alarm aggregation.
[0099]
[0100] wherein, ε t represents the dynamic density threshold of time ε t , ε0 represents the basic density threshold, the default value is 1.0, α represents the sensitivity coefficient, the default value is 0.5, N t represents the number of alarms in the current time window, N avg represents the number of historical alarms.
[0101] The Bayesian network is used for root cause positioning.
[0102]
[0103] wherein, P(R|E) represents the posterior probability of the root cause E under the evidence E, P(E|R) represents the likelihood probability of the root cause R generating the evidence E, P(R) represents the prior probability of the root cause R, and k represents the total number of possible root causes.
[0104] Priority determination model:
[0105]
[0106] wherein, Priority represents the normalized priority level score (0-1), L represents the alarm level (1-5), I represents the influence range coefficient (1-3), C represents the continuity index (0-1), and T represents the timeliness decay factor (0-1).
[0107] Step S3: The local real-time alarm information is reported to the multi-point control unit with a timestamp, and the multi-point control unit aggregates and removes the alarm based on the entry dimension and time window, and further reports to the conference management platform. The conference management platform analyzes the real-time and historical alarm information according to different dimensions such as conference and site, forms an intuitive visual global first-level alarm view and a second-level alarm detail view, and provides system operation reliability analysis support for the video conference system administrator. At the same time, based on the alarm details, the system administrator provides non-entry technical support for non-professional personnel in the front-end conference site to ensure the business continuity of the video conference system.
[0108] Abnormal fault tolerance mechanism:
[0109] Network jitter processing: dynamic retransmission interval: t retry = min(2 n-1 , 30s), where n represents the current retry number, the maximum retransmission interval is limited to 30 seconds, and the current cache queue is greater than or equal to 1000;
[0110] Device offline processing: heartbeat detection interval 5s, data temporary storage MCU local disk ≥24h, compensation synchronization algorithm:
[0111]
[0112] Where S new represents the corrected state value, S local represents the locally recorded state value, represents the sequence number of the remote i-th event, represents the corresponding time sequence number recorded locally
[0113] Data consistency guarantee: RAFT protocol synchronization state, CRC-32 check mechanism, allowing a maximum of 3s inconsistent window.
[0114] Example two
[0115] Suppose a certain enterprise holds a video conference involving A1, B1, C1 three sites, and the system detects:
[0116] Audio howling occurs at A1 site 15:00:00 (S=0.8, lasts for 3 minutes)
[0117] L=0.6x0.8+0.3x3+0.1x2=1.28→2
[0118] Where the signal distortion S is 0.8 (severe distortion), the duration D is 3 (minutes), the influence range F is 2 (affecting the audio subsystem), and the weighting coefficient is the default value (0.6, 0.3, 0.1).
[0119] The MPCU aggregates alarms in the time window 15:00-15:05: 3 identical alarms are merged, D is updated to 3→5, and the final severity is calculated:
[0120]
[0121] wherein the alarm severity L i is 2, the bandwidth h = 0.8, and d i denotes the geographical distance weight (after normalization, the value is 0-1) of each venue to A1.
[0122] Root cause analysis result: microphone hardware failure probability: 72%, acoustic feedback probability: 65%, network envelope distortion probability: 23%
[0123] Calculation basis:
[0124]
[0125] System automatically executes:
[0126] A1 venue: prompt to replace the microphone (confidence 85%).
[0127] All venues: temporarily reduce the audio gain by 6dB.
[0128] Plan a second diagnosis at 15:20.
[0129] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one from another entity or action without necessarily requiring or implying that there is any such actual relationship or order between such entities or actions. Also, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0130] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A cross-domain collaborative monitoring method based on the inherent characteristics of enterprise-level video conferencing systems, characterized in that, Includes the following steps: Step S1: The local meeting room monitors the health status of audio and video equipment in real time and outputs alarm information through video display devices; Step S2: Push the local alarm information containing the meeting room name to other meeting rooms in the same meeting and display it in real time in the remote meeting room; Step S3: Alarm information from different meeting venues across the entire network is centrally reported to the multi-point control unit. After aggregation and deduplication, it is pushed to the meeting management platform to generate a global visual alarm status interface.
2. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 1, characterized in that: In step S1, local real-time monitoring specifically includes: Step S101: Send a status detection command to the audio subsystem through the centralized control system to obtain the health status of the audio equipment; Step S102: Send a status detection command to the video subsystem through the centralized control system to obtain the health status of the video equipment; Step S103: Send the detection results to the local video conferencing terminal via the video conferencing control protocol.
3. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 1, characterized in that: In step S2, the alarm information display specifically includes: Step S201: The video conferencing terminal sends the detection results to the video splicing processing unit in the form of streaming media data; In step S202, the video splicing processing unit sends the video stream containing the alarm text image to the local meeting room video display device; Step S203: The alarm information is displayed in real time in the local meeting room in the form of video footage; In step S204, after the alarm fault is handled, the alarm information disappears from the video screen.
4. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 1, characterized in that: In step S2, the local alarm push specifically includes: Step S211: The local video conferencing terminal encapsulates the video data containing the alarm content into a real-time transmission protocol packet. Step S212: Send the real-time transmission protocol packet along with the unique identifier of the venue to the multipoint control unit via the user datagram protocol stream; Step S213: After the multi-point control unit mixes the streams, it generates a single video stream and sends the high-level alarm data to the remote video conferencing terminal of the same conference in real time. Step S214: The remote video conferencing terminal of the same meeting controls the display of alarm information in the form of text and screen in the remote meeting room through the remote splicing processor. In step S215, after the alarm fault is handled, the alarm information disappears from the video screen.
5. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 1, characterized in that: In step S3, the centralized reporting of alarm information across the entire network specifically includes: Step S301: Each video conferencing terminal reports its local alarm information, along with the timestamp, the unique identifier of the meeting venue, and the unique identifier of the meeting, to the multipoint control unit. Step S302: The multi-point control unit aggregates and deduplicates alarms based on the membership dimension and time window; In step S303, the multi-point control unit reports the processed alarm information to the conference management platform.
6. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 1, characterized in that: In step S3, the generation of the global visual alarm status interface specifically includes: Step S311: The meeting management platform receives the global alarm event stream in real time through a distributed message queue; Step S312: Clean, normalize, and perform multi-dimensional correlation analysis on the alarm data; Step S313: Construct a dynamic visualization interface based on the vector topology layer to realize the visualization mapping of alarm hot zone distribution and temporal evolution trend, which includes alarm level standardization and meeting room information.
7. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 2, characterized in that: The status detection command includes: Control subsystem status detection commands; Audio subsystem channel status detection command; Signal stability detection command for video subsystem; Device connection status detection command.
8. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 3, characterized in that: The alarm text displayed on the local screen includes: Faulty equipment; Equipment fault type identification; Time of failure; Recommended measures.
9. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 4, characterized in that: The real-time transmission protocol packet encapsulation includes: Alarm content coding; Local venue signage information; Timestamp information.
10. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 4, characterized in that: The alarm text on the remote meeting room screen includes: Venue Name; Fault type and fault level; Timestamp information.
11. The cross-domain collaborative monitoring method based on the inherent characteristics of an enterprise-level video conferencing system as described in claim 5, characterized in that: The aggregation deduplication process includes: Alarm classification based on terminal ID and conference ID; Merge duplicate alarms within a time window; Alarm priority determination.