Remote Conference Intelligent Operation and Maintenance System and Operation and Maintenance Method
By conducting real remote meeting simulation and comparative analysis in the remote meeting system, the problems of low manual inspection efficiency and inability to comprehensively detect in the existing technology are solved, and a comprehensive, accurate and efficient intelligent diagnosis of the remote meeting system is achieved.
Patent Information
- Application Number
- CN202410751292.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-06-12
AI Technical Summary
The existing remote conferencing system is inefficient during manual inspection and cannot fully detect equipment operation and data transmission, and there is a problem of incomplete inspection.
Design a remote conference intelligent operation and maintenance system, and conduct real remote conference simulations between the main venue and the sub-venue, collect video and audio information, and conduct comparative analysis to determine whether there are abnormalities in the system.
It realizes comprehensive, accurate and efficient intelligent diagnosis of the remote meeting system, can accurately detect equipment operation abnormalities and data transmission problems, and improves the accuracy and comprehensiveness of the detection.
Smart Images

Figure CN118474294B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote conferencing, and particularly to an intelligent operation and maintenance system and an operation and maintenance method for remote conferencing. Background Art
[0002] Remote conferencing includes various communication methods such as video conferencing, audio conferencing, and instant messaging, enabling participants to conduct visual and auditory communication in real time at different locations. A remote conferencing system includes an MCU control unit and conference terminals. Each venue joins the remote conference by connecting the conference terminal to the MCU to hold a conference at each venue. To fully participate in a remote conference, each venue often includes complex devices and connection methods such as a sound amplification subsystem, a display subsystem, a sound pickup subsystem, and a conference access subsystem. Among them, the access subsystem is connected to other conference terminals and the MCU through a network. During use, each subsystem cooperates with each other to transmit audio and video signals through the access subsystem. In the whole process, it is necessary to ensure that each subsystem operates well and the network condition is good to enable normal communication.
[0003] In actual use, in order to ensure the normal operation of the system, it is often necessary to conduct manual inspections in advance to ensure that the system operates normally during use. However, if manual inspections are carried out on each venue, it is not only time-consuming and laborious, wasting human resources, but also the inspection efficiency is low. Moreover, manual inspections can usually only check whether the equipment is complete and whether the wiring between each device is accurate and good, etc., external conditions, and cannot check whether each device is operating normally and whether the data transmission between each device is normal, resulting in incomplete inspections. Summary of the Invention
[0004] One of the purposes of the present invention is to provide an intelligent operation and maintenance system for remote conferencing with high efficiency, accuracy, and comprehensiveness.
[0005] An intelligent operation and maintenance system for remote conferencing, which is applied to a remote conferencing system, includes an operation and maintenance subsystem. The operation and maintenance subsystem cooperates with the remote conferencing system to control. By sequentially switching each venue to the main venue and conducting real remote conference simulations between the main venue and each sub-venue after each switch, collecting the video and audio information of the main venue, as well as the video and audio information transmitted to each sub-venue, and comparing and analyzing the video of the main venue with the video received by the sub-venue, and the audio of the main venue and the audio received by the sub-venue respectively, so as to determine whether there is an abnormality in the remote conferencing system.
[0006] The beneficial effects of the present invention are as follows: 1. The present invention can perform a real remote meeting simulation between the main venue and each branch venue. By comparing and analyzing the video of the main venue collected during the simulation with the video received by the branch venue, as well as the audio of the main venue and the audio received by the branch venue respectively, it is possible to determine whether there are abnormalities in the remote meeting system. In this way, it can accurately detect whether there are any operating abnormalities in each device in the remote meeting system or whether there are any communication abnormalities between devices. If there is a device failure, the operation and maintenance terminal cannot receive the corresponding video and audio. If the device in the branch venue cannot receive the video and / or audio transmitted from the main venue, or the devices within the venue cannot communicate with each other, or although communication is possible, there are problems with the quality of the transmitted video and / or audio. That is, after comparing and analyzing the video of the main venue with the video received by the branch venue, and the audio of the main venue with the audio received by the branch venue respectively, if the differences in the video and / or audio before and after transmission do not meet the requirements, it indicates that there may be abnormalities in each device of the remote meeting system and / or communication abnormalities between devices. In this way, it can achieve comprehensive, accurate, and efficient intelligent diagnosis of the remote meeting system.
[0007] 2. The present invention collaboratively controls through the operation and maintenance subsystem and the remote meeting system, sequentially switches each venue to the main venue, and performs a real remote meeting simulation between the main venue and each branch venue after each switch. In this way, it can ensure that two-way communication is carried out between every two venues, and any abnormalities existing during both the forward transmission and the reverse transmission between venues can be detected, further improving the accuracy and comprehensiveness of the detection.
[0008] A preferred embodiment of the present invention is that the remote meeting system includes an MCU control unit, as well as a meeting terminal, a display device, a camera, and a speaker set in each venue. The operation and maintenance subsystem includes an operation and maintenance server, an operation and maintenance terminal, an operation and maintenance display device, and an operation and maintenance camera.
[0009] During the remote meeting simulation, the operation and maintenance server sends a setting signal to the MCU control unit. Through the MCU control unit, the main venue is set to the broadcast mode. The operation and maintenance server sends a control signal to the MCU control unit, and the MCU control unit establishes a remote meeting connection and calls the meeting terminals of all venues to join the meeting.
[0010] The operation and maintenance server sends a control signal to the operation and maintenance terminal of the current main venue, and the operation and maintenance terminal controls the operation and maintenance display device to play the first video according to the control signal.
[0011] The operation and maintenance terminal in the main venue controls the operation and maintenance camera to collect the first video, and transmits the collected first video to the MCU control unit. The MCU control unit then transmits the first video to the conference terminals in other branch venues; the second video is displayed through the display devices in each branch venue, and the sound in the second video is emitted through the speakers in each branch venue.
[0012] The operation and maintenance terminals located in each branch venue control the cameras in each branch venue to collect the second video, and transmit the collected second video to the operation and maintenance server through the operation and maintenance terminals.
[0013] The first video collected by the operation and maintenance camera is also transmitted to the operation and maintenance terminal. The operation and maintenance terminal transmits the first video to the operation and maintenance server. The operation and maintenance server is used to detect the pictures in the first video and the second video, and compare whether the picture contents are consistent. If they are inconsistent, an abnormal signal is sent.
[0014] Switch the main venue and perform the detection again by the same means to detect whether the video reverse transmission is normal.
[0015] In a preferred embodiment of the present invention, the remote conference system further includes a microphone. The operation and maintenance subsystem further includes a sound-emitting device and a sound-pickup device.
[0016] The operation and maintenance terminal is further used to control the sound-emitting device in the main venue to emit sound.
[0017] The microphone is used to amplify the sound.
[0018] The sound-pickup device in the main venue is used to collect the amplified sound and transmit the collected first audio information to the operation and maintenance terminal and the MCU control unit respectively; the operation and maintenance terminal transmits the first audio information to the operation and maintenance server.
[0019] The MCU control unit is further used to transmit the first audio information to the conference terminals in each branch venue. The speakers in each branch venue are used to play the second audio information formed after the conference terminals receive the first audio information. The sound-pickup devices in each branch venue are used to collect the second audio information played by the speakers and transmit it to the operation and maintenance terminal. The operation and maintenance terminal transmits the second audio information to the operation and maintenance server.
[0020] The operation and maintenance server detects the first audio information and the second audio information and compares whether they are consistent. If they are inconsistent, an abnormal signal is sent.
[0021] Switch the main venue and perform the detection again by the same means to detect whether the audio reverse transmission is normal.
[0022] In a preferred embodiment of the present invention, a daily inspection plan is stored in the operation and maintenance server. The operation and maintenance server starts a system timer, and when the trigger condition is reached, the inspection process is automatically started. Thus, unattended operation is realized, and the inspection is triggered autonomously throughout the process, which is more intelligent and efficient.
[0023] In a preferred embodiment of the present invention, the operation and maintenance server detects the IP addresses of the conference terminals in all the venues to be inspected. If it is found that a conference terminal is offline, an instruction is sent to the operation and maintenance terminal. The operation and maintenance terminal controls the power-on of the serial server in the venue, and automatically starts all the devices in the venue, including the remote conference terminals. After an operation and maintenance inspection is completed, an instruction is sent to the operation and maintenance terminal. The operation and maintenance terminal sends a shutdown instruction to the devices according to the control protocols of the respective devices to make them shut down, and then controls the serial server to power off in a preset order, so as to turn off all the devices in the venue, including the remote conference terminals.
[0024] In a preferred embodiment of the present invention, an exception handling plan is stored in the operation and maintenance server. When the operation and maintenance server detects an exception, the automatic handling plan is automatically triggered according to the configured conditions when data is returned during the inspection. The automatic handling plan includes restarting the device, adjusting the volume of the sound-producing device and the sound-pickup device, and adjusting the camera. The operation and maintenance server has a self-learning function, and self-learns according to each detected exception and the automatic handling plan, so as to automatically generate a new handling plan when a similar exception is detected.
[0025] In a preferred embodiment of the present invention, the operation and maintenance server detects the pictures in the first video and the second video in the same way. The principle is to detect the image content of the two pictures and the relative positions where the content appears in the pictures respectively, and then compare the detection results of the two pictures. If the content of the two pictures and the relative positions where each content appears are substantially the same, it is considered that the pictures are under the same scene.
[0026] In a preferred embodiment of the present invention, the steps of sound content detection are as follows:
[0027] Sound signal acquisition: The operation and maintenance terminal captures the sound signal in real time through the sound-pickup device;
[0028] Signal preprocessing: The captured sound signal is subjected to noise suppression and echo cancellation to improve the quality of subsequent processing;
[0029] Audio signal segmentation: According to the playing duration of the audio, the sound signal is cut into segments of equal length;
[0030] Sampling rate unification: The sampling rates of all audio segments are converted into a unified standard sampling rate;
[0031] Feature extraction: Mel-frequency cepstral coefficients are used to extract features from each audio frame;
[0032] Feature data standardization, standardize the extracted MFCC feature data for the input of the neural network;
[0033] It also includes:
[0034] Unify the frame length, if the audio frame length is insufficient or exceeds the standard length, perform padding or truncation to ensure that all input data lengths are consistent;
[0035] Neural network input, use the standardized feature data as input and pass it to the neural network;
[0036] Convolution layer processing, the convolution layer in the neural network performs multiple convolution calculations on the input audio features to extract deeper features;
[0037] Pooling layer, use the pooling layer after the convolution layer to reduce the feature dimension and extract the most important features;
[0038] Fully connected layer processing, pass the feature data extracted by the convolution layer and the pooling layer into the fully connected layer for further feature fusion and classification decision-making;
[0039] Activation function application, use the activation function in the fully connected layer to introduce non-linearity;
[0040] Output layer mapping, apply the Sigmoid function after the last fully connected layer to map the output to a probability value between 0 and 1;
[0041] Threshold judgment, set a threshold (such as 0.8), convert the output of the Sigmoid function into a binary judgment to determine whether the audio content matches the expectation;
[0042] Result output, output the detection result, including the similarity score and the judgment of whether the target audio content is detected.
[0043] A preferred embodiment of the present invention is that it further includes an operation and maintenance trend analysis module, and the operation and maintenance trend analysis module is used for: recording the detailed data of each inspection into the system, and the system incorporates the historical data into a neural network for trend analysis at regular intervals for neural network model training, for upgrading and iterating the neural network, and this neural network is directly used for the trend analysis of inspection parameters, for obtaining the possible future detection result values, and displaying this value to the user for reference, and giving an alarm when the warning condition is reached during the trend analysis of each parameter.
[0044] A preferred embodiment of the present invention is that the steps for training the trend analysis neural network model are:
[0045] Ensure the quality and integrity of the data, clean the data, and remove noise and outliers;
[0046] Perform data standardization or normalization to bring the data to the same scale;
[0047] Supplement data features, including collection time, meeting room parameters, environmental parameters, spatial parameters, etc.;
[0048] Divide the data into a training set, a validation set, and a test set;
[0049] Construct a trend analysis neural network model;
[0050] Use the training set data to train the neural network and adjust the hyperparameters;
[0051] Use the validation set to optimize the model and use the test set to evaluate the model performance.
[0052] A preferred embodiment of the present invention is that the method for constructing a trend analysis neural network model includes:
[0053] Input layer: This layer receives the input data, and the number of nodes is the same as the number of data features;
[0054] Convolutional layer: Use a 1D convolutional layer to extract local features in the time series data, set multiple convolutional kernels, and each convolutional kernel is responsible for extracting different features;
[0055] Pooling layer: After the convolutional layer, use a pooling layer to reduce the spatial dimension of the features and increase the invariance to data displacement;
[0056] Recurrent layer: Use an LSTM or GRU layer to handle the long-term dependencies in the data;
[0057] Fully connected layer: After the recurrent layer, include multiple fully connected layers for further feature learning and decision-making;
[0058] Dropout layer: Add a Dropout layer between or after the fully connected layers to reduce the risk of model overfitting;
[0059] Batch normalization layer: After the fully connected layer, a batch normalization layer can be added to accelerate the training process and improve the stability of the model;
[0060] Output layer: The output layer has only one node, and a linear activation function is used to output the prediction result.
[0061] The second object of the present invention is to provide a method for intelligent operation and maintenance of remote conferences. By sequentially switching each venue to the main venue and conducting real remote conference simulations between the main venue and each branch venue after each switch, collecting the video and audio information of the main venue, as well as the video and audio information transmitted to each branch venue, and comparing and analyzing the video of the main venue with the video received by the branch venue, and the audio of the main venue and the audio received by the branch venue respectively, to determine whether there are abnormalities in the remote conference system.
[0062] The overall advantages of the present invention are as follows:
[0063] By setting up automatic pre - plans to promptly handle faults occurring during inspections, automatically performing inspections according to preset strategies without manual intervention.
[0064] During the inspection process, problems existing in the system can be discovered in advance, including whether the sound and images are normal, whether the device control is normal, etc.
[0065] When problems are found, automatically remind the relevant person in charge to handle them according to the preset strategy, and at the same time record the processing results and timeliness.
[0066] Statistically analyze in multiple aspects such as failure rate and problem - handling efficiency based on the collected data.
[0067] Collect images of the operation and maintenance display device through an operation and maintenance camera and conduct inspections to simulate the inspection effect of human vision.
[0068] Pick up sound signals through a sound - picking device and conduct inspections to simulate the inspection effect of human hearing.
[0069] Simulate human voice by sending sounds to the microphone to check the sound - picking system.
[0070] Analyze the collected images through computer vision to determine whether the display system is normal.
[0071] Analyze the collected sound signals through a multi - layer neural network to determine whether the pronunciation of the sound amplification subsystem is normal. Brief Description of the Drawings
[0072] Appendix Figure 1 Shown is the hardware layout diagram of the venue.
[0073] Appendix Figure 2 Shown is the connection topology diagram of the operation and maintenance subsystem of the present invention.
[0074] Appendix Figure 3 Shown is the flow chart of the intelligent operation and maintenance system for remote conferences of the present invention. Detailed Embodiments
[0075] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described below are only used to explain the present invention and do not limit the protection scope of the present invention.
[0076] The terms "first", "second", etc. in the specification, claims, and embodiments of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0077] The present invention will be further described in detail below through preferred specific embodiments:
[0078] The reference numerals in the accompanying drawings of the specification include: operation and maintenance display device 1, operation and maintenance camera 2, sound pickup device 3, sound output device 4, screen 5, camera 6, speaker 7, microphone 8, seat 9, operation and maintenance terminal 10, operation and maintenance server 11.
[0079] The intelligent operation and maintenance system for remote conferences disclosed in this embodiment is applied to a remote conference system and includes an operation and maintenance subsystem. The operation and maintenance subsystem and the remote conference system are controlled in coordination. By sequentially switching each venue to the main venue and conducting real remote conference simulations between the main venue and each branch venue after each switch, the video and audio information of the main venue, as well as the video and audio information transmitted to each branch venue, are collected. Then, the video of the main venue is compared and analyzed with the video received by the branch venue, and the audio of the main venue is compared and analyzed with the audio received by the branch venue to determine whether there is an abnormality in the remote conference system.
[0080] The remote conference system includes an MCU control unit, as well as a conference terminal, a display device, a camera, and a speaker set in each venue. As shown in the attached Figure 2 The connection topology diagram of the operation and maintenance subsystem of the present invention is shown as follows. The operation and maintenance subsystem includes an operation and maintenance server, an operation and maintenance terminal, an operation and maintenance display device, an operation and maintenance camera, a sound pickup device, and a sound output device. The operation and maintenance display device, operation and maintenance camera, sound pickup device, and sound output device are the adopted hardware, which are arranged in each venue. The operation and maintenance terminal is deployed in each venue and is responsible for docking with various peripherals of the operation and maintenance system, reading peripheral data, and controlling the peripherals. The operation and maintenance terminal is connected to the operation and maintenance server, receives the control instructions of the operation and maintenance server, and uploads data to the operation and maintenance server. The operation and maintenance server is the central control computer of the operation and maintenance system, deploys the operation and maintenance system platform, and is connected to the operation and maintenance terminal, sending operation and maintenance instructions to the operation and maintenance terminal and receiving the data transmitted back by the operation and maintenance terminal.
[0081] As shown in the attached Figure 1The figure shows the hardware layout diagram of the meeting venue. The display device in the venue is a screen. The cameras in the venue are located directly in front of all seats. The microphones, voice devices, and sound pickup devices are all set on the seats. The voice is emitted through the voice devices to simulate real meeting speeches. The operation and maintenance display device is one of a display, a signal lamp, or a fixed marker. In this embodiment, a display is preferably used. When simulating a real meeting scenario, specific video content can be played, such as a QR code video. The operation and maintenance camera is installed facing the operation and maintenance display device for collecting the video played on the operation and maintenance display device.
[0082] The remote conference system includes an MCU control unit and conference terminals. In the remote conference system, MCU (Multipoint Control Unit) is the abbreviation of the Multipoint Control Unit. The MCU control unit is a key network device used to manage and control multiple participants in a video conference. The MCU allows users from different locations to communicate and collaborate through the video conference system. A conference terminal (Video Conference Terminal) refers to the terminal device used in the video conference system, which allows users to communicate in real time with other participants through video, audio, and data transmission. The video conference terminal can be a hardware device, a software application, or a combination of both.
[0083] In this embodiment, the peripheral device is connected to the operation and maintenance terminal to transmit data. The connection methods include network connection, RS232 serial port connection, etc. The operation and maintenance terminal is connected to the operation and maintenance server through the network to transmit data and receive operation and maintenance terminal instructions.
[0084] As shown in the appendix Figure 3 The figure shows the flow chart of the intelligent operation and maintenance system for the remote conference of the present invention.
[0085] Detect the IP addresses of the conference terminals in all the meeting venues to be inspected soon. If it is found that a conference terminal is offline, an instruction is sent to the operation and maintenance terminal to control the power-on of the serial port server in the venue so as to start all the devices in the venue including the remote conference terminal.
[0086] During the remote conference, the operation and maintenance server sends a setting signal to the MCU control unit to set the main venue to the broadcast mode through the MCU control unit. The MCU control unit controls the routing and distribution of signals during the simulation of the remote conference to ensure the synchronization and continuity of the signals.
[0087] Among them, the detection method for the video transmission in the venue is as follows:
[0088] The operation and maintenance server sends a control signal to the MCU control unit. The MCU control unit establishes a remote conference connection and calls all the conference terminals in the venues to join the conference;
[0089] The operation and maintenance server sends a control signal to the operation and maintenance terminal of the current main venue, and the operation and maintenance terminal controls the operation and maintenance display device to play the first video according to the control signal;
[0090] The operation and maintenance terminal of the main venue controls the operation and maintenance camera to collect the first video, and transmits the collected first video to the MCU control unit. The first video is transmitted to the conference terminals of other branch venues through the MCU control unit; the second video is displayed through the display devices of each branch venue, and the sound in the second video is emitted through the speakers of each branch venue;
[0091] The operation and maintenance terminals located in each branch venue control the cameras in each branch venue to collect the second video, and transmit the collected second video to the operation and maintenance server through the operation and maintenance terminals;
[0092] The first video collected by the operation and maintenance camera is also transmitted to the operation and maintenance terminal, and the operation and maintenance terminal transmits the first video to the operation and maintenance server. The operation and maintenance server is used to detect the pictures in the first video and the second video, and compare whether the picture contents are the same. If they are not the same, an abnormal signal is sent;
[0093] Switch the main venue and perform detection again by the same means to detect whether the video reverse transmission is normal.
[0094] After an operation and maintenance inspection is completed, an instruction is sent to the operation and maintenance terminal to let the operation and maintenance terminal send a shutdown instruction to the device according to the control protocol of each device to shut it down, and then the controlled serial server is powered off in a preset order to turn off all devices in the venue including remote conference terminals.
[0095] Since the picture captured by the operation and maintenance terminal is the live picture of the entire venue, rather than the display signal of a specific screen. Therefore, it is necessary to perform content recognition and processing on the picture through computer vision technology for further analysis. In this solution, it is necessary to detect the display screen, and then cut out the content it displays according to the detected display screen for further recognition. The operation and maintenance server detects the pictures in the first video and the second video in the same way, as follows:
[0096] Video stream capture, the operation and maintenance terminal uses the streaming media server to pull the real-time video stream from the RTSP stream addresses of the cameras in each branch venue and the operation and maintenance camera, and intercepts specific frames from the RTSP stream as pictures through the streaming media processing component;
[0097] Picture preprocessing, preprocess the intercepted picture data, such as scaling it to the input size required by the model (416*416 pixels), and perform normalization operations;
[0098] Feature extraction: Use convolutional neural networks to perform multi-level convolution operations on the input image, abstract image features layer by layer, and generate multi-scale feature maps;
[0099] Generate candidate regions for target detection. For some models, such as Faster R-CNN, RPN is used to quickly generate candidate target regions.
[0100] Bounding box regression, which refines the candidate region and adjusts the position and size of the bounding box to match the target more accurately;
[0101] Category prediction: Based on bounding box regression, the category of each candidate region is predicted to determine its category.
[0102] Non-maximum suppression, which handles overlapping prediction bounding boxes, selects the best detection results, and removes redundant and low-confidence predictions;
[0103] Post-processing of the results: threshold screening is performed on the processed detection results to retain only the object detection boxes that are higher than the preset confidence threshold;
[0104] Specific object processing, further processing the image for specific object categories detected;
[0105] Content recognition and comparison: decode or identify the content of the cut image and compare it with the image detected at the branch venue;
[0106] The comparison result is fed back. If the image content matches, the transmission is confirmed to be successful; if it does not match, an abnormal signal is issued.
[0107] The specific method of feature extraction is:
[0108] Select or initialize a filter, a small matrix that slides over the input data to extract features;
[0109] Input data preparation, the input data is usually a multidimensional array;
[0110] Local receptive field convolution multiplies each element of the filter by the local area of the corresponding position in the input data, and then sums the products to obtain a single output value;
[0111] Sliding window operation, sliding the filter over the entire input data, performing local receptive field convolution each time, thereby generating a new element for the output feature map;
[0112] Feature map padding: In order to control the size reduction of the feature map, padding can be added to the edge of the input data so that the size after the convolution operation is consistent with the input data or adjusted as needed;
[0113] Step size, which defines the step size for the filter to slide. A step size of 1 means sliding one pixel each time, and a step size greater than 1 can reduce the size of the feature map.
[0114] Feature map generation: Repeat the local receptive field convolution and sliding window operations until all regions of the input data are covered, generating a complete feature map to capture the local features in the input data.
[0115] The feature extraction also includes:
[0116] Multiple filters: Use multiple filters to process the input data in parallel. Each filter is responsible for extracting different features, and the final feature map is the set of outputs of all filters.
[0117] Activation function: Apply a non-linear activation function after the convolution operation to increase the non-linear expression ability of the model;
[0118] Size adjustment: Adjust the feature map according to the size after the convolution operation to meet the input requirements of the subsequent layers;
[0119] Pooling: After the convolutional layer, follow a pooling layer to reduce the spatial dimension of the feature map and extract more abstract features.
[0120] Compare the content of the first video image and the second video image to check if the content is consistent. The principle is to detect the image content and the relative position of the content in the two images respectively through the previously described video content detection method. Then compare the detection results of the two images. If the content of the two images and the relative positions of the content are generally the same, it is considered that the images are from the same scene.
[0121] In this embodiment, the method for video quality detection is as follows:
[0122] Sharpness: The clarity of edges and details in the image.
[0123] Texture: The complexity of repeated patterns in the image.
[0124] Color: The richness and saturation of colors in the image.
[0125] Contrast: The difference between bright and dark regions in the image.
[0126] Noise: Unwanted random variations or grains in the image.
[0127] Distortion: Compression, blurring, or other types of distortion that may occur in the image.
[0128] This solution adopts the BRISQUE detection algorithm. The BRISQUE algorithm uses the changes in the statistical characteristics of natural scenes to evaluate image quality. Features are extracted from the image, and then the learned model is used to predict the image quality score.
[0129] The detection method for audio transmission in the venue is as follows:
[0130] The operation and maintenance terminal is also used to control the sound-emitting device in the main venue to emit sound;
[0131] The microphone is used to amplify the sound;
[0132] The sound-pickup device in the main venue is used to collect the amplified sound and transmit the collected first audio information to the operation and maintenance terminal and the MCU control unit respectively; the operation and maintenance terminal transmits the first audio information to the operation and maintenance server;
[0133] The MCU control unit is also used to transmit the first audio information to the conference terminals in each branch venue. The speakers in each branch venue are used to play the second audio information formed after the conference terminal receives the first audio information. The sound-pickup devices in each branch venue are used to collect the second audio information played in the speakers and transmit it to the operation and maintenance terminal. The operation and maintenance terminal transmits the second audio information to the operation and maintenance server,
[0134] The operation and maintenance server detects the first audio information and the second audio information and compares whether they are consistent. If they are not consistent, an abnormal signal is sent;
[0135] Switch the main venue and perform detection again by the same means to detect whether the audio reverse propagation is normal.
[0136] Among them, after the operation and maintenance terminal captures the sound signal through the sound-pickup device, a series of processes are performed on the sound before the content can be recognized. The steps for sound content detection are as follows:
[0137] Sound signal acquisition: The operation and maintenance terminal captures the sound signal in real time through the sound-pickup device.
[0138] Signal preprocessing: Perform noise suppression and echo cancellation on the captured sound signal to improve the quality of subsequent processing.
[0139] Audio signal segmentation: According to the playing duration of the audio, the sound signal is cut into segments of equal length, such as 2 seconds per segment.
[0140] Sampling rate unification: Convert the sampling rates of all audio segments to a unified standard sampling rate, such as 16000 Hz.
[0141] Feature extraction: Use Mel Frequency Cepstral Coefficients (MFCC) to extract features from each audio frame, usually extracting 39-dimensional feature data.
[0142] Feature data normalization: Normalize the extracted MFCC feature data for the input of the neural network.
[0143] Unify frame length: If the audio frame length is insufficient or exceeds the standard length, pad or truncate it to ensure that all input data has the same length. In this solution, it is supplemented to a length of 400 frames.
[0144] Neural network input: Pass the normalized feature data as input to the neural network.
[0145] Convolution layer processing: The convolution layer in the neural network performs multiple convolution calculations on the input audio features to extract deeper features.
[0146] Pooling layer: Use the pooling layer after the convolution layer to reduce the feature dimension and extract the most important features.
[0147] Fully connected layer processing: Pass the feature data extracted by the convolution layer and the pooling layer into the fully connected layer for further feature fusion and classification decision-making.
[0148] Activation function application: Use activation functions such as ReLU in the fully connected layer to introduce non-linearity.
[0149] Output layer mapping: Apply the Sigmoid function after the last fully connected layer to map the output to a probability value between 0 and 1.
[0150] Threshold judgment: Set a threshold (such as 0.8) to convert the output of the Sigmoid function into a binary judgment to determine whether the audio content matches the expectation.
[0151] Result output: Output the detection result, including the similarity score and the judgment of whether the target audio content is detected.
[0152] The present invention analyzes the collected sound signals through a multi-layer neural network to determine whether the pronunciation of the sound amplification subsystem is normal. The multi-layer neural network can be various network models such as a convolutional neural network or a recurrent neural network.
[0153] The following are the steps for MFCC feature calculation:
[0154] Preprocessing: Perform pre-emphasis on the input audio signal to enhance the high-frequency part and highlight the details of the speech signal.
[0155] Frame segmentation: The audio signal is segmented into short-time frames, usually with a frame length of 20 - 40 milliseconds, and there is a certain overlap between frames to maintain the continuity of the signal.
[0156] Windowing: Apply a window function, such as the Hamming window, to each frame of data to reduce the discontinuity at the frame boundaries.
[0157] Fast Fourier Transform (FFT): Perform the fast Fourier transform on the windowed signal of each frame to obtain the frequency-domain representation of that frame.
[0158] Mel filter bank: Use a set of Mel filters to filter the frequency-domain signal. The Mel filters simulate the human auditory perception and map the signal to the Mel frequency scale.
[0159] Energy calculation: Calculate the energy of the signal passing through the Mel filters to obtain the energy of each filter.
[0160] Logarithmic processing: Apply logarithmic processing to the energy of each Mel filter to increase the sensitivity to low-energy signals.
[0161] Discrete Cosine Transform (DCT): Perform the discrete cosine transform on the logarithmic energy sequence to obtain the MFCC coefficients. The DCT transforms the signal from the time domain to the frequency domain, and usually only takes the first few coefficients because they contain most of the energy information.
[0162] Regularization and dimensionality reduction: Optional steps, perform regularization processing on the MFCC coefficients, such as normalization, and dimensionality reduction processing, such as principal component analysis (PCA), to reduce the computational complexity and improve the performance.
[0163] Feature vector: The finally obtained MFCC coefficients form a multi-dimensional feature vector that captures the time-frequency characteristics of the audio signal and can be used for tasks such as speech recognition and sentiment analysis.
[0164] The operation and maintenance server stores a daily inspection plan. The inspection plan, such as setting the inspection cycle, automatically triggers the inspection when the inspection time is reached. The operation and maintenance server starts the system timer and automatically starts the inspection process when the trigger condition is met.
[0165] In this embodiment, the design of the inspection plan includes two parts. One is the inspection time, which is used to specify when to conduct the inspection. The other is the inspection venue, which refers to which venues need to be inspected in this inspection. These two together constitute an inspection plan. When the time condition is met, the system automatically starts the inspection and includes the specified venues in the inspection one by one.
[0166] The operation and maintenance server stores an exception handling plan. When the operation and maintenance server detects an exception, it automatically executes the handling plan for disposal.
[0167] The system of the present invention supports corresponding responses and dispositions according to the detection parameters returned by the operation and maintenance terminal to automatically solve the problems found in the inspection. The automatic disposition plan will be automatically triggered according to the configured conditions when data is returned during the inspection. The disposition methods of the plan include all control instructions for the devices accessed by the operation and maintenance terminal. Example of the configuration method of the automatic disposition plan:
[0168] Example 1: When the operation and maintenance terminal detects that the remote conference terminal is offline through a command, it automatically sends a power-on instruction to the serial server to turn on the venue devices.
[0169] Example 2: When the operation and maintenance terminal detects that the volume of the venue is less than the specified threshold (e.g., 60 db), it sends a control instruction to the mixer to increase the audio output volume of the speakers.
[0170] This embodiment also discloses a method for intelligent operation and maintenance of remote conferences. By sequentially switching each venue to the main venue and conducting real remote conference simulations between the main venue and each branch venue after each switch, collecting the video and audio information of the main venue, as well as the video and audio information transmitted to each branch venue, and comparing and analyzing the video of the main venue with the video received by the branch venues, and the audio of the main venue with the audio received by the branch venues respectively, to determine whether there are abnormalities in the remote conference system.
[0171] This operation and maintenance system can automatically conduct inspections according to the preset strategy without manual intervention.
[0172] During the inspection process, problems existing in the system can be discovered in advance, including whether the sound and image are normal, whether the device control is normal, etc.
[0173] When a problem is discovered, it automatically reminds the relevant person in charge to handle it according to the preset strategy, and at the same time records the handling result and timeliness.
[0174] Statistical analysis is carried out in multiple aspects such as the failure rate and problem handling efficiency based on the collected data, specifically including:
[0175] The operation and maintenance system of the present invention will record the detailed data of each inspection into the system. The system will incorporate the historical data into a neural network for trend analysis for neural network model training at regular intervals (usually one week) to upgrade and iterate the neural network. And this neural network will be directly used for trend analysis of the inspection parameters to obtain the possible future detection result values, and display these values to the user for reference. The steps for training the trend analysis neural network model are:
[0176] Ensure the quality and integrity of the data, clean the data, and remove noise and outliers.
[0177] Perform data standardization or normalization to bring the data to the same scale. Usually, the data is scaled to a specific range, most commonly the interval [0, 1].
[0178] Supplement data features, including collection time, meeting room parameters, environmental parameters, spatial parameters, etc.
[0179] Divide the data into a training set, a validation set, and a test set. Usually, the ratio is 70% for the training set, 15% for the validation set, and 15% for the test set.
[0180] Construct a neural network model. See the following text for the specific model structure.
[0181] Train the neural network using the training set data. Adjust hyperparameters such as the learning rate, batch size, number of iterations, etc.
[0182] Use the validation set to optimize the model. Use the test set to evaluate the model performance.
[0183] Check metrics such as the accuracy, recall, F1-score, etc. of the model.
[0184] Deploy the trained model to the production environment.
[0185] Ensure that the model can receive real-time data and give predictions.
[0186] Method for designing the neural network structure for trend analysis:
[0187] Input layer: This layer receives the input data, and the number of nodes is the same as the number of data features.
[0188] Convolutional layer (optional): Use a 1D convolutional layer to extract local features in the time series data. Set multiple convolutional kernels, and each kernel is responsible for extracting different features.
[0189] Pooling layer (optional): After the convolutional layer, use a pooling layer to reduce the spatial dimension of the features while increasing the invariance to data displacement.
[0190] Recurrent layer: Use an LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) layer to handle the long-term dependencies in the data.
[0191] Fully connected layer: After the recurrent layer, include multiple fully connected layers (Dense layers) for further feature learning and decision-making.
[0192] Dropout layer (optional): Add a Dropout layer between or after the fully connected layers to reduce the risk of model overfitting.
[0193] Batch normalization layer (optional): After the fully connected layer, a batch normalization layer can be added to accelerate the training process and improve the stability of the model.
[0194] Output layer: The output layer has only one node and uses a linear activation function to output the prediction result.
[0195] Operation warning: In addition to giving a warning when the inspection return parameters reach the set threshold during the inspection, the system will also give a warning when analyzing the trends of various parameters. The warning methods include displaying a warning icon on the interface, sending an in-station message to the administrator, and sending a mobile phone text message to the administrator through a third-party SMS interface for notification.
[0196] The preferred embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. Typical well-known structures and common general knowledge technologies in the preferred embodiments are not described in detail here. Those of ordinary skill in the art can, under the inspiration given by this embodiment, complete and implement the technical solution of the present invention in combination with their own capabilities. Some typical well-known structures, well-known methods or common general knowledge technologies should not be an obstacle for those of ordinary skill in the art to implement the present application.
[0197] The scope of protection required by the present application shall be subject to the content of its claims, and the content recorded in the description of the invention, the specific implementation manners and the drawings of the specification is used to interpret the claims.
[0198] Within the scope of the technical concept of the present application, several modifications can also be made to the specific implementation manners of the present application, and these modified specific implementation manners should also be regarded as within the scope of protection of the present application.
Claims
1. The remote conference intelligent operation and maintenance system is applied to the remote conference system, which is characterized by: The invention comprises an operation and maintenance subsystem, wherein the operation and maintenance subsystem and the remote conference system are controlled cooperatively, and each venue is switched to the main venue in turn to ensure that two-way communication is carried out between each two venues, and after each switch, a real remote conference simulation is carried out between the main venue and each branch venue, and the video and audio information of the main venue are collected, and the video and audio information transmitted to each branch venue are collected, and the video of the main venue is compared with the video received by the branch venue, and the audio of the main venue is compared with the audio received by the branch venue, so as to judge whether there is an abnormality in the remote conference system. The remote conference system comprises an MCU control unit and conference terminals, display devices, cameras, and speakers arranged in each venue. The operation and maintenance subsystem It includes an operation and maintenance server, an operation and maintenance terminal, an operation and maintenance display device and an operation and maintenance camera. The operation and maintenance terminal is deployed in each venue and is responsible for docking with various peripherals of the operation and maintenance system, reading peripheral data and controlling peripherals. The operation and maintenance terminal is connected to the operation and maintenance server, accepts control instructions from the operation and maintenance server and uploads data to the operation and maintenance server. The operation and maintenance server sends operation and maintenance instructions to the operation and maintenance terminal and receives data sent back by the operation and maintenance terminal. In the remote conference simulation, the operation and maintenance server sends a setting signal to the MCU control unit, and the main venue is set to broadcast mode through the MCU control unit. The operation and maintenance server sends a control signal to the MCU control unit, and the MCU control unit establishes a remote conference connection and calls the conference terminals of all venues into the conference. The operation and maintenance server sends a control signal to the operation and maintenance terminal of the current main venue, and the operation and maintenance terminal controls the operation and maintenance display device to play the first video according to the control signal; The operation and maintenance terminal at the main venue controls the operation and maintenance camera to capture the first video, and transmits the captured first video to the MCU control unit, which transmits the first video to the conference terminals at other branch venues through the MCU control unit; displays the second video through the display devices at each branch venue, and emits the sound in the second video through the speakers at each branch venue; The operation and maintenance terminal located at each branch venue controls the camera of each branch venue to collect the second video, and transmits the collected second video to the operation and maintenance server through the operation and maintenance terminal; The first video captured by the operation and maintenance camera is also transmitted to the operation and maintenance terminal, and the operation and maintenance terminal transmits the first video to the operation and maintenance server, and the operation and maintenance server is used to detect the pictures in the first video and the second video, and compare whether the picture contents are consistent, and if they are inconsistent, an abnormal signal is issued; Switch to the main venue and perform the test again using the same method to check whether the video reverse propagation is normal.
2. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: The remote conference system also includes a microphone, and the operation and maintenance subsystem also includes a sound-generating device and a sound-collecting device. In the remote conference simulation: the operation and maintenance terminal is also used to control the sound device at the main venue to make sound; The microphone is used to amplify the sound; The sound pickup device at the main venue is used to collect the amplified sound and transmit the collected first audio information to the operation and maintenance terminal and the MCU control unit respectively; The operation and maintenance terminal transmits the first audio information to the operation and maintenance server; The MCU control unit is further used to transmit the first audio information to the conference terminals of each branch venue, and the speakers of each branch venue are used to play the second audio information formed after the conference terminals receive the first audio information. The sound pickup devices of each branch venue are used to collect the second audio information played in the speakers and transmit it to the operation and maintenance terminal, and the operation and maintenance terminal transmits the second audio information to the operation and maintenance server. The operation and maintenance server detects the first audio information and the second audio information, and compares whether they are consistent, and if they are inconsistent, issues an abnormal signal; Switch to the main venue and perform the test again using the same method to check whether the audio reverse propagation is normal.
3. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance server stores a daily inspection plan. The operation and maintenance server starts a system timer, and automatically starts the inspection process plan when a trigger condition is met.
4. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance server detects the IP addresses of conference terminals in all the venues to be inspected. If any conference terminal is found to be offline, a command is sent to the operation and maintenance terminal. The operation and maintenance terminal controls the serial port server of the venue to power on and automatically starts all the equipment in the venue including the remote conference terminal. After an operation and maintenance inspection is completed, a command is sent to the operation and maintenance terminal. The operation and maintenance terminal sends a shutdown command to the device according to the control protocol of each device to shut it down, and then controls the serial port server to power off in a preset order to shut down all the equipment in the venue including the remote conference terminal.
5. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance server stores an exception handling plan. When the operation and maintenance server detects an exception, the exception handling plan is automatically triggered according to the configured conditions. The exception handling plan includes restarting the device, adjusting the volume of the sound device and the sound pickup device, and adjusting the camera. The operation and maintenance server has a self-learning function, which performs self-learning according to each detected exception and the exception handling plan, so as to automatically generate a new handling plan when a similar exception is detected.
6. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance server detects the pictures in the first video and the second video in the same way. The principle is to detect the image content of the two pictures and the relative positions of the contents in the pictures respectively, and then compare the detection results of the two pictures. If the contents of the two pictures and the relative positions of the contents are roughly the same, they are considered to be pictures in the same scene.
7. The remote conference intelligent operation and maintenance system according to claim 2, characterized in that: Also includes sound content detection, The steps for sound content detection are as follows: Sound signal collection: the operation and maintenance terminal captures the sound signal in real time through the sound pickup device; Signal preprocessing: noise suppression and echo elimination of captured sound signals to improve the quality of subsequent processing; Audio signal segmentation: cut the sound signal into segments of equal length according to the audio playback duration; Unify the sampling rate and convert the sampling rate of all audio segments to a unified standard sampling rate; Feature extraction, using Mel-frequency cepstral coefficients (MFCC) to extract features from each audio frame; Feature data standardization: standardize the extracted MFCC feature data to facilitate the input of the neural network; The frame length is unified. If the audio frame length is less than or exceeds the standard length, it is padded or truncated to ensure that all input data lengths are consistent; Neural network input, passing the standardized feature data as input to the neural network; Convolutional layer processing: The convolutional layer in the neural network performs multiple convolution calculations on the input audio features to extract deeper features; Pooling layer processing, using pooling layer after convolution layer to reduce feature dimension and extract the most important features; Fully connected layer processing, the feature data extracted by the convolution layer and the pooling layer are passed to the fully connected layer for further feature fusion and classification decision; Application of activation function: Use activation function in fully connected layers to introduce nonlinearity; Output layer mapping, applying the Sigmoid function after the last fully connected layer to map the output to a probability value between 0 and 1; Threshold judgment: Set the threshold and convert the output of the Sigmoid function into a binary judgment to determine whether the audio content is consistent with expectations; Result output: outputs the detection results, including the similarity score and whether the target audio content is detected.
8. The remote conference intelligent operation and maintenance system according to claim 1, characterized in that: It also includes an operation and maintenance trend analysis module, which is used to: record detailed data of each inspection in the system. The system incorporates historical data into a neural network for trend analysis at regular intervals to train the neural network model, which is used to upgrade and iterate the neural network. The neural network is directly used for trend analysis of inspection parameters to derive possible values of detection results in the future, and display these values to users for reference. When performing trend analysis on each parameter, it will issue an early warning when the early warning conditions are met.
9. The remote conference intelligent operation and maintenance system according to claim 8, characterized in that: The training steps of the trend analysis neural network model are: Ensure data quality and integrity, clean data, remove noise and outliers; Standardize or normalize the data so that the data are on the same scale; Supplementary data features, including collection time, conference room parameters, environmental parameters, and space parameters; Divide the data into training, validation and test sets; Constructing a neural network model for trend analysis; Use the training set data to train the neural network and adjust the hyperparameters; Use the validation set to fine-tune the model and the test set to evaluate the model performance.
10. The remote conference intelligent operation and maintenance system according to claim 9, characterized in that: Trend analysis neural network models include: Input layer: This layer receives input data, and the number of nodes is consistent with the number of data features; Convolutional layer: Use 1D convolutional layer to extract local features in time series data. Set multiple convolution kernels, each of which is responsible for extracting different features. Pooling layer: After the convolutional layer, the pooling layer is used to reduce the spatial dimension of the features while increasing the invariance to data displacement; Recurrent layers: Use LSTM or GRU layers to handle long-term dependencies in the data; Fully connected layers: After the recurrent layers, multiple fully connected layers are included for further feature learning and decision making; Dropout layer: Add a Dropout layer between or after the fully connected layers to reduce the risk of overfitting the model; Batch Normalization Layer: After the fully connected layer, a batch normalization layer is added to speed up the training process and improve the stability of the model; Output layer: The output layer has only one node and uses a linear activation function to output the prediction results.
Citation Information
Patent Citations
Remote operation and maintenance video conference system based on image recognition
CN114363552A
AI automatic test method and system based on multi-user video conference system
CN115623187A
Multi-party conference test method and device, equipment and storage medium
CN116886890A
Cited By
Remote conference operation and maintenance equipment
CN223194767U