Smart power plant cooperative consultation communication method and system based on audio and video data compression

Through cross-modal feature fusion and dynamic encoding, traditional audio and video transmission is solved, and the problems of low efficiency and high redundancy in smart power plants are realized, high-fidelity and low-latency diagnostic data stream transmission is realized, and intelligent operation and maintenance of smart power plants is supported.

CN120475154APending Publication Date: 2025-08-12润电能源科学技术有限公司
View PDF 0 Cites -1 Cited by

Patent Information

Application Number
CN202510593738.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-12

Smart Images

  • Figure CN120475154A_ABST
    Figure CN120475154A_ABST
Patent Text Reader

Abstract

The invention discloses a smart power plant cooperative consultation communication method and system based on audio and video data compression, and relates to the technical field of smart power plant remote cooperative diagnosis, and the method comprises the steps: taking an equipment vibration feature as a core drive, fusing audio and video signal features, obtaining multi-modal data, and carrying out the preprocessing; performing cross-modal feature joint modeling according to the preprocessed multi-modal data, and then performing audio and video dynamic coding and compression; based on the encoded and compressed data, multi-modal code rate collaborative allocation is carried out, and after standardized data output is carried out through a hardware acceleration and fault-tolerant processing mechanism, cloud collaborative diagnosis is executed. According to the invention, enhanced compression and noise separation of key features are realized; meanwhile, code rate distribution is dynamically optimized in combination with diagnosis requirements, the fidelity of key information is further improved on the basis of guaranteeing low-delay transmission, diagnosis data streams with higher precision are provided for remote experts, and the intelligent operation and maintenance requirements of a power plant are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote collaborative diagnosis of smart power plants, and in particular to a collaborative consultation communication method and system for smart power plants based on audio and video data compression. Background Art

[0002] With the advancement of smart power plant construction, remote collaborative diagnosis of equipment status and expert consultation systems have become crucial for ensuring safe and efficient power plant operation. The quality of real-time transmission of on-site audio and video data directly impacts the accuracy of remote experts' fault diagnosis and the efficiency of emergency response. However, the complex electromagnetic environment, high noise levels, and limited network bandwidth of power plants present challenges for traditional audio and video transmission solutions, and existing technical solutions still have room for improvement. Typical video coding systems improve efficiency by optimizing motion search algorithms, but they do not fully integrate the vibration spectrum characteristics of the equipment. In dynamic vibration scenarios, macroblock partitioning strategies may cause bitrate fluctuations. While audio coding schemes can reduce steady-state noise, due to a lack of correlation modeling with video signals, the sound of equipment movement and visual information are encoded independently, resulting in a high level of data redundancy. Furthermore, the independent processing architecture for audio and video signals lacks a comprehensive spatiotemporal correlation model, which may lead to a certain degree of synchronization error, affecting the spatiotemporal consistency analysis of fault scenarios.

[0003] These issues primarily stem from the design limitations of traditional compression technologies. Existing audio noise reduction and video compression algorithms fail to fully leverage prior knowledge of the device's physical state, potentially losing key features during the encoding process. Furthermore, the dynamic noise spectrum and vibration interference found in industrial scenarios are not effectively modeled. Fixed bitrate allocation strategies struggle to dynamically adjust quantization parameters based on bandwidth fluctuations and diagnostic priorities, potentially compromising signal fidelity under limited bandwidth conditions.

[0004] Existing technical solutions have limitations in audio and video transmission scenarios in smart power plants:

[0005] Traditional video encoding methods rely on temporal motion estimation models, which fail to fully incorporate the high-frequency vibration characteristics of the device. Under mechanical vibration interference, inter-frame prediction efficiency decreases, resulting in a decrease in compression rate and potential distortion of edge details, which affects the accuracy of visual recognition of surface defects. Regarding audio processing, while existing noise reduction algorithms can suppress steady-state background noise, using a fixed threshold strategy can weaken the characteristics of abnormal device soundprints, thereby reducing the discernibility of faulty audio features.

[0006] Existing technologies use an independent audio and video encoding architecture and lack a cross-modal signal correlation modeling mechanism. This may lead to repeated encoding of physical related information such as device action video and driving sound, increasing data redundancy. In addition, the fixed bit rate allocation strategy does not fully consider the priority differences in diagnostic scenarios. When bandwidth is limited, the fidelity of key information may be difficult to effectively guarantee. At the hardware implementation level, the computing resource utilization of traditional encoding schemes still has room for improvement. When processing high-resolution video, the hardware load may be high, resulting in longer-than-expected delays in key alarm data, affecting the real-time requirements of industrial applications. These problems have, to a certain extent, restricted the accuracy and response efficiency of remote diagnostic systems. Summary of the Invention

[0007] In order to solve the above problems, the purpose of the present invention is to provide a real-time communication technology for collaborative consultation of smart power plants based on audio and video data compression, aiming to provide remote experts with more accurate diagnostic data streams to support the intelligent operation and maintenance needs of power plants.

[0008] To achieve the above technical objectives, the present application provides a smart power plant collaborative consultation communication method based on audio and video data compression, comprising the following steps:

[0009] Driven by device vibration characteristics, it integrates audio and video signal features to obtain multimodal data and perform preprocessing.

[0010] Based on the pre-processed multimodal data, cross-modal feature joint modeling is performed, and then audio and video dynamic encoding and compression are performed;

[0011] Based on the encoded and compressed data, multi-modal code rate collaborative allocation is carried out, and after standardized data output through hardware acceleration and fault-tolerant processing mechanisms, cloud-based collaborative diagnosis is performed.

[0012] Preferably, when preprocessing the multimodal data, a three-level optimization strategy is used to preprocess the video data by gamma correction, non-local mean filtering and dynamic region of interest extraction;

[0013] Audio data preprocessing is performed through adaptive echo cancellation and preliminary spectral subtraction.

[0014] Preferably, when performing cross-modal feature joint modeling, the preprocessed audio signal is subjected to multi-layer feature analysis in sequence through short-time Fourier transform, Mel frequency extraction and differential feature calculation to obtain audio feature analysis results.

[0015] Preferably, when performing cross-modal feature joint modeling, adaptive segmentation, frequency domain transformation, energy gradient analysis and frequency domain motion vector field construction are sequentially performed, and a dynamic macroblock partitioning strategy is adopted to extract video features from the preprocessed video signal to obtain video feature analysis results.

[0016] Preferably, when performing cross-modal feature joint modeling, based on the audio feature analysis results and the video feature analysis results, cross-modal feature joint modeling is performed in sequence through feature dimension alignment, multi-head attention mechanism, redundant signal recognition and correlation feature enhancement.

[0017] Preferably, when performing dynamic encoding and compression of audio and video, audio dynamic encoding is performed by sequentially performing selective noise filtering, frequency domain gain compensation, improved linear prediction coding and dynamic bit rate control;

[0018] Video dynamic encoding is performed through frequency domain motion compensation, lightweight variational autoencoder, industrial scene adaptability enhancement, layered quantization strategy and hardware acceleration optimization.

[0019] Preferably, when performing dynamic encoding and compression of audio and video, compression parameters are dynamically adjusted according to device status, diagnostic requirements and network conditions, and the transmission quality of key information is diagnosed with priority under limited bandwidth.

[0020] Preferably, when performing multimodal bitrate collaborative allocation, a priority weight model is established through device weight mapping, scenario priority evaluation and bandwidth allocation strategy in sequence; and a graded degradation strategy is designed for bandwidth fluctuations. At the same time, through network status feedback, emergency handling, elastic scheduling mechanism and quality assurance strategy in sequence, the network status and diagnosis needs are continuously monitored and parameters are dynamically optimized to complete multimodal bitrate collaborative allocation.

[0021] Preferably, when hardware acceleration and fault-tolerant processing mechanisms are used to output standardized data, an FPGA+DSP heterogeneous computing architecture is adopted. After setting up a parallel processing flow and a fault-tolerant processing mechanism, standardized data output is performed in sequence through audio and video encapsulation, timestamp binding, and device status association.

[0022] The present invention also discloses a smart power plant collaborative consultation communication system based on audio and video data compression, which is used to implement the above-mentioned smart power plant collaborative consultation communication method based on audio and video data compression, including:

[0023] The data acquisition and preprocessing module is used to acquire and preprocess multimodal data based on the vibration characteristics of the equipment, integrating the characteristics of audio and video signals;

[0024] The data processing module is used to perform cross-modal feature joint modeling based on the pre-processed multimodal data, and then perform dynamic encoding and compression of audio and video;

[0025] The data output module is used to perform multi-modal bit rate coordination based on the encoded and compressed data, and to output standardized data through hardware acceleration and fault-tolerant processing mechanisms;

[0026] The collaborative diagnosis module is used to obtain standardized data and perform cloud-based collaborative diagnosis through multi-party consultation, interactive annotation, diagnosis report generation and historical data comparison.

[0027] The present invention discloses the following technical effects:

[0028] 1. Cross-modal high-fidelity compression: By integrating audio spectrum, video motion features, and device vibration data through a cross-modal attention mechanism, this technology accurately preserves key diagnostic information in scenarios with strong noise and high-frequency vibration. Compared to traditional independent encoding schemes, this effectively eliminates redundant encoding of device motion video and drive sound, while also preventing excessive attenuation of abnormal soundprint features (such as bearing friction) during the noise reduction process, significantly improving the discernibility of fault features and the fidelity of visual details.

[0029] 2. Dynamic Scenario Adaptive Compression: Based on a frequency-domain motion compensation algorithm and a dynamic noise spectrum learning model, encoding parameters are dynamically adjusted based on the vibration characteristics of power plant equipment and the industrial noise spectrum. A priority-aware bitrate allocation strategy prioritizes the encoding quality of critical information such as instrument readings and abnormal voiceprints when bandwidth is limited, avoiding the degradation of diagnostic information caused by traditional equal compression strategies.

[0030] 3. Lightweight and efficient encoding: Adopting a hardware-aware parallel acceleration architecture, the computational efficiency of the frequency domain transformation and feature extraction modules is optimized, reducing the hardware resource utilization of high-resolution video encoding to less than 30% of traditional solutions. This supports low-latency real-time compression on embedded devices, meeting the stringent real-time and reliability requirements of industrial sites. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 This is a schematic diagram of the architecture of the smart power plant collaborative consultation communication method system based on audio and video data compression according to the present invention;

[0033] Figure 2 It is a schematic flow chart of the method described in the present invention. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0035] like Figure 1-Figure 2 As shown, the present invention provides a smart power plant collaborative consultation communication method based on audio and video data compression. This method integrates audio, video, and equipment vibration data, builds a cross-modal attention mechanism and a dynamic noise model, and achieves enhanced compression of key features and noise separation. At the same time, it dynamically optimizes bit rate allocation based on diagnostic needs, further improving the fidelity of key information while ensuring low-latency transmission, providing remote experts with a more accurate diagnostic data stream, and supporting the intelligent operation and maintenance needs of power plants. Specifically, it includes the following:

[0036] This paper proposes a multimodal joint compression method to address the audio and video data compression needs in the complex industrial environments of smart power plants. This method, driven by equipment vibration characteristics, integrates audio and video signal features to build a complete real-time communication system, addressing data transmission issues under high noise and vibration interference. The system architecture is divided into four major modules: multimodal data acquisition and preprocessing, cross-modal feature joint modeling, dynamic audio and video coding and compression, and multimodal bit rate collaborative allocation.

[0037] 1. Multimodal data acquisition and preprocessing:

[0038] 1.1 Data Collection

[0039] Deploy sensing terminals on-site at power plant equipment to simultaneously collect multimodal data:

[0040] Industrial camera: Captures 1080p@30fps or 4K@15fps video streams, features a dust-proof and waterproof housing, and is temperature-resistant from -20°C to 80°C.

[0041] Directional microphone array: uses a 6-channel microphone with a sampling rate of 96kHz and a frequency response range of 20Hz-45kHz, covering the characteristic frequency band of abnormal sound patterns of the device;

[0042] Three-axis vibration sensor: sampling rate 1kHz, measurement range ±16g, frequency response 0.5Hz-5kHz, real-time acquisition of equipment vibration spectrum;

[0043] Precision clock synchronization module: realizes timestamp synchronization of multi-source data based on the precise time protocol, with synchronization accuracy ≤1ms;

[0044] The acquisition system utilizes a modular design, allowing flexible configuration of sensor quantities and parameters based on different device types. The signal acquisition terminal features a local cache, storing 30 minutes of high-quality raw data during network instability and automatically retransmitting it upon network recovery, ensuring the continuity and integrity of diagnostic data. Furthermore, all sensors utilize an industrial-grade design, ensuring stable operation in harsh environments such as high temperature, high humidity, and strong electromagnetic interference, significantly reducing maintenance costs.

[0045] 1.2 Preprocessing process:

[0046] Video preprocessing adopts a three-level optimization strategy:

[0047] Gamma correction: Apply the inverse gamma function to correct grayscale distortion caused by uneven industrial lighting, with a γ value range of 1.8-2.2;

[0048] Non-local mean filtering: An improved non-local mean filtering algorithm is used with a search window size of 13×13, a block size of 3×3, and a filter strength parameter of h=10 to jointly remove salt and pepper noise and Gaussian noise.

[0049] Dynamic region of interest extraction: Detect motion regions based on background difference models, use morphological operations to optimize the boundaries of image regions of interest, and apply downsampling preprocessing to non-critical areas;

[0050] Audio preprocessing includes two key steps:

[0051] Adaptive echo cancellation: uses a normalized least mean square algorithm (step size parameter μ = 0.05) to estimate the acoustic feedback path in real time, with a filter order of 256-512 (dynamically adjusted according to the ambient reverberation time) to eliminate microphone self-oscillation;

[0052] Preliminary spectral subtraction: The audio signal is framed (20ms frame length, 10ms frame shift), converted to the frequency domain using piecewise fast Fourier transform, and a modified spectral subtraction method (oversubtraction factor α = 1.5, noise floor β = 0.02) is applied to dynamically suppress steady-state noise.

[0053] The preprocessed data is output in a standardized format:

[0054] Video stream: YUV420 format, resolution adaptively adjusted according to network status (4K / 1080p / 720p);

[0055] Audio stream: PCM format, 24-bit depth, retaining the original 6-channel beam data or synthesizing 16kHz sampling rate mono;

[0056] Vibration data: CSV format, records three-axis (X / Y / Z) acceleration values, 1ms sampling interval;

[0057] 2. Joint modeling of cross-modal features:

[0058] Traditional audio and video processing methods typically treat signals independently, ignoring their physical connections, resulting in information oscillation and feature loss. This paper establishes a joint feature model for audio, video, and device vibration data. By exploring the inherent connections between multimodal signals, it achieves complementary information enhancement and provides a theoretical foundation for compression layering.

[0059] 2.1 Audio feature extraction:

[0060] The pre-processed audio signal undergoes multi-layer feature analysis:

[0061] Short-time Fourier transform: 20ms frame length, 10ms frame shift (50% overlap), Hanning window function (α = 0.5) is applied to suppress spectrum leakage, and 512-point FFT is used to generate a high-resolution time-frequency spectrum matrix;

[0062] Mel frequency extraction: A Mel filter bank consisting of 40 triangular bandpass filters is used to map the linear spectrum to the perceptually relevant Mel domain, extracting 13th-order Mel frequency coefficients to highlight the soundprint characteristics of mechanical faults such as gear meshing and bearing friction.

[0063] Differential feature calculation: The first-order difference window length is 5 frames, extracting Δ Mel frequency features. The second-order difference window length is 3 frames, extracting Δ Mel frequency features, dynamically capturing the time domain change rate of acoustic features, and enhancing the ability to detect transient impact noise.

[0064] The audio feature extraction process is specifically optimized for the sound characteristics of industrial equipment, effectively distinguishing normal operating sounds from abnormal sound patterns. By analyzing Mel-frequency cepstral coefficients, the system can capture subtle acoustic changes that are imperceptible to the human ear, such as the high-frequency impact sound produced by early-stage bearing damage. Furthermore, feature calculation further enhances sensitivity to dynamic changes in sound, enabling the system to accurately identify transient abnormal events even in strong environments.

[0065] 2.2 Video feature extraction:

[0066] The video frame adopts dynamic macroblock division strategy:

[0067] Adaptive segmentation: Dynamically adjusts the macroblock size based on the input spectrum characteristics of the vibration sensor. Low-frequency vibration scenarios use 32×32 macroblocks, while high-frequency vibration scenarios automatically switch to 8×8 macroblocks.

[0068] Frequency domain transform: 8×8 discrete cosine transform is applied to each macroblock to generate a frequency domain coefficient matrix;

[0069] Energy gradient analysis: Calculate the energy distribution and gradient of discrete cosine transform coefficients and extract the main frequency components in the horizontal, vertical and diagonal directions;

[0070] Frequency domain motion vector field construction: Combined with the three-axis vibration spectrum, a frequency domain motion prediction model is established, improving accuracy by 15-20%;

[0071] This video extraction and parsing method breaks through the constraints of traditional fixed macroblock partitioning. Through an adaptive segmentation strategy driven by vibration data, it provides an optimal coding structure for devices with different vibration frequencies. The frequency-domain motion vector field, combined with actual physical vibration parameters, enables more accurate motion prediction, significantly improving compression efficiency.

[0072] 2.3 Cross-modal feature fusion:

[0073] Deep correlation modeling based on extracted features:

[0074] Feature dimension alignment: The audio Mel frequency feature sequence is reduced to the same temporal length as the video discrete cosine transform feature through one-dimensional convolution (kernel size 5, stride 2). The video frequency domain energy features are mapped to a time series vector through spatial average pooling (4×4). The vibration spectrum is extracted using a two-layer LSTM network (hidden dimension 128, dropout rate 0.2) to extract temporal dynamic features.

[0075] Multi-head attention mechanism: Build an 8-head self-attention model (head dimension 64), use audio, video, and vibration features as query / key / value inputs, calculate the cross-modal association matrix, and dynamically assign attention weights;

[0076] Redundant signal identification: Dynamic redundant signal identification is implemented based on the attention weight matrix (threshold 0.7), strongly correlated signals are marked, and their overlapping features are jointly encoded;

[0077] Correlation feature enhancement: The high-weight region feature enhancement ratio is 1.2-1.5, generating the noise mask matrix and frequency domain motion compensation parameters to guide subsequent hierarchical compression decisions;

[0078] Cross-modal feature fusion is the core innovation of this invention. It breaks down one of the information silos in traditional audio and video processing and establishes a unified feature space based on physical connections. Through a multi-head attention mechanism, the system automatically discovers the inherent connections between data from different modalities, such as the proximity of device vibration to images and videos, and the mechanical modeling of these connections not only reduces the data model but also improves information variance and reliability.

[0079] 3. Dynamic encoding and compression of audio and video:

[0080] Based on a cross-modal feature model, this invention implements a module-based adaptive encoding and compression strategy tailored to specific industrial scenarios. Unlike traditional fixed-parameter encoding, this approach dynamically adjusts compression parameters based on device status, diagnostic requirements, and network conditions, prioritizing the transmission quality of critical information within limited bandwidth.

[0081] 3.1, Audio dynamic encoding:

[0082] Based on the cross-modal feature fusion results, a dynamic audio coding strategy is designed:

[0083] Selective noise filtering: Using a noise mask generated by cross-modal attention, frequency-domain selective filtering is performed on industrial background noise, with a filtering depth of 20-30dB and dynamic adjustment of the frequency band suppression threshold.

[0084] Frequency domain gain compensation: Provides protective enhancement for key frequency bands of abnormal device sound patterns, with the compensation amount strictly controlled within 3dB to avoid attenuation of key fault characteristics.

[0085] Improved linear predictive coding: The prediction order is 12-16, and the adaptive encoder dynamically allocates the bit rate based on the voiceprint feature matching degree, increasing the bit rate of key fault frequency bands by 20-30%;

[0086] Dynamic bitrate control: Dynamically adjusts output bitrate based on network bandwidth and diagnostic priority, with end-to-end processing latency less than 100ms;

[0087] Dynamic audio coding addresses the challenge of preserving the fidelity of complex acoustic features in industrial environments. By identifying key correlates of abnormal sound patterns and enhancing their protection, the system can preserve the acoustic signatures of equipment failures even in high-noise environments. Dynamic parameter adjustment in linear predictive coding enables the compression algorithm to adapt to different types of acoustic events, such as persistent noise and transient impact sounds. Compared to traditional fixed-parameter coding, this method improves the recognition rate of abnormal sound patterns within the same bandwidth, enabling remote experts to more clearly hear abnormal equipment sounds and assist in fault diagnosis.

[0088] 3.2 Video dynamic encoding:

[0089] Optimizing video compression performance for high-frequency vibration scenarios:

[0090] Frequency domain motion compensation: A frequency domain motion model is established using vibration sensor input, with a prediction coefficient confidence threshold of 0.85 and a compensation accuracy error of <2%;

[0091] Lightweight variational autoencoder: 4-level adjustable downsampling layer (64-128-256-512 filters) and residual enhancement module, latent space dimension 256, regularization coefficient β = 0.01;

[0092] Enhanced adaptability to industrial scenarios: Introducing a priori knowledge base of device textures and improving encoder adaptability through adversarial training (discriminator learning rate 0.0002);

[0093] Hierarchical quantization strategy: key areas use low quantization parameter values (18-22) for fine quantization, and secondary background areas use high quantization parameter values (28-32) for coarse quantization to achieve differentiated encoding;

[0094] Hardware acceleration optimization: FPGA acceleration unit is used to achieve parallel computing of discrete cosine transform and variational autoencoder inference, with a processing speed of ≥30fps@1080p and a peak signal-to-noise ratio (PSNR) ≥40dB.

[0095] Video dynamic encoding overcomes the performance bottleneck of traditional encoders in high-frequency vibration scenarios. Frequency-domain motion compensation technology directly utilizes physical domain data sensors to guide motion prediction, significantly improving encoding efficiency. A lightweight variational autoencoder, combined with industry prior knowledge, enables the system to "understand" the surface texture characteristics of surface devices, preserving key details during the compression process. A layered optimization strategy allocates differentiated encoding resources to different intensity regions, ensuring that critical information such as instrument readings and device status indicators can be read within specified bandwidth conditions. An optimized hardware acceleration solution enables the system to achieve real-time encoding on edge devices, meeting the low-latency requirements of remote diagnostics.

[0096] 4. Multimodal code rate collaborative allocation:

[0097] In practical application scenarios, network bandwidth is often the bottleneck resource for remote diagnosis. This invention achieves intelligent allocation of bandwidth resources between different signal sources and content areas by establishing a priority model based on diagnostic requirements, ensuring that the transmission quality of core diagnostic information can be guaranteed under bandwidth conditions.

[0098] 4.1. Priority model construction:

[0099] Establish a priority weight model based on device type and diagnostic scenario:

[0100] Equipment weight mapping: key equipment such as generators and steam turbines are given high weights (0.8-1.0), and auxiliary system equipment is given medium and low weights (0.3-0.7);

[0101] Scenario priority assessment: Alarm state weight is 1.0, abnormal state weight is 0.8, and normal state weight is 0.5. Quantification parameters are adjusted dynamically.

[0102] Bandwidth allocation strategy: 70-80% of available bandwidth is allocated to high-priority information, 10-20% to medium-priority information, and the remainder to low-priority information;

[0103] The priority model is designed based on the knowledge base of military operations and maintenance experts and historical diagnostic analysis. The system comes with pre-configured priority configurations for over 50 common fault scenarios and supports customization based on actual application needs. This model not only considers the criticality of the equipment but also incorporates multiple factors such as fault type and severity, providing a scientific basis for resource allocation.

[0104] 4.2 Adaptive Bitrate Adjustment

[0105] Design a hierarchical degradation strategy to address bandwidth fluctuations:

[0106] Dynamic video stream adaptation: The resolution is gradually reduced from 4K to 720p, and intelligent area cropping technology is used to preserve key device areas;

[0107] Motion blur reduction: Apply a Wiener filter (window size 5×5) when reducing resolution to reduce detail loss;

[0108] Audio frequency band selection: When bandwidth is severely limited, the 300Hz-8kHz core frequency band is retained and non-essential high-frequency bands are disabled;

[0109] Band energy compensation: Enhances feature recognition by compressing the dynamic range (compression ratio 2:1, threshold -20dB);

[0110] An adaptive encoding rate adjustment strategy enables the system to smoothly detect network fluctuations, avoiding transmission interruptions or severe quality degradation. The upgrade and degradation strategy is based on human access characteristics and diagnostic priority, prioritizing the printability of critical information when resources are allocated. Unlike traditional approaches that simply reduce overall quality, this system implements quality adjustments for content acquisition, ensuring access to core diagnostic areas even under extremely low bandwidth conditions.

[0111] 4.3 Real-time monitoring and adjustment:

[0112] The system continuously monitors network status and diagnoses needs to dynamically optimize parameters:

[0113] Network status feedback: Based on the Real-time Transport Control Protocol, network status data is collected every 200ms to respond to bandwidth fluctuations in a timely manner;

[0114] Emergency handling: When a high-frequency vibration alarm is detected, the corresponding video stream bit rate is automatically increased, while the coding resources in the non-alarm area are reduced;

[0115] Flexible scheduling mechanism: By predicting bandwidth models, encoding parameters are adjusted in advance to reduce the impact of jitter;

[0116] Quality assurance strategy: In extreme bandwidth-constrained situations, priority is given to ensuring video quality and abnormal voiceprint integrity in the alarm area. The frame rate in non-critical areas can be reduced to 15fps.

[0117] The emergency priority handling mechanism ensures that the system can immediately adjust resource allocation at critical moments and input high-quality diagnostic information as soon as possible.

[0118] 5. Hardware acceleration and real-time output:

[0119] 5.1 Heterogeneous Hardware Architecture

[0120] The system adopts FPGA+DSP heterogeneous computing architecture:

[0121] FPGA video processing unit: Xilinx Artix-7 series, 50K-100K logic units, 200MHz frequency, responsible for video frequency domain transformation and VAE inference;

[0122] DSP audio processing unit: TIC6678 series, 8 cores, main frequency 1.0-1.2GHz, floating-point performance 16GFLOPS, responsible for FFT transformation, attention calculation and dynamic encoding;

[0123] Optimized resource allocation: Through a parallel computing architecture, it supports simultaneous processing of four 1080p video streams while keeping power consumption within 10-15W.

[0124] The multi-layer computing architecture leverages the strengths of various hardware components. The FPGA's sophisticated computation and fixed-point calculations are well-suited for modification and adjustment operations in video encoding. The DSP, with its strengths in floating-point calculations and complex algorithm implementation, is well-suited for audio processing and feature extraction. Through sophisticated hardware design, the system maintains processing performance while keeping overall power consumption within the range of industrial edge devices.

[0125] 5.2 Real-time processing pipeline:

[0126] Enable efficient parallel processing:

[0127] Pipeline design: acquisition → preprocessing → feature extraction → encoding → packaging, 5-stage pipeline design, processing delay <150ms;

[0128] Dynamic frequency adjustment: adjusts the FPGA clock frequency (100-300MHz) in real time according to video complexity, balancing encoding efficiency and energy consumption;

[0129] Task scheduling mechanism: Flexible scheduling strategy based on priority preemption to ensure real-time transmission of alarm information;

[0130] Processing simulation significantly improves system throughput through refined task grouping and threading. Latency balancing across all simulation stages ensures maximum resource utilization and avoids dynamic performance bottlenecks. Frequency scaling adjusts processor frequency based on actual load, meeting performance requirements while minimizing capacity.

[0131] 5.3. Fault-tolerant processing mechanism:

[0132] Improve system reliability:

[0133] Load monitoring: Real-time monitoring of hardware resource utilization, with a threshold of 80% automatically triggering task migration;

[0134] Backup switching: The core encoding unit is configured with redundant resources, and switching is completed within 50ms in the event of a failure;

[0135] Error recovery: A state recovery mechanism based on differential updates allows for rapid recovery from the most recent state after an abnormal interruption.

[0136] Degraded operation: Automatically switches to low-power mode under extreme conditions to ensure continuous operation of core functions;

[0137] The fault-tolerant mechanism provides the system with multi-layered reliability, enabling stable operation in challenging environments and under abnormal conditions. Load monitoring and task migration prevent system stalls caused by single-point performance bottlenecks. Iterative design and rapid state transitions ensure serviceability, while intermittent state recovery technology significantly enhances system recovery from crashes.

[0138] 6. Data output and collaborative diagnosis:

[0139] 6.1 Standardized Data Output

[0140] The compressed data is output in standard format:

[0141] Audio and video encapsulation: Uses WebM network video format, VP9 encoding for video, and Opus encoding for audio;

[0142] Timestamp binding: Accurate synchronization of multimodal data is achieved based on the network time protocol, with a timestamp accuracy of 1ms;

[0143] Device status association: Embed device status metadata in compressed data streams to achieve data association;

[0144] Standardized data output ensures system interoperability and compatibility. The WebM network video format, as an open standard, supports cross-platform playback and processing, allowing integration with existing enterprise systems. Precise time synchronization ensures temporal consistency of multimodal data, allowing experts to accurately correlate consistent waveform, sound, and vibration information. The embedded device status metadata enables the diagnostic platform to automatically correlate relevant parameters, providing experts with comprehensive diagnostic context.

[0145] 6.2. Collaborative diagnosis support:

[0146] Cloud diagnostic platform functions:

[0147] Multi-party consultation: Establish a point-to-point low-latency channel based on real-time web communication, supporting 30 people to conduct collaborative diagnosis online at the same time;

[0148] Interactive annotation: Experts can use the touch interface to annotate abnormal areas in real time, and the system automatically associates the audio features of the corresponding time points;

[0149] Diagnostic report generation: Multimodal data fusion analysis, automatic generation of time-space consistent diagnostic reports, support PDF / HTML export;

[0150] Historical data comparison: Automatically match current anomalies with historical cases, providing reference cases ranked by similarity;

[0151] A robust diagnostic platform transforms densely transmitted, high-quality data into practical diagnostic value. The multi-party consultation function breaks down geographical constraints, enabling experts in diverse locations to share on-site information in real time and provide multi-angle professional opinions. The case tagging tool supports experts in interpreting key points of interest, and the system intelligently correlates multimodal data to facilitate comprehensive analysis. Automatically generated diagnostic reports preserve key information from the consultation process for future reference and knowledge accumulation. The historical case comparison function provides experts with accumulated experience, accelerating problem resolution.

[0152] This invention effectively solves the problems of feature loss, compression inefficiency, and lack of real-time performance in traditional solutions in industrial environments through delayed fusion and joint compression of multimodal data. While ensuring high fidelity of key diagnostic information, the system significantly improves data redundancy and transmission load, meeting the stringent technical requirements of remote coordinated diagnosis in smart factories.

[0153] This invention addresses the need for remote collaborative diagnosis of smart power plant equipment and aims to address key technical bottlenecks in audio and video data compression in complex industrial environments. It proposes a high-fidelity, low-redundancy multimodal joint compression method. This method addresses shortcomings in traditional approaches, such as insufficient adaptability of video encoding to equipment vibration scenarios, loss of abnormal features due to audio noise reduction, and inefficient multimodal data collaboration. Through innovative cross-modal feature fusion and dynamic encoding mechanisms, this method achieves efficient compression of audio and video data while accurately retaining key diagnostic information.

[0154] The core goal of this invention is to build an audio and video compression system that is adaptive to industrial scenarios: through the joint modeling of audio spectrum, video motion characteristics and equipment vibration data, the data redundancy caused by single-modal independent encoding is eliminated; dynamic noise suppression and frequency-domain motion compensation algorithms are designed to maintain the integrity of key features under strong noise and high-frequency vibration interference; priority-aware bit rate allocation strategies are developed to dynamically optimize compression parameters according to diagnostic needs, providing remote experts with high-fidelity, spatiotemporally consistent multimodal diagnostic data streams.

[0155] In summary, the present invention has the following characteristics:

[0156] (1) Cross-modal industrial feature fusion mechanism: The equipment vibration spectrum is deeply integrated into the audio and video compression decision, the frequency domain motion compensation model driven by vibration data is used to optimize the video inter-frame prediction, and the cross-modal attention mechanism is combined to establish a joint coding space of audio spectrum, video texture and vibration features, effectively solving the problem of feature loss and redundant coding in traditional single-modal processing in complex industrial scenarios.

[0157] (2) Dynamic priority-aware coding strategy: Adaptive bitrate allocation technology based on diagnostic requirements and network status, defining a weight mapping model between key areas of the device and abnormal voiceprints, dynamically adjusting quantization parameters and sub-band bit allocation, and prioritizing the fidelity of core features when bandwidth fluctuates.

[0158] (3) Industrial-grade lightweight acceleration architecture: Design a hardware-aware parallel processing pipeline, accelerate frequency domain transformation and lightweight variational autoencoder inference through FPGA, and combine it with embedded DSP to achieve real-time audio processing, achieving higher encoding rates and lower power consumption on edge devices, meeting the real-time and reliability requirements in power plant environments.

[0159] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0160] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0161] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A smart power plant collaborative consultation communication method based on audio and video data compression, characterized in that: The following steps are involved: Driven by device vibration characteristics, it integrates audio and video signal features to obtain multimodal data and perform preprocessing. Based on the pre-processed multimodal data, cross-modal feature joint modeling is performed, and then audio and video dynamic encoding and compression are performed; Based on the encoded and compressed data, multi-modal code rate collaborative allocation is carried out, and after standardized data output through hardware acceleration and fault-tolerant processing mechanisms, cloud-based collaborative diagnosis is performed.

2. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 1 is characterized by: When preprocessing multimodal data, a three-level optimization strategy is used to preprocess the video data through gamma correction, non-local mean filtering and dynamic region of interest extraction; Audio data preprocessing is performed through adaptive echo cancellation and preliminary spectral subtraction.

3. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 2 is characterized by: When performing cross-modal feature joint modeling, the preprocessed audio signal is subjected to multi-layer feature analysis through short-time Fourier transform, Mel frequency extraction and differential feature calculation to obtain the audio feature analysis results.

4. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 3 is characterized by: When performing cross-modal feature joint modeling, adaptive segmentation, frequency domain transformation, energy gradient analysis and frequency domain motion vector field construction are carried out in sequence, and a dynamic macroblock partitioning strategy is adopted to extract video features from the preprocessed video signal to obtain video feature analysis results.

5. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 4 is characterized by: When performing cross-modal feature joint modeling, based on the audio feature analysis results and the video feature analysis results, cross-modal feature joint modeling is performed in sequence through feature dimension alignment, multi-head attention mechanism, redundant signal recognition and correlation feature enhancement.

6. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 5 is characterized by: When performing dynamic encoding and compression of audio and video, audio dynamic encoding is performed through selective noise filtering, frequency domain gain compensation, improved linear prediction coding and dynamic bit rate control; Video dynamic encoding is performed through frequency domain motion compensation, lightweight variational autoencoder, industrial scene adaptability enhancement, layered quantization strategy and hardware acceleration optimization.

7. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 6 is characterized by: When performing dynamic encoding and compression of audio and video, the compression parameters are dynamically adjusted according to the device status, diagnostic requirements and network conditions, and the transmission quality of key information is prioritized under limited bandwidth.

8. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 7 is characterized by: When performing multimodal bitrate collaborative allocation, a priority weight model is established through device weight mapping, scenario priority evaluation, and bandwidth allocation strategy. A graded degradation strategy is designed for bandwidth fluctuations. At the same time, network status feedback, emergency handling, flexible scheduling mechanism, and quality assurance strategy are used to continuously monitor network status and diagnosis needs and dynamically optimize parameters to complete multimodal bitrate collaborative allocation.

9. The smart power plant collaborative consultation communication method based on audio and video data compression according to claim 8 is characterized by: When hardware acceleration and fault-tolerant processing mechanisms are used to output standardized data, an FPGA+DSP heterogeneous computing architecture is adopted. After setting up parallel processing flows and fault-tolerant processing mechanisms, standardized data output is performed through audio and video encapsulation, timestamp binding, and device status association.

10. A smart power plant collaborative consultation communication system based on audio and video data compression, used to implement the smart power plant collaborative consultation communication method based on audio and video data compression according to any one of claims 1 to 9, characterized in that: include: The data acquisition and preprocessing module is used to acquire and preprocess multimodal data based on the vibration characteristics of the equipment, integrating the characteristics of audio and video signals; The data processing module is used to perform cross-modal feature joint modeling based on the pre-processed multimodal data, and then perform dynamic encoding and compression of audio and video; The data output module is used to perform multi-modal bit rate coordination based on the encoded and compressed data, and to output standardized data through hardware acceleration and fault-tolerant processing mechanisms; The collaborative diagnosis module is used to obtain standardized data and perform cloud-based collaborative diagnosis through multi-party consultation, interactive annotation, diagnosis report generation and historical data comparison.