An operating condition semantic cooperative driving multi-modal video intelligent transmission method for oil and gas well surface testing
By using condition-guided prediction and semantic substitution transmission in a closed-loop control system, the problems of abnormal monitoring and video integrity under low bandwidth conditions in oil and gas well surface testing sites were solved. This resulted in steady-state low load, abnormal low latency, and post-event traceability video transmission effects at oil and gas well surface testing sites.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENGDU SICHUAN OIL SAFETY TECHNOLOGY ENGINEERING CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-03
AI Technical Summary
Video monitoring at oil and gas well surface testing sites is difficult to simultaneously achieve anomaly monitoring, low-bandwidth operation, and the integrity of the original video under conditions of unstable wireless links, large bandwidth fluctuations, and limited power supply. Existing technologies have failed to effectively utilize operating data for prediction and coordinated transmission, resulting in missing key frames or excessive resource consumption before anomalies occur.
A closed-loop control method is adopted, which includes working condition-guided prediction, semantic substitution transmission, network-wide resource scheduling, and original video retrospective transmission. Video nodes and sensor nodes are uniformly managed through edge gateways, multimodal data acquisition and spatiotemporal alignment are performed, a working condition-video association model is established, transmission modes are adaptively switched, and pre-coding enhancement and network-wide bandwidth scheduling are performed at the edge to achieve intelligent video transmission.
It improves the integrity of video capture and post-event traceability for anomaly monitoring under low bandwidth conditions, reduces resource consumption, and ensures timely transmission of key frames and integrity of incident analysis.
Smart Images

Figure CN122340260A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of oil and gas well surface testing, industrial video surveillance, edge computing, and wireless intelligent transmission technology. Specifically, it relates to a multimodal video intelligent transmission method, system, electronic device, and computer-readable storage medium that utilizes on-site working condition data, video semantic information, and scene digital twin collaborative driving. Background Technology
[0002] Surface testing of oil and gas wells is a critical operational step in the exploration and development of oil and gas fields. It typically requires the deployment of video monitoring nodes and operational data acquisition nodes in areas such as wellheads, choke and kill manifolds, separators, metering devices, storage tanks, flare zones, and operational access routes to continuously monitor on-site production status, safety risks, and personnel operations.
[0003] Since the test sites are mostly located in remote areas in the wild, they generally suffer from problems such as limited power supply, unstable wireless links, large fluctuations in network bandwidth, and complex on-site environments. Traditional solutions that continuously upload full bitrate videos cannot balance bandwidth, latency, energy consumption, and evidence integrity.
[0004] Existing solutions typically only perform adaptive encoding from the video content side, or uniformly reduce resolution, frame rate, and bit rate when the network is congested. However, they have not fully utilized the characteristics of working condition data such as pressure, temperature, flow rate, liquid level, vibration, and combustible gas concentration as a precursor to changes in video. Therefore, they are prone to missing critical transition frames before anomalies occur.
[0005] On the other hand, oil and gas well testing is in steady-state operation for most of the time. If a complete video stream is continuously transmitted, it will consume a large amount of wireless bandwidth and node power, and also squeeze the transmission resources of alarm information, process data and control commands.
[0006] Furthermore, when multiple regions experience anomalies simultaneously, if each node independently competes for link resources, it is often impossible to prioritize video uploads from high-risk regions. On the other hand, if a low-bandwidth solution only reduces the upload frequency, it may make it difficult to trace back the original video before the anomaly occurred, thereby affecting incident analysis and liability determination.
[0007] Therefore, a new intelligent video transmission solution is needed for oil and gas well surface testing sites to establish a collaborative closed loop between low-bandwidth operation, real-time monitoring of anomalies, and the traceability of original video. Summary of the Invention
[0008] The technical problem to be solved by this invention is to provide a multimodal intelligent video transmission method driven by working condition semantic collaboration for oil and gas well surface testing. This method is not simply about optimizing encoding or simply compressing video, but rather integrates working condition prediction, semantic substitution transmission, network-wide resource scheduling, and original video backtracking and retransmission into a single closed-loop control link.
[0009] To achieve the above objectives, the present invention adopts the following technical solution. For the overall system architecture, please refer to the appendix. Figure 1 As shown, the system of this invention includes video nodes, sensor nodes, an edge gateway, and a remote monitoring center. The video nodes are responsible for video acquisition, edge detection, encoding control, and local loop caching; the sensor nodes are responsible for acquiring operating parameters such as pressure, temperature, flow rate, liquid level, vibration, and combustible gas concentration; the edge gateway is responsible for multimodal spatiotemporal alignment, operating condition-video leader correlation modeling, mode switching control, bandwidth scheduling, and backtracking / retransmission control; the remote monitoring center is responsible for digital twin rendering, keyframe verification, manual viewing, and alarm linkage.
[0010] See attached general process reference. Figure 2 As shown, the method of the present invention includes eight steps: multimodal acquisition and spatiotemporal alignment, working condition-video leader correlation modeling, scene semantic state modeling, transmission mode switching, precoding enhancement within the prediction window, network-wide resource scheduling, original video retrospective transmission, and remote digital twin calibration. These steps do not operate in isolation, but form a closed loop of "working condition prediction - semantic substitution - resource guarantee - evidence back-up" through unified control of the edge gateway.
[0011] Furthermore, in S1 (Multimodal Acquisition and Spatiotemporal Alignment), the edge gateway provides unified time synchronization for each video node and sensor node, mapping the sensor time series to the video frame time axis. It also establishes a "region-device-video frame" correspondence based on camera pose, monitored area boundaries, and device spatial layout, ensuring that subsequent operating parameters match the corresponding video region and video time. (See Appendix for details.) Figure 1 and attached Figure 2 .
[0012] Furthermore, in S2 (Operating Condition-Video Lead Correlation Modeling), the edge gateway integrates physical mechanism rules with historical data fitting results. Based on the degree of anomaly, trend of change, and multi-parameter coupling relationship of parameters such as pressure surge, vibration increase, flow change, liquid level change, and combustible gas concentration change, it generates an operating condition risk level and prediction lead time. The prediction lead time is used to characterize the time window in which anomalies in the operating condition precede the appearance of abnormal video footage, so as to adjust the transmission and encoding strategies in advance before significant changes in the footage occur. (See Appendix) Figure 2 .
[0013] Furthermore, in S3 (Scene Semantic State Modeling), the system pre-registers static scene information, which includes at least device layout, pipeline topology, camera installation pose, and area boundaries. During runtime, it continuously fuses visual recognition results with sensor inference results to construct a scene semantic state vector representing object state, anomaly category, severity level, and verification fields. This scene semantic state vector is then used as the primary transmission object under low bandwidth conditions. (See Appendix...) Figure 1 and attached Figure 2 .
[0014] Furthermore, in S4 (transmission mode switching), the system adaptively switches between Level-0 (semantic transmission mode), Level-1 (keyframe verification mode), and Level-2 (real-time video mode) based on the operational risk level, semantic consistency, and manual viewing requests. Level-0 is used to send incremental scene status messages, Level-1 is used to append periodic keyframes to the semantic messages, and Level-2 is used to send full-resolution or high-fidelity video streams. Upgrading and downgrading use different thresholds and hold durations to suppress mode oscillations. (See attached diagram.) Figure 3 .
[0015] Furthermore, in S5 (precoding enhancement within the prediction window), after the edge gateway issues a coding switch prediction based on the prediction lead, the video node adjusts the coding parameters within the prediction window, including at least one or more of the following: shortening the GOP (Group of Pictures) length, increasing keyframe density, improving the coding quality of the ROI (Region of Interest), increasing the bitrate cap, or increasing the frame rate cap; wherein, the ROI is jointly determined based on the location of key devices, anomaly type, and historical high-risk areas, thereby obtaining higher fidelity original video at the anomaly initiation stage, see Appendix. Figure 2 and attached Figure 3 .
[0016] Furthermore, in S6 (full network resource scheduling), the edge gateway aggregates the operating condition level, link quality, remaining power, historical anomaly frequency, and current backtracking requests of each region, uniformly calculates the bandwidth allocation weight for each region, and performs preemptive scheduling according to the priority order of combustible gas exceeding limits, pressure exceeding limits, equipment anomalies, and process fluctuations. When the total bandwidth is insufficient, low-priority regions are downgraded from real-time video mode to key frame verification mode or semantic transmission mode, as shown in the appendix. Figure 4 .
[0017] Furthermore, in S7 (raw video circular buffer and backtracking), video nodes continuously store full-resolution raw video and time indexes in their local circular buffer. Backtracking is automatically triggered when the operating condition escalates to a warning or alarm level. Remote monitoring centers are also allowed to initiate manual backtracking according to specified time windows, camera parameters, and video parameters. The backtracking data is sent via an independent backtracking stream queue, with a scheduling priority higher than normal and attention-level services but lower than alarm-level real-time video services. (See attached diagram). Figure 4 .
[0018] Furthermore, in S8 (Remote Digital Twin Rendering and Calibration Feedback), the remote monitoring center reconstructs the scene based on the static scene model and scene semantic state vector, and performs deviation verification on the digital twin results through periodically received keyframes or real-time video. When the verification result is lower than a preset threshold, the edge gateway reduces the semantic confidence and triggers an upgrade or parameter correction, thus forming a closed-loop calibration mechanism of "edge prediction, remote verification, and parameter write-back." (See Appendix) Figure 1 To be continued Figure 3 .
[0019] Through the above technical solution, the present invention does not only adjust a single video coding parameter, but organically couples working condition prediction, semantic substitution transmission, mode switching, precoding enhancement, network-wide scheduling and backtracking transmission, so as to simultaneously take into account three objectives in the bandwidth-constrained oil and gas well surface test scenario: steady-state low load, abnormal low latency and post-event traceability.
[0020] Compared with the prior art, the present invention has at least the following beneficial effects: Compared to schemes that adaptively encode based solely on video motion features, this invention utilizes the leading nature of operating parameters to switch encoding strategies before abnormal scenes appear, which helps improve the integrity of video capture during the anomaly initiation stage. Attached Figure Description
[0021] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments of the invention and, together with the specification, further serve to explain the principles of the invention and enable those skilled in the art to practice and use the invention.
[0022] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the overall process of the method of the present invention; Figure 3 This is a schematic diagram of the transmission mode switching logic of the present invention; Figure 4 This is a schematic diagram illustrating the resource scheduling and backtracking priority of the present invention. Detailed Implementation
[0023] The present invention will be further described below with reference to preferred embodiments. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0024] System Deployment and Multimodal Spatiotemporal Alignment For system deployment methods, please refer to the appendix. Figure 1As shown, each video node and the corresponding sensor node in the area are mapped to a unified area number. The edge gateway performs unified time synchronization on the video timestamp and sensor timestamp, and establishes a "region-device-video frame" association relationship by combining the camera pose, field of view, and area boundary.
[0025] In one embodiment, the edge gateway performs clock drift correction on each node every second and interpolates the sensor sequence onto the video frame timeline, thereby providing a unified time reference for subsequent operational condition-video leader correlation calculations. Camera pose can be registered during the system deployment phase using calibration boards, device outlines, or manual annotation.
[0026] Working Condition - Video Lead Association Modeling For reference on the modeling and triggering process, see Appendix Figure 2 As shown, the edge gateway uses a hybrid model combining physical mechanism rules and historical data fitting to evaluate the leading effect of current operating conditions on subsequent changes in video content. Physical mechanism rules are used to describe process causal relationships such as pressure surges possibly preceding gushing, abnormal vibrations possibly preceding equipment shaking, and sudden flow changes possibly preceding liquid level changes; historical data fitting is used to correct the sensitivity of various operating condition parameters to different video events.
[0027] In one embodiment, the operating condition risk score R_i (representing the operating condition risk score of the i-th monitoring area) is used as the basis for mode switching and resource scheduling. The operating condition risk score can be calculated by formula (1): R_i=Σ(α_k·N_k)+β·T_i+γ·U_i+δ·L_i (1) Where R_i represents the operational risk score of the i-th monitoring area; N_k represents the degree of exceeding or approximating the normalized parameters of the k-th operational condition; α_k represents the weight coefficient of the parameters of the k-th operational condition; T_i represents the parameter change trend term of the i-th monitoring area; U_i represents the multimodal inconsistency penalty term; L_i represents the link vulnerability term; and β, γ, and δ are the corresponding weight coefficients.
[0028] In one optional embodiment, the normal level, attention level, early warning level, and alarm level can correspond to R_i values of less than 0.35, 0.35 to 0.55, 0.55 to 0.75, and not less than 0.75, respectively. These ranges can be recalibrated based on different well sites, different testing processes, and historical sample distributions.
[0029] The prediction lead Δt_i (representing the prediction lead for the i-th monitoring area) is used to characterize the time window in which the abnormal operating condition signal precedes the appearance of the abnormal video image. Δt_i can be determined by correlation analysis between the operating condition sequence and the video event sequence, statistical analysis of process propagation delay, or empirical configuration, with a preferred value of 3 to 30 seconds.
[0030] Scene semantic state modeling and digital twin For a reference on the relationship between scene perception and rendering, please see the appendix. Figure 1 and attached Figure 2 As shown, the system pre-registers static elements, including equipment spatial layout, process connectivity, camera installation position, and monitoring area boundaries; and continuously updates dynamic elements during operation, including valve status, instrument reading range, fluid gushing status, personnel location, and equipment operating status.
[0031] In one embodiment, the system constructs a scene semantic state vector V_i (representing the semantic state vector of the i-th monitoring area), which includes at least the area number, timestamp, object identifier, object status, status confidence, anomaly category, severity level, prediction lead time, and verification field.
[0032] The semantic consistency score C_i (representing the degree of consistency between the visual recognition result and the sensor inference result in the i-th monitoring area) is used to determine whether it is necessary to upgrade from semantic mode to keyframe verification mode or real-time video mode. When the visual judgment and the sensor judgment are consistent, C_i is increased; when the two conflict, C_i is decreased, and keyframe verification is triggered first.
[0033] Transmission mode switching and hysteresis control For example, refer to the attached mode switching logic. Figure 3 As shown, this invention sets up three transmission modes: Level-0 semantic transmission mode, Level-1 keyframe verification mode, and Level-2 real-time video mode. In Level-0 mode, only semantic state incremental messages are transmitted; in Level-1 mode, low-resolution keyframes are sent periodically based on incremental messages; in Level-2 mode, full-resolution or high-fidelity real-time video streams are uploaded.
[0034] To suppress frequent mode oscillations, the system employs hysteresis control. Specifically, when R_i crosses the upper cut-off threshold and remains for a first holding duration, a ramp-up is performed; when R_i falls below the lower cut-off threshold and remains for a second holding duration while keyframe calibration passes, a ramp-down is performed. Preferably, the first holding duration can be set to 2 to 5 seconds, the second holding duration can be set to 10 to 30 seconds, and the second holding duration is longer than the first holding duration.
[0035] In one embodiment, when C_i is less than 0.80, the system upgrades from Level-0 to Level-1; when C_i is less than 0.60 or R_i enters the warning or alarm level, the system upgrades to Level-2. Manual inspection commands can also directly trigger the upgrade to Level-2.
[0036] Precoding enhancement within the prediction window For example, refer to the appendix for precoding-enhanced trigger relationships. Figure 2 and attached Figure 3 As shown, after the edge gateway generates a coding switch prediction signal based on Δt_i, the video node performs precoding enhancement within the prediction window before the image changes significantly.
[0037] In this invention, the length of a group of images (GOP) can be shortened, the keyframe density can be increased, the coding quality of the region of interest (ROI) can be improved, and the upper limits of bitrate and frame rate can be increased as needed. The ROI is preferably determined jointly by the locations of key devices in a static scene, high-risk areas corresponding to anomaly categories, and the spatial distribution of historical events.
[0038] For example, when the system detects a rapid rise in wellhead pressure, the wellhead valve assembly and separator inlet area can be set as ROI in advance to improve the keyframe density and coding quality of the above areas; when the system detects increased equipment vibration, the coding quality of the equipment body and the area near its connecting pipelines can be improved in advance.
[0039] Joint bandwidth scheduling across the entire network For a reference to the network-wide resource allocation process, please see the attached document. Figure 4 As shown, the edge gateway aggregates the operating conditions, link status, remaining power, historical anomaly frequency, and backtracking requests of each region, and performs unified scheduling of the total bandwidth.
[0040] In one embodiment, the bandwidth scheduling weight W_i (representing the bandwidth allocation weight of the i-th monitoring area) can be calculated by formula (2): W_i=λ_1·P_i+λ_2·Q_i+λ_3·E_i+λ_4·H_i+λ_5·R_i (2) Where W_i represents the bandwidth allocation weight of the i-th monitoring area; P_i represents the operating condition priority; Q_i represents the wireless link quality; E_i represents the remaining power; H_i represents the historical abnormal frequency; R_i represents the current operating condition risk score; and λ_1 to λ_5 are the corresponding weight coefficients.
[0041] When multiple areas experience concurrent anomalies, the preferred priority order is: combustible gas exceeding limits, pressure exceeding limits, equipment malfunction, and process fluctuations. Alarm-level real-time video streams can utilize preemptive bandwidth allocation, while lower-priority areas are downgraded to keyframe verification mode or semantic transmission mode.
[0042] Original video loop buffer and backtracking For priority references regarding backtracking and retransmission, please see the appendix. Figure 4 As shown, the video node continuously saves the original video file and time index in a local circular cache. Preferably, the local circular cache is saved for 30 to 120 minutes, more preferably 60 minutes.
[0043] Retrospective retransmission can be initiated in two ways: automatic and manual. Automatic triggering means that the edge gateway automatically sends a retransmission request when it detects that the working condition level has been raised to the warning level or alarm level; manual triggering means that the remote monitoring center sends a retransmission request after specifying the retrospective time window, camera identifier, and video parameters.
[0044] In a preferred embodiment, the default backtracking time window is 30 seconds before the anomaly to 120 seconds after the anomaly. Retransmission uses a backtracking stream queue independent of the real-time video stream. In queue scheduling, alarm-level real-time video streams take priority over the backtracking stream queue, which in turn takes priority over normal-level and attention-level semantic messages and keyframes.
[0045] Remote digital twin rendering and calibration closed loop The remote monitoring center performs digital twin rendering based on a static scene model and scene semantic state vectors, continuously displaying the status of equipment, valves, personnel, and fluids even under weak network conditions. To avoid model drift caused by long-term reliance on semantic messages, the system periodically receives keyframes or real-time video and compares them with the rendered images to identify discrepancies.
[0046] In one embodiment, SSIM (Structural Similarity Index) or object state-based consistency index can be used as the basis for deviation evaluation. When the structural similarity between the rendered image and the actual video is lower than a preset threshold, the edge gateway reduces the semantic confidence, and if necessary, forces an upgrade to keyframe verification mode or real-time video mode, and updates the threshold parameters and object state mapping relationship. Specific Implementation Taking a shale gas well test site as an example, three video nodes and several sensor nodes are deployed in the wellhead area, separator area, and storage tank area, respectively, and are managed uniformly by the same edge gateway. The system is initially in Level-0 semantic transmission mode, with an average semantic message length of 80 to 300 bytes.
[0048] When the pressure in the wellhead area rises rapidly within 6 seconds, and the vibration signal rises synchronously, the edge gateway calculates Δt_i as 8 seconds based on historical samples and issues a pre-encoding switch warning signal before the pressure exceeds the alarm threshold. The wellhead area video node then shortens the GOP, improves the ROI encoding quality, and enters Level-1 keyframe verification mode; if a gushing state is subsequently detected or R_i enters the warning level, it immediately upgrades to Level-2 real-time video mode.
[0049] If a medium-level process fluctuation occurs simultaneously in the separator area, the edge gateway prioritizes bandwidth protection for the wellhead area based on the W_i calculation result, downgrading the separator area to keyframe verification mode. After the anomaly ends, the system automatically retransmits the original video clips from 30 seconds before the anomaly to 120 seconds after the anomaly in the wellhead area, thus achieving a balance between low-bandwidth operation and evidence integrity.
Claims
1. A working condition semantic collaborative driving method for intelligent transmission of multimodal video for surface testing of oil and gas wells, characterized in that: The method, applied to a monitoring transmission system consisting of video nodes, sensor nodes, edge gateways, and a remote monitoring center, includes at least a video mode, a working condition sensing mode, and a scene structure mode, and comprises: (1) Collect video streams and operating parameters of the target area, and perform spatiotemporal alignment of video timestamps, sensor timestamps, area numbers and camera poses; (2) Based on the testing process mechanism rules of oil and gas wells, historical synchronous samples and real-time operating conditions trends, construct the operating condition-video leader association model, output the operating condition risk level, semantic consistency results and prediction lead of the current area, and generate the coded switching warning signal; (3) Construct scene semantic state vectors based on static scene models and dynamic state perception results; (4) Based on the working condition risk level, semantic consistency results, prediction lead time and manual viewing request, adaptively switch between semantic transmission mode, key frame verification mode and real-time video mode; (5) Upon receiving the encoding switching warning signal, perform precoding enhancement on the video stream within the time window corresponding to the prediction lead; (6) The edge gateway aggregates the working condition level, link status, remaining power and backtracking requests of multiple regions and performs joint bandwidth scheduling across the entire network; (7) The video node continuously saves the original video in the local circular cache, and extracts the original video segments as needed for backtracking and retransmission when there is an abnormal upgrade or when a backtracking instruction is received; (8) The remote monitoring center performs digital twin rendering based on the scene semantic state vector and performs closed-loop calibration on the scene semantic state vector based on key frames or real-time video.
2. The method for intelligent transmission of multimodal video driven by semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, The operating condition-video leader correlation model is a hybrid model that combines a physical mechanism rule engine and a historical data fitting model; the prediction lead is determined based on at least one of the correlation between the operating condition sequence and the video event sequence, process propagation delay, and empirical configuration, and the value of the prediction lead ranges from 3 seconds to 30 seconds.
3. The method for intelligent transmission of multimodal video driven by working condition semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, The scene semantic state vector includes at least a portion of the following: region number, timestamp, object identifier, object state, state confidence, anomaly category, severity level, prediction lead, and verification field. In semantic transmission mode, the video node only sends incremental messages of the scene semantic state vector; in keyframe verification mode, it sends the incremental messages and periodic keyframes; and in real-time video mode, it sends the real-time video stream.
4. The method for intelligent transmission of multimodal video driven by semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, Hysteresis control is used when switching between the semantic transmission mode, keyframe verification mode and real-time video mode; when the working condition risk level crosses the upper cut-off condition and continues for a first holding time, the level is upgraded; when the working condition risk level falls back to below the lower cut-off condition and continues for a second holding time and the calibration is passed, the level is downgraded, wherein the second holding time is longer than the first holding time.
5. The method for intelligent transmission of multimodal video driven by semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, The precoding enhancements include one or more of the following measures: increasing keyframe density, shortening image group length, improving region of interest coding quality, increasing bitrate cap, increasing frame rate cap, and improving buffer fidelity.
6. The method for intelligent transmission of multimodal video driven by working condition semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, The network-wide bandwidth joint scheduling calculates the regional bandwidth allocation weights based on operating condition priority, wireless link quality, remaining power, historical abnormal frequency, and current risk level. When multiple regions experience concurrent anomalies, bandwidth is allocated in a preemptive or weighted manner according to the priority order of combustible gas exceeding limits, pressure exceeding limits, equipment anomalies, and process fluctuations.
7. The method for intelligent transmission of multimodal video driven by semantic collaboration for surface testing of oil and gas wells according to claim 1, characterized in that, The video node local cyclic cache stores at least the most recent 30 to 120 minutes of original video; the backtracking and retransmission supports both automatic and manual triggering, and adopts a backtracking stream queue independent of the real-time video stream. The scheduling priority of the backtracking stream queue is higher than that of normal and attention-level service streams, but lower than that of alarm-level real-time video streams.
8. A working condition semantic collaborative multimodal video intelligent transmission system for surface testing of oil and gas wells, characterized in that, The system includes video nodes, sensor nodes, edge gateways, and a remote monitoring center. The video nodes are used to acquire video streams, perform edge detection, establish local loop buffers, and perform precoding enhancement and multi-level transmission mode switching according to control commands. The sensor nodes are used to acquire operating parameters at the oil and gas well testing site. The edge gateway is used to perform spatiotemporal alignment of video data and operating data, construct an operating condition-video lead correlation model, establish a scene semantic state model, generate encoding switching warning signals, perform joint bandwidth scheduling across the entire network, and issue backtracking and retransmission requests. The remote monitoring center is used to receive semantic messages, keyframes, real-time video, and backtracking video, and to perform digital twin rendering, deviation verification, manual viewing, and alarm linkage.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a program executable by the processor, which, when executed, implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method according to any one of claims 1 to 7.