Multimodal synergistic and environment knowledge augmented power edge monitoring method and system
Patent Information
- Application Number
- CN202610780143.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-02
AI Technical Summary
[0007]本发明提供一种多模态协同与环境知识增强的电力边缘监测方法,解决了现有电力边缘监测中因缺乏环境语义理解与动态适应能力,导致复杂环境下误报漏报率高、动态隐患识别难及无源场景下算力能耗失衡的问题
[0051]By employing a cascaded inference mechanism of "wide-area initial screening + precise re-judgment," combined with spatiotemporal feature enhancement technology using gimbal zoom and video keyframe extraction, the false alarm rate in complex scenarios such as dense fog in mountainous areas and urban lighting is significantly reduced. By dynamically injecting multimodal large models with RAG-based environmental knowledge, edge terminals can achieve scene adaptation capabilities without retraining, eliminating persistent false alarms in specific environments. Simultaneously, by constructing an optimization objective function that includes dynamic energy penalties, heterogeneous tasks are accurately mapped to corresponding computing units and a smooth degradation strategy is executed. This ensures the accuracy of hazard assessment while effectively avoiding computing power contention and power consumption runaway, perfectly balancing the engineering contradiction between "accurate detection" and "long lifespan" for passive monitoring terminals.
Smart Images

Figure CN122313405B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent power inspection technology, specifically to a power edge monitoring method and system that combines multimodal collaboration and environmental knowledge enhancement. Background Technology
[0002] With the continuous advancement of intelligent power grid inspection, edge computing-based visual monitoring terminals have been widely used in power grid operation and maintenance. These terminals typically deploy lightweight target detection algorithms on the edge to achieve local real-time identification and rapid response to potential hazards and defects. However, in the complex and ever-changing power field environment, monitoring solutions relying solely on traditional computer vision technology have gradually revealed many limitations, making it difficult to meet the stringent requirements of the power system for safe and reliable operation.
[0003] False alarms are particularly prominent in existing monitoring devices under complex natural environments. Dense fog common in mountainous areas, industrial emissions around substations, and complex light reflections in cities at night can all severely interfere with pure vision algorithms based on visible light imaging. Due to a lack of deep understanding of environmental semantics, the algorithms are prone to misinterpreting fog as smoke or fire, water reflections or lights as fire, and even water stains or shadows on equipment surfaces as oil leaks. These frequent false alarms not only generate massive amounts of invalid data, increasing the burden of manual review in the cloud, but also cause terminal devices to frequently trigger energy-intensive review processes and data transmission operations, resulting in a huge waste of valuable edge computing resources and battery power.
[0004] Meanwhile, existing technologies are inadequate in addressing dynamic hazards with spatiotemporal evolution characteristics. For example, conductor galloping on transmission lines, the hanging of large foreign objects, and the diffusion of smoke all require analysis of continuous time series for accurate detection. However, most current edge monitoring terminals rely solely on single-frame static images for judgment, severing the evolutionary logic of hazards over time. This leads to numerous missed risks and makes it difficult to effectively identify dynamic threats in their nascent stages or those constantly changing in form.
[0005] Besides the shortcomings at the perception level, the rigidity of edge algorithms is another major pain point. Once a model is trained and deployed to the terminal, its detection thresholds and judgment rules become fixed and cannot be dynamically adjusted according to the specific environment. This leads to persistent false alarms or missed alarms in specific scenarios. For example, when there are specific industrial steam emissions around a substation year-round, even if cloud-based manual verification has repeatedly confirmed it as normal environmental interference, the edge terminal, lacking the ability to remember and integrate environmental context, will stubbornly trigger alarms the next time it encounters the same scenario. Existing edge-cloud collaboration often only stays at the level of model retraining or parameter distribution, lacking dynamic management and real-time utilization of local environmental knowledge.
[0006] Finally, existing technologies also have significant shortcomings in the coordinated scheduling of computing resources. When the edge device needs to run multiple tasks simultaneously, such as object detection models and multimodal large models, the lack of fine-grained allocation strategies and task-level energy consumption management mechanisms for heterogeneous computing units often leads to computing power contention among CPUs, GPUs, and NPUs. This chaotic resource competition not only causes a surge in inference latency but also results in runaway power consumption, which greatly threatens the long-term stable operation of edge monitoring terminals that typically use solar power or passive designs. Summary of the Invention
[0007] This invention provides a power edge monitoring method with multimodal collaboration and environmental knowledge enhancement, which solves the problems of high false alarm and false negative rates, difficulty in identifying dynamic hidden dangers, and imbalance of computing power and energy consumption in passive scenarios caused by the lack of environmental semantic understanding and dynamic adaptation capabilities in existing power edge monitoring.
[0008] This invention is achieved through the following technical solution:
[0009] In a first aspect, this application provides a power edge monitoring method with multimodal collaboration and environmental knowledge enhancement, comprising the following steps:
[0010] Deploy a heterogeneous computing architecture at the edge and use a lightweight target detection model to perform wide-area initial screening of the acquired visible light images;
[0011] When a suspected potential hazard target is initially identified, a spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features.
[0012] The lightweight multimodal large model on the client side is invoked, and deep semantic re-judgment is performed by combining the high-resolution zoom image and key frame sequence. Environmental knowledge based on retrieval enhancement is introduced to dynamically enhance the re-judgment process and obtain structured judgment information.
[0013] Based on structured analysis information, combined with the current energy state and heterogeneous computing load, an optimization objective function is constructed that includes the benefits of hazard assessment and dynamic energy penalties. The optimal task combination is obtained by solving the function, and the activated tasks are mapped to the corresponding computing units for execution.
[0014] A further optimized solution is that the triggering spatiotemporal verification mechanism specifically includes:
[0015] Based on the coordinates and size of the suspected potential hazard target in the image, calculate the horizontal rotation angle, vertical rotation angle and zoom magnification of the gimbal, and control the gimbal to perform zoom shooting to obtain the high-resolution zoom image.
[0016] Switch the camera module to video capture mode, calculate the motion intensity of the image through inter-frame difference, dynamically adjust the sampling frame rate based on the motion intensity, and extract the frame with the highest motion intensity from the captured video segment as the key frame sequence.
[0017] A further optimization scheme is that the introduction of environmental knowledge based on retrieval enhancement to dynamically enhance the re-judgment process specifically includes:
[0018] When the cloud confirms that it is a false alarm, the system receives the false alarm environment context knowledge sent by the cloud, vectorizes it and stores it in the lightweight vector database on the edge, and configures a time decay-based weight model for the knowledge item.
[0019] Before performing the deep semantic re-judgment, the features of the current suspected hidden danger are converted into query vectors, and the most matching environmental knowledge is retrieved from the vector database.
[0020] The most matching environmental knowledge and basic prompt word template are combined to form an enhanced prompt word, and the enhanced prompt word is input into the lightweight multimodal large model on the edge.
[0021] A further optimization scheme is as follows: the solution to obtain the optimal task combination specifically includes:
[0022] Define a Boolean variable for task execution, whereby the tasks include lightweight object detection, gimbal control, video encoding, and multimodal large model inference.
[0023] Based on the current energy score and remaining energy at the edge, a dynamic energy penalty coefficient is set, and the optimization objective function is solved to obtain the optimal task combination for the current time slot;
[0024] When the current energy score is lower than a preset threshold, a task degradation strategy is executed: in medium power mode, high-energy-consuming video encoding tasks are turned off, and only gimbal zoom and single-frame VLM inference are retained; in low power mode, deep task degradation is triggered, freezing VLM and gimbal control tasks, and only low-frequency YOLO detection and simplified backhaul are maintained.
[0025] A further optimization scheme is that the calculation of the horizontal rotation angle, vertical rotation angle, and zoom magnification of the gimbal specifically includes:
[0026] Based on the offset of the target center point relative to the image center, and combined with the horizontal and vertical field of view of the camera, the horizontal rotation angle and the vertical rotation angle are calculated respectively.
[0027] An exponential smoothing filter is introduced to preprocess the calculated rotation angle to reduce control oscillations;
[0028] The target zoom factor is calculated based on the ratio of the target frame area to the preset standard area, combined with the maximum optical zoom factor limit.
[0029] A further optimized solution is that the formula for calculating the target zoom factor is:
[0030] ;
[0031] in, For the target zoom magnification, This is the maximum optical zoom magnification. Based on the zoom magnification, and These are the width and height of the target bounding box, respectively. and These are the width and height of the preset standard area, respectively.
[0032] A further optimization scheme is that extracting the frame with the highest motion intensity as the keyframe sequence specifically includes:
[0033] The difference image is obtained by performing absolute difference operation on two adjacent frames, and then Gaussian filtering is used to denoise the difference image;
[0034] The percentage of pixels whose difference values exceed a preset threshold in the denoised difference image is statistically analyzed to quantify the motion intensity of the image.
[0035] The video segment is evenly divided into multiple sub-segments, and the frame with the highest motion intensity in each sub-segment is selected as the keyframe of that sub-segment.
[0036] A further optimization scheme is proposed, in which the time-decay-based weight model is specifically defined as follows:
[0037] The weight of a knowledge item decays exponentially over time, and its calculation formula is as follows:
[0038] ;
[0039] in, Let be the weight of the knowledge item at time t. As the initial weights, The attenuation coefficient is... Time for knowledge creation;
[0040] When the weight of a knowledge item is lower than the preset retention threshold, expiration cleanup is automatically triggered. When new environmental context knowledge is received, its cosine similarity with existing knowledge is calculated. If the similarity is greater than the preset deduplication threshold, an overwrite update is performed; otherwise, append storage is performed.
[0041] A further optimization scheme is as follows: the lightweight multimodal large model on the calling side, combined with the high-resolution zoom image and keyframe sequence for deep semantic re-judgment, specifically includes:
[0042] A visual-language adapter is used to reduce the dimensionality and align the visual features of high-resolution zoom images and keyframe sequences to generate an aligned visual token sequence.
[0043] The visual token sequence is concatenated with the text token sequence corresponding to the structured prompt words to form a multimodal input sequence;
[0044] The multimodal input sequence is input into a quantized and compressed language model to generate a JSON-formatted output containing the judgment result, category, confidence level, and reason in an autoregressive manner.
[0045] Secondly, this application provides a multimodal collaborative and environmental knowledge-enhanced power edge monitoring system, comprising:
[0046] The wide-area initial screening module is used to deploy a heterogeneous computing architecture at the edge and use a lightweight target detection model to perform wide-area initial screening of the acquired visible light images.
[0047] The spatiotemporal verification module is communicatively connected to the wide-area initial screening module. When a suspected hidden danger target is detected during the initial screening, the spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features.
[0048] The semantic review and knowledge enhancement module is connected to the spatiotemporal review module. It is used to call the lightweight multimodal large model on the end side, combine the high-resolution zoom image and key frame sequence to perform deep semantic review, and introduce environmental knowledge generated based on retrieval enhancement to dynamically enhance the review process and obtain structured judgment information.
[0049] The heterogeneous scheduling and energy management module communicates with the semantic review and knowledge enhancement module. It is used to construct an optimization objective function that includes the benefits of hidden danger assessment and dynamic energy penalties based on structured judgment information, combined with the current energy state and heterogeneous computing load. The optimal task combination is solved and the activated tasks are mapped to the corresponding computing units for execution.
[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] By employing a cascaded inference mechanism of "wide-area initial screening + precise re-judgment," combined with spatiotemporal feature enhancement technology using gimbal zoom and video keyframe extraction, the false alarm rate in complex scenarios such as dense fog in mountainous areas and urban lighting is significantly reduced. By dynamically injecting multimodal large models with RAG-based environmental knowledge, edge terminals can achieve scene adaptation capabilities without retraining, eliminating persistent false alarms in specific environments. Simultaneously, by constructing an optimization objective function that includes dynamic energy penalties, heterogeneous tasks are accurately mapped to corresponding computing units and a smooth degradation strategy is executed. This ensures the accuracy of hazard assessment while effectively avoiding computing power contention and power consumption runaway, perfectly balancing the engineering contradiction between "accurate detection" and "long lifespan" for passive monitoring terminals. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0053] Figure 1 A flowchart illustrating a multimodal collaborative and environmental knowledge-enhanced power edge monitoring method provided in this embodiment;
[0054] Figure 2 This embodiment provides a functional block diagram of a power edge monitoring system with multimodal collaboration and environmental knowledge enhancement. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0056] Firstly, such as Figure 1 As shown, this application provides a power edge monitoring method with multimodal collaboration and environmental knowledge enhancement, including the following steps:
[0057] Step S1: Deploy a heterogeneous computing architecture at the edge and use a lightweight target detection model to perform wide-area initial screening of the acquired visible light images;
[0058] Step S2: When a suspected potential hazard is initially identified, a spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features.
[0059] Step S3: Call the lightweight multimodal large model on the terminal side, combine the high-resolution zoom image and key frame sequence to perform deep semantic re-judgment, and introduce environmental knowledge based on retrieval enhancement to dynamically enhance the re-judgment process and obtain structured judgment information;
[0060] Step S4: Based on the structured assessment information, combined with the current energy state and heterogeneous computing load, construct an optimization objective function that includes the benefits of hidden danger assessment and dynamic energy penalties, solve for the optimal task combination, and map the activated tasks to the corresponding computing units for execution.
[0061] This invention utilizes a cascaded reasoning mechanism that combines wide-area initial screening with precise re-judgment, deeply integrating the spatial detail capture capability of gimbal zoom with video temporal analysis technology, to significantly suppress false alarm interference caused by complex environmental factors such as dense fog in mountainous areas and city lights.
[0062] By using retrieval-enhanced generation technology, proprietary knowledge of power scenarios is dynamically injected into a multimodal large model, enabling edge terminals to have the ability to adapt to different scenarios without relying on full retraining, thus eradicating the persistent problem of false alarms in fixed scenarios with extremely low computing power cost.
[0063] Furthermore, by constructing an optimization objective function that includes a dynamic energy penalty coefficient, precise allocation and smooth degradation of heterogeneous computing tasks are achieved. This ensures high accuracy in hazard assessment while effectively avoiding competition for computing resources and unnecessary power consumption, thus solving the engineering challenge of the mutual constraint between detection efficiency and energy consumption in passive monitoring terminals.
[0064] In one embodiment, step S1: Deploying a heterogeneous computing architecture at the edge and using a lightweight target detection model to perform wide-area initial screening of the acquired visible light images, specifically includes the following steps:
[0065] Step S11: Configure heterogeneous computing hardware including a neural network processor, a graphics processor and a central processing unit, load the pre-trained lightweight object detection model into the neural network processor, and output the initialized heterogeneous computing architecture.
[0066] Specifically, the lightweight target detection model adopts the YOLOv8 architecture, and the neural network processor is the NPU unit built into the RK3588 chip.
[0067] Step S12: Control the camera module to acquire visible light images of the scene at a preset fixed frequency (e.g., 1 to 2 frames per second), transmit the image data stream to the neural network processor in the heterogeneous computing architecture, and output the image data frames to be processed.
[0068] Specifically, the preset frequency is set to 1 to 2 frames per second to balance the real-time performance of detection with the power consumption of the edge.
[0069] Step S13: Drive the neural network processor to call the lightweight target detection model to perform inference operations on the image data frame, identify potential targets in the image and output a list of suspected hidden danger targets with location coordinates and confidence levels.
[0070] Specifically, the potential targets include flames, smoke, foreign objects in the wiring, and areas of equipment oil leakage.
[0071] This embodiment executes the YOLO model embedded in the NPU, which not only frees up CPU resources for processing business logic but also significantly reduces the inference energy consumption per unit image. At the same time, in conjunction with a low frame rate acquisition strategy, it greatly extends the battery life of edge terminals in passive scenarios while ensuring that no sudden hidden dangers are missed, thus reserving sufficient power reserves for subsequent in-depth analysis.
[0072] In one embodiment, step S2: When a suspected potential hazard is initially detected, a spatiotemporal verification mechanism is triggered to acquire a high-resolution zoom image containing spatial details and a keyframe sequence of temporal dynamic features, specifically including the following steps:
[0073] Step S21: Analyze the location coordinates and size information of the suspected hidden danger targets in the list, calculate the horizontal rotation angle, vertical rotation angle and lens zoom of the gimbal, and output the gimbal control command set.
[0074] Specifically, based on the offset of the target center point relative to the image center, and combined with the horizontal and vertical field of view angles of the camera, the horizontal and vertical rotation angles are calculated respectively. Let the width of the image captured by the camera be W, the height be H, and the coordinates of the top-left corner of the target detection box be... The coordinates of the lower right corner are Then the coordinates of the target center point The calculation is as follows:
[0075] Equation (1)
[0076] width of the target box and height The calculation is as follows:
[0077] Equation (2)
[0078] Horizontal rotation angle and vertical rotation angle The calculation formula is:
[0079] Equation (3)
[0080] Equation (4) In the formula, For horizontal field of view, This refers to the vertical field of view.
[0081] Since directly calculated angle values may contain slight fluctuations, directly driving the motor can easily lead to mechanical wear and screen shaking. To improve the smoothness of the control response and avoid frequent motor starts and stops, an exponential smoothing filter is introduced to preprocess the control commands.
[0082] Equation (5)
[0083] Equation (6)
[0084] in For smoothing coefficients, and From the perspective of the target, and This represents the actual angle at the previous moment. When the deviation between the target and the current viewing angle is small, smoothing filtering can effectively reduce control oscillations. Preferably, an exponential smoothing filter is introduced to preprocess the calculated rotation angle to reduce control oscillations of the gimbal motor.
[0085] Step S22: Adjust the camera module's posture and focal length according to the gimbal control command set, capture close-up images of the target area, and output a high-resolution zoom image containing rich texture details.
[0086] To ensure that the acquired image covers the entire target while preserving the texture details required for recognition, the zoom level needs to be dynamically determined based on the target's proportion in the image. The zoom level calculation must comprehensively consider the ratio of the target bounding box area to a preset standard area, avoiding excessive magnification while maintaining detail clarity. The standard target bounding box size is set as ( , The base zoom magnification is... The maximum optical zoom is The target zoom factor Z is calculated using the following formula:
[0087] Equation (7)
[0088] To further capture the dynamic evolution characteristics of potential hazards (such as the speed of smoke diffusion or the trajectory of foreign objects) and compensate for the lack of temporal perception in single-frame static images, the system immediately executes the following temporal verification process after completing zoom shooting:
[0089] Step S23: Switch the camera module to video capture mode to capture dynamic images with a set duration of 10 seconds, quantify the motion intensity of the image through inter-frame difference operation, extract key frame sequences with significant motion features, and output key frame sequences representing temporal dynamic features.
[0090] Specifically, the video segment is evenly divided into multiple sub-segments, and the frame with the highest motion intensity in each sub-segment is selected as the keyframe for that sub-segment. The difference calculation and Gaussian noise reduction process is as follows: Let... This represents the pixel coordinates of frame t at time t. The grayscale value at that point, the difference image between two adjacent frames. We obtain the following through absolute difference operations:
[0091] Equation (8)
[0092] Next, Gaussian filtering is used to preprocess the difference image to suppress noise interference:
[0093] Equation (9)
[0094] in, This is a two-dimensional Gaussian kernel function, where σ is the standard deviation parameter and k is the kernel radius. (Image motion intensity) Defined as the proportion of pixels whose difference value exceeds a threshold τ out of the total number of pixels:
[0095] Equation (10)
[0096] in, This is an indicator function that takes the value 1 when the condition is met, and 0 otherwise. The value range is [0,1], and a larger value indicates a higher proportion of the moving area in the image. Based on Dynamically adjust the frame rate of video encoding to balance dynamic detail capture with edge computing power and storage consumption:
[0097] Equation (11)
[0098] When exercise is intense At that time, use a high frame rate Capture rapidly changing details;
[0099] When the movement is weak At that time, use a low frame rate Save resources;
[0100] If the state is intermediate, linear interpolation is performed. Combining uniform sampling and motion detection, the video segment is uniformly divided into N sub-segments, and the frame with the highest motion intensity in each sub-segment is selected as the candidate keyframe.
[0101] Equation (12)
[0102] in, This represents the total number of frames in the video. For the first The time interval of each segment. The final output keyframe sequence is This sequence eliminates a large number of redundant static frames, retains the core dynamic features of the evolution of hidden dangers, and will be used as input for subsequent multimodal large model re-judgment.
[0103] This embodiment solves the misjudgment caused by the lack of texture of small distant targets by gimbal adaptive zoom, while the key frame extraction algorithm based on motion intensity effectively eliminates a large number of static redundant frames and accurately locks the most representative dynamic features in the process of smoke diffusion or wire dancing, thereby greatly improving the accuracy of subsequent semantic re-judgment at the input end.
[0104] In one embodiment, step S3: The lightweight multimodal large model on the edge is invoked, and deep semantic re-judgment is performed by combining the high-resolution zoom image and keyframe sequence. Environmental knowledge generated based on retrieval enhancement is introduced to dynamically enhance the re-judgment process, obtaining structured judgment information. Specifically, this includes the following steps:
[0105] Step S31: Use a visual-language adapter to extract features and transform dimensions of the high-resolution zoom image and keyframe sequence, generate a visual token sequence aligned with the text semantic space, and output the aligned multimodal feature vector.
[0106] Specifically, the vision-language adapter consists of two layers of perception mechanisms, used to reduce the dimensionality of visual features and map them to the embedding dimension of the language model. The input set of visual images is... We use ViT to divide each image into blocks and extract features. Let the first block be... The number of patches for each image is The hidden layer dimension is ViT outputs the original visual feature matrix. To reduce the computational overhead of the language model, this invention employs a two-layer perceptron model to construct a vision-language adapter for feature dimensionality reduction and alignment:
[0107] Equation (13)
[0108] Where GELU(·) is a smooth nonlinear activation function. This is the bias vector for the first fully connected layer. This is the bias vector for the second fully connected layer. Let be the original visual feature matrix of the i-th image. This is the weight matrix of the first fully connected layer. , This is the weight matrix of the second fully connected layer. , For the intermediate hidden layer dimension of the visual-language adapter, The embedding dimension of the language model is used to obtain the aligned visual token sequence of the i-th image after processing by the visual-language adapter. .
[0109] Visual features alone are often insufficient to distinguish interference in specific environments (such as substation steam and fire smoke), therefore the following environmental context knowledge needs to be introduced to assist in the judgment;
[0110] Step S32: Convert the current suspected hidden danger features into query vectors, retrieve matching environmental knowledge from the edge vector database, concatenate and fuse the retrieved knowledge text with the basic prompt word template, and output enhanced prompt words.
[0111] The environmental knowledge includes the characteristics of industrial emissions and seasonal environmental disturbance patterns around the substation. Once a false alarm is confirmed by manual review on the cloud, the environmental context knowledge text will be... The data is sent to the monitoring terminal, segmented and quantized, and then stored in a local lightweight vector database.
[0112] To prevent the knowledge base from expanding excessively and accumulating noise over time, a dynamic update mechanism based on time decay is introduced for knowledge items. The initial weights are The decay model of its weights over time t is as follows:
[0113] Equation (14)
[0114] in, The attenuation coefficient is... Time for knowledge creation.
[0115] When the weight of knowledge items Below the preset retention threshold When expired knowledge is automatically cleaned up, strict deduplication checks are performed during knowledge entry and updates. Specifically, when new contextual knowledge is received, its cosine similarity to existing knowledge is calculated (using vector dot product normalization). If the similarity is greater than a preset deduplication threshold, the deduplication is removed. If the value is 0.85, then an overwrite update operation is performed; otherwise, append storage is performed, thus effectively avoiding knowledge base redundancy.
[0116] Furthermore, when conducting environmental knowledge retrieval, it is necessary to comprehensively consider the historical validity of the knowledge (time weight) and its relevance to the current scene (semantic similarity) to ensure that the output enhanced prompts are both timely and accurate.
[0117] When the detection algorithm deployed by the terminal monitoring device detects a suspected potential hazard, it transforms the target's category characteristics and coordinate information into a query vector. Approximate nearest neighbor retrieval is performed in the local vector database to recall the Top-K set of environmental knowledge most relevant to the current scene. Calculate the attention matching score between the query and each knowledge item. :
[0118] Equation (15)
[0119] in, For the feature vector of knowledge, Assign it a time decay weight. Select the knowledge text with the highest score. As a prior environment context.
[0120] By combining the retrieved most matching environmental knowledge With basic prompt word template To form enhanced prompts :
[0121] Equation (16)
[0122] The enhanced prompt word is input into the multimodal large model along with the image features, so that VLM inference is not based on fixed training rules, but rather on dynamic judgment of the current specific environmental context, thereby achieving adaptive cognitive update and further reducing false positives.
[0123] Step S33: Input the enhanced prompt word and the aligned multimodal feature vector into the quantized and compressed edge-side lightweight multimodal large model to perform deep semantic reasoning operations. The Prompt structure of the multimodal input is as follows:
[0124] <system> You are a power line inspection expert, and you need to determine whether the target in the picture is a real hidden danger or defect.< / system>
[0125] {Visual Token Placeholder}
[0126] <question> Please determine whether the image represents a real production hazard or environmental interference. Please output strictly according to the JSON format: {"judgment": "True / False", "category": "specific category of hazard / interference", "confidence": 0.00-1.00, "reason": "brief judgment criteria"}< / question>
[0127] The final output of the model must strictly adhere to the preset JSON format (e.g., {"judgment":"True","category":"smoke","confidence":0.92,"reason":"The white gas in the image shows no diffusion trend, consistent with the characteristics of steam in a substation"}). Adding any additional non-JSON content is prohibited; otherwise, the output will be considered invalid.
[0128] Considering the computing power and memory limitations of edge terminals (such as RK3588 / RK3576) (NPU provides 6 TOPS computing power, 8-16GB memory), the base model of this solution uses a lightweight vision-language model with a parameter scale between 2B and 4B (such as MiniCPM-V-2.6 or Qwen2.5-VL-3B).
[0129] To meet the real-time requirements of the edge devices, a hybrid precision quantization strategy was adopted for model deployment: the language model weights were compressed using Weight-Only INT4 quantization (based on the GPTQ algorithm), while the visual encoder (ViT) retained FP16 precision. This approach strictly controls the overall memory usage of the model to within 4GB while ensuring that the loss of inference accuracy is controllable.
[0130] Specifically, during the inference phase, the system concatenates the visual tokens of all images in sequence and combines them with the text Prompt sequence encoded by the Tokenizer. (Where L is the text length) are combined to form a complete multimodal input sequence. Its specific structure is shown in equation (17):
[0131] Equation (17)
[0132] The multimodal input sequence is fed into a quantized and compressed language model, which generates output tokens character by character in an autoregressive manner until a terminator is triggered or the output length reaches its limit. The text generated by the large model strictly adheres to a preset JSON format.
[0133] After extracting the target field using regular expressions, logical judgment is then performed. A VLM confidence threshold is introduced. (Settings of this invention) If the "judgment" field value is True and confidence ≥ If the alarm is detected, the device will immediately upload the alarm information, the original zoomed image, the key frame, and the model output reason to the cloud; otherwise, it will be considered a false alarm caused by environmental interference.
[0134] Given the energy-constrained nature of edge terminals, not all analysis tasks require the same computing power. The system needs to dynamically allocate resources based on the aforementioned structured analysis information and the current energy status.
[0135] This embodiment leverages the deep semantic understanding capabilities of a multimodal large model (VLM) combined with retrieval-enhanced generation (RAG) technology to effectively address the semantic gap problem of pure visual algorithms in complex environments. Specifically, by dynamically enhancing the Prompt function with environmental knowledge, the model can distinguish between "industrial steam around a substation" and "real fire smoke," eliminating persistent false alarms in specific scenarios. Simultaneously, by employing INT4 quantization and visual adapter technology, the large model is successfully deployed on resource-constrained edge computing environments with controllable accuracy loss, achieving a qualitative leap from "blind alarms" to "cognitive judgment."
[0136] In one embodiment, step S4: Based on structured assessment information, combined with the current energy state and heterogeneous computing load, an optimization objective function is constructed that includes the benefits of hazard assessment and dynamic energy penalties. The optimal task combination is solved, and the activated tasks are mapped to the corresponding computing units for execution. Specifically, this includes the following steps:
[0137] Step S41: Collect the current energy score and remaining energy at the edge in real time through the energy sensing module, and simultaneously obtain the real-time load, temperature and energy efficiency ratio data of the neural network processor, graphics processor and central processing unit, and output the current energy status and heterogeneous computing load information.
[0138] Specifically, the energy sensing module is integrated into the main control chip of the edge monitoring terminal and is used to monitor the battery voltage and remaining capacity in real time.
[0139] Step S42: Define a task set including lightweight target detection, gimbal control, video encoding and multimodal large model inference, estimate the execution energy consumption vector of each task on different computing units, and construct an optimization objective function that includes the benefit of hazard assessment and dynamic energy penalty.
[0140] Specifically, define a Boolean variable for task execution. ( , For the i-th task, use a Boolean variable to construct an optimization objective function that includes "benefits from hazard assessment" and "energy penalty":
[0141] Equation (18)
[0142] in, To evaluate the benefit function, we characterize the expected accuracy improvement in identifying real hidden dangers under different task execution Boolean variables δ (e.g., activating the VLM task can improve accuracy by 20%, and activating the video encoding task can improve accuracy by 15%). This is a dynamic energy penalty coefficient that is negatively correlated with the current energy score (the lower the energy level, the heavier the penalty). The current energy score at the edge. This is an energy consumption vector. , To reduce the energy consumption per execution of a lightweight target detection task, The energy consumption per execution of a gimbal control task. The energy consumption per execution of a video encoding task. The energy consumption per execution for multimodal large model inference tasks. This represents the remaining energy.
[0143] Step S43: Solve the optimization objective function to obtain the optimal task combination decision for the current time slot, and accurately map the activated tasks to the corresponding computing units for execution based on the decision results.
[0144] Specifically, the lightweight object detection task is mapped to a neural network processor for execution, the visual encoding part of the multimodal large model inference task is mapped to a graphics processor for execution, and the language model inference part is mapped to a central processing unit for execution.
[0145] Step S44: Based on the current energy score's threshold range and the predicted replenishment energy (such as expected solar charging), execute a smooth service level adjustment strategy, dynamically freeze or release high-energy-consuming tasks, and output the final task scheduling instruction set.
[0146] Specifically, when in high power mode, the dynamic energy penalty coefficient λ is extremely small, the optimal task combination δ=[1,1,1,1], and all tasks are activated (lightweight object detection, gimbal control, video encoding, multimodal large model inference).
[0147] When in medium power mode, some tasks are downgraded, and high-power-consuming video encoding tasks are shut down. =0), only gimbal zoom and single-frame VLM inference are retained;
[0148] When in low power mode, trigger deep task degradation, δ=[1,0,0,0], freeze multimodal large model and gimbal control tasks, and only maintain low-frequency YOLO detection and minimal backhaul.
[0149] This embodiment addresses the pain points of heterogeneous computing power preemption and power consumption runaway by constructing a multi-objective optimization function with dynamic energy penalty coefficients at the edge. Specifically, utilizing real-time power data fed back by the energy sensing module, the system can adaptively adjust task combinations: running at full speed when the power is high to ensure the accuracy of hazard assessment, and effectively avoiding system crashes caused by power depletion through smooth degradation strategies (such as disabling video encoding or freezing VLM tasks) when the power is low to medium. This mechanism maximizes the lifespan of passive monitoring terminals while ensuring priority capture of real hazards, solving the industry problem of balancing detection accuracy and battery life in passive monitoring terminals, and making the engineering implementation of multimodal large models in passive power scenarios possible.
[0150] Secondly, such as Figure 2 As shown, this application provides a power edge monitoring system with multimodal collaboration and environmental knowledge enhancement, including a wide-area initial screening module 100, a spatiotemporal verification module 200, a semantic verification and knowledge enhancement module 300, and a heterogeneous scheduling and energy consumption management module 400.
[0151] The wide-area preliminary screening module 100 is used to deploy a heterogeneous computing architecture at the edge and use a lightweight target detection model to perform wide-area preliminary screening of the acquired visible light images.
[0152] The spatiotemporal verification module 200 is communicatively connected to the wide-area initial screening module 100. When a suspected hidden danger target is detected during the initial screening, the spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features.
[0153] The semantic review and knowledge enhancement module 300 is communicatively connected to the spatiotemporal review module 200. It is used to call the lightweight multimodal large model on the end side, combine the high-resolution zoom image and key frame sequence to perform deep semantic review, and introduce environmental knowledge based on retrieval enhancement to dynamically enhance the review process and obtain structured judgment information.
[0154] The heterogeneous scheduling and energy management module 400 is communicatively connected to the semantic review and knowledge enhancement module 300. It is used to construct an optimization objective function that includes the benefits of hidden danger assessment and dynamic energy penalties based on structured assessment information, combined with the current energy state and heterogeneous computing load. The optimal task combination is solved and the activated tasks are mapped to the corresponding computing units for execution.
[0155] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A power edge monitoring method with multimodal collaboration and environmental knowledge enhancement, characterized in that, Includes the following steps: A heterogeneous computing architecture is deployed at the edge, and a lightweight target detection model is used to perform wide-area initial screening of the acquired visible light images; When a suspected potential hazard target is initially identified, a spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features. The lightweight multimodal large model on the client side is invoked, and deep semantic re-judgment is performed by combining the high-resolution zoom image and key frame sequence. Environmental knowledge based on retrieval enhancement is introduced to dynamically enhance the re-judgment process and obtain structured judgment information. The specific steps of dynamically enhancing the review process by introducing environmental knowledge based on retrieval enhancement include: When the cloud confirms that it is a false alarm, the system receives the false alarm environment context knowledge sent by the cloud, vectorizes it and stores it in the lightweight vector database on the edge, and configures a time decay-based weight model for the knowledge item. Before performing the deep semantic re-judgment, the features of the current suspected hidden danger are converted into query vectors, and the most matching environmental knowledge is retrieved from the vector database. The most matching environmental knowledge and basic prompt word template are combined to form enhanced prompt words, and the enhanced prompt words are input into the edge lightweight multimodal large model; Based on structured assessment information, combined with the current energy state and heterogeneous computing load, an optimization objective function is constructed that includes the benefits of hazard assessment and dynamic energy penalties. The optimal task combination is solved, and the activated tasks are mapped to the corresponding computing units for execution. The solution to obtain the optimal task combination specifically includes: Define a Boolean variable for task execution, whereby the tasks include lightweight object detection, gimbal control, video encoding, and multimodal large model inference. Based on the current energy score and remaining energy at the edge, a dynamic energy penalty coefficient is set, and the optimization objective function is solved to obtain the optimal task combination for the current time slot; When the current energy score is lower than a preset threshold, a task degradation strategy is executed: in medium power mode, high-energy-consuming video encoding tasks are turned off, and only gimbal zoom and single-frame VLM inference are retained; in low power mode, deep task degradation is triggered, freezing VLM and gimbal control tasks, and only low-frequency YOLO detection and simplified backhaul are maintained.
2. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 1, characterized in that, The triggering spatiotemporal verification mechanism specifically includes: Based on the coordinates and size of the suspected potential hazard target in the image, calculate the horizontal rotation angle, vertical rotation angle and zoom magnification of the gimbal, and control the gimbal to perform zoom shooting to obtain the high-resolution zoom image. Switch the camera module to video capture mode, calculate the motion intensity of the image through inter-frame difference, dynamically adjust the sampling frame rate based on the motion intensity, and extract the frame with the highest motion intensity from the captured video segment as the key frame sequence.
3. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 2, characterized in that, The calculation of the horizontal rotation angle, vertical rotation angle, and zoom magnification of the gimbal specifically includes: Based on the offset of the target center point relative to the image center, and combined with the horizontal and vertical field of view of the camera, the horizontal rotation angle and the vertical rotation angle are calculated respectively. An exponential smoothing filter is introduced to preprocess the calculated rotation angle to reduce control oscillations; The target zoom factor is calculated based on the ratio of the target frame area to the preset standard area, combined with the maximum optical zoom factor limit.
4. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 3, characterized in that, The formula for calculating the target zoom level is: ; in, For the target zoom magnification, This is the maximum optical zoom magnification. Based on the zoom magnification, and These are the width and height of the target bounding box, respectively. and These are the width and height of the preset standard area, respectively.
5. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 2, characterized in that, The extraction of the frame with the highest motion intensity as the keyframe sequence specifically includes: The difference image is obtained by performing absolute difference operation on two adjacent frames, and then Gaussian filtering is used to denoise the difference image; The percentage of pixels whose difference values exceed a preset threshold in the denoised difference image is statistically analyzed to quantify the motion intensity of the image. The video segment is evenly divided into multiple sub-segments, and the frame with the highest motion intensity in each sub-segment is selected as the keyframe of that sub-segment.
6. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 1, characterized in that, The weighting model based on time decay is as follows: The weight of a knowledge item decays exponentially over time, and its calculation formula is as follows: ; in, Let be the weight of the knowledge item at time t. As the initial weights, The attenuation coefficient is... Time for knowledge creation; When the weight of a knowledge item is lower than the preset retention threshold, expiration cleanup is automatically triggered. When new environmental context knowledge is received, its cosine similarity with existing knowledge is calculated. If the similarity is greater than the preset deduplication threshold, an overwrite update is performed; otherwise, append storage is performed.
7. The power edge monitoring method with multimodal collaboration and environmental knowledge enhancement according to claim 1, characterized in that, The lightweight multimodal large model on the calling side, combined with the high-resolution zoom image and keyframe sequence for deep semantic re-judgment, specifically includes: A visual-language adapter is used to reduce the dimensionality and align the visual features of high-resolution zoom images and keyframe sequences to generate an aligned visual token sequence. The visual token sequence is concatenated with the text token sequence corresponding to the structured prompt words to form a multimodal input sequence; The multimodal input sequence is input into a quantized and compressed language model to generate a JSON-formatted output containing the judgment result, category, confidence level, and reason in an autoregressive manner.
8. A power edge monitoring system with multimodal collaboration and environmental knowledge enhancement, characterized in that, A power edge monitoring method for implementing multimodal collaboration and environmental knowledge enhancement as described in any one of claims 1-7; the system comprises: The wide-area initial screening module is used to deploy a heterogeneous computing architecture at the edge and use a lightweight target detection model to perform wide-area initial screening of the acquired visible light images. The spatiotemporal verification module is communicatively connected to the wide-area initial screening module. When a suspected hidden danger target is detected during the initial screening, the spatiotemporal verification mechanism is triggered to obtain a high-resolution zoom image containing spatial details and a key frame sequence of temporal dynamic features. The semantic review and knowledge enhancement module is connected to the spatiotemporal review module. It is used to call the lightweight multimodal large model on the end side, combine the high-resolution zoom image and key frame sequence to perform deep semantic review, and introduce environmental knowledge generated based on retrieval enhancement to dynamically enhance the review process and obtain structured judgment information. The heterogeneous scheduling and energy management module communicates with the semantic review and knowledge enhancement module. It is used to construct an optimization objective function that includes the benefits of hidden danger assessment and dynamic energy penalties based on structured judgment information, combined with the current energy state and heterogeneous computing load. The optimal task combination is solved and the activated tasks are mapped to the corresponding computing units for execution.
Citation Information
Patent Citations
Multi-modal evolutionary knowledge injection method and system
CN120235247A
Power operation and maintenance risk prediction method based on multi-modal fusion large model
CN120767799A