Image-text multi-modal training reasoning method and system based on dynamic mixing precision
By dynamically adjusting the accuracy of the graphic multimodal model of the power inspection equipment, the problem that static configuration schemes cannot balance efficiency and reliability is solved, and the accuracy and resource matching at different inspection stages is achieved, thereby improving inspection efficiency and endurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID HEBEI ELECTRIC POWER CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
AI Technical Summary
The existing static mixed precision configuration scheme cannot adapt to the dynamic nature of power inspection operations, resulting in a tradeoff between inspection reliability and efficiency. In particular, during high-speed inspections and fine-scale testing, there are problems of excessive computational overhead or insufficient precision.
By acquiring multi-source status information of inspection equipment in real time, the accuracy configuration scheme of each functional module in the large graphic and textual multimodal model is dynamically generated. The accuracy is dynamically matched with the resources by adaptively adjusting the scheme according to the task status, model structure and computing resource status.
While ensuring the reliability of key scenario identification, we balance processing speed and calculation accuracy, improve the efficiency of inspection operations and equipment endurance, and achieve synergistic optimization of identification reliability, operation efficiency and energy utilization.
Smart Images

Figure CN121901844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system automation technology, and in particular to a method and system for training and reasoning using a multimodal graph and text model based on dynamic hybrid precision. Background Technology
[0002] In the field of intelligent power system inspection, the use of drones or robots equipped with multimodal sensors such as visible light and infrared for automatic defect identification has become a key technological means to ensure the safe operation of the power grid. To achieve efficient and accurate defect identification on inspection terminal equipment, complex image-text multimodal deep learning models are typically deployed. However, these models generally have high computational demands and memory consumption, while the inspection equipment itself faces strict limitations in terms of computing power, power consumption, and heat dissipation. To address this contradiction, existing technologies mainly adopt a static mixed-precision inference optimization strategy. This involves assigning different computational precisions (e.g., FP32, FP16, INT8) to each functional module in the model before deployment, aiming to improve inference speed and reduce power consumption while keeping accuracy loss controllable.
[0003] However, power line inspection operations are inherently highly dynamic and multi-stage. The entire inspection process typically includes two significantly different task modes: high-speed corridor inspection and fixed-point detailed inspection. High-speed corridor inspection requires rapid coverage of a wide area of lines with high throughput, emphasizing efficiency and endurance; fixed-point detailed inspection requires close-range, multi-angle detailed examination of specific equipment, demanding extremely high accuracy and reliability in defect identification. Existing static accuracy configuration schemes are unchanging compromise strategies that cannot adaptively adjust to real-time changes in task stages, environmental conditions, and equipment resource availability. If high-precision calculations are used globally to meet the accuracy requirements of detailed inspection, unnecessary computational overhead will occur during the high-speed inspection phase, severely sacrificing inspection efficiency and equipment endurance. Conversely, if lower accuracy is used globally to pursue efficiency, the insufficient model representation capability will significantly increase the risk of missed detections and misjudgments in critical detailed inspection stages, creating potential safety hazards. Summary of the Invention
[0004] This invention provides a method and system for training and inference of image and text multimodal based on dynamic mixed precision, which solves the problem that the reliability and efficiency of inspection cannot be balanced due to static precision configuration scheme.
[0005] In a first aspect, the present invention provides a method for training and inference of image and text multimodal models based on dynamic mixed precision. The method includes: acquiring multi-source state information of inspection equipment during the inspection process in real time, the multi-source state information including task state information, model structure information, and computing resource state information; dynamically generating a precision configuration scheme for each functional module in the large image and text multimodal model under the current inspection state based on the multi-source state information; configuring each functional module in the large image and text multimodal model based on the precision configuration scheme to obtain a large image and text multimodal model with real-time configuration; and obtaining the inspection recognition result through model recognition and inference based on the large image and text multimodal model with real-time configuration and the real-time inspection data of the inspection equipment.
[0006] Secondly, the present invention provides a multimodal training and inference device for graphics and text based on dynamic mixed precision. The device includes: a communication module for acquiring multi-source status information of the inspection equipment during the inspection process in real time, the multi-source status information including task status information, model structure information, and computing resource status information; a processing module for dynamically generating a precision configuration scheme for each functional module in the large graphics and text model under the current inspection state based on the multi-source status information; configuring each functional module in the large graphics and text model based on the precision configuration scheme to obtain a large graphics and text model with real-time configuration; and obtaining the inspection recognition result through model recognition and inference based on the large graphics and text model with real-time configuration and the real-time inspection data of the inspection equipment.
[0007] Thirdly, embodiments of the present invention provide an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.
[0008] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in the first aspect and any possible implementation thereof.
[0009] This invention provides a method and system for training and inference of image and text multimodal systems based on dynamic hybrid precision. This invention acquires and fuses three types of multi-source state information—task state, model structure, and computing resources—in real time to perceive and understand the current inspection scenario and resource boundaries. Based on this, it dynamically generates and executes differentiated precision configuration schemes for each functional module, avoiding the drawbacks of a single static precision configuration. This allows the computational precision of the large image and text multimodal model to accurately match real-time task requirements and equipment status. While ensuring the reliability of key scenario recognition, it intelligently balances processing speed and computational precision for different inspection stages, improving the efficiency of inspection operations and equipment endurance, and achieving synergistic optimization of recognition reliability, operational efficiency, and energy utilization. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the architecture of a graph-text multimodal training and inference system based on dynamic hybrid precision provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a multimodal training and inference method for text and image based on dynamic mixed precision provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of a graph-text multimodal training and inference device based on dynamic mixed precision provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0013] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0014] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.
[0015] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0016] As described in the background section, directly deploying high-performance inference models to actual inspection operations is constrained by the stringent resource limitations of the inspection equipment itself. The airborne or vehicle-mounted computing platforms of drones and inspection robots are subject to strict limitations in terms of computing power, power consumption, memory, and heat dissipation. Furthermore, current advanced multimodal inference models typically have a large parameter scale and computational complexity. If deployed indiscriminately, the high computational overhead of the model will lead to excessively high inference latency, making it difficult to meet real-time requirements. Simultaneously, it will significantly increase power consumption, shorten the runtime of a single operation, and thus reduce inspection efficiency.
[0017] To deploy models at resource-constrained edge environments, existing technologies typically employ inference optimization methods, such as mixed precision. The basic idea is to reduce the computation of precision-insensitive parts during model inference from 32-bit floating-point (FP32) to 16-bit floating-point (FP16) or even 8-bit integer (INT8), thereby accelerating computation and reducing power consumption and memory usage. The mainstream approach is to perform a one-time analysis and configuration of all parts of the model before deployment, forming a fixed mixed-precision inference strategy, i.e., a static precision configuration scheme. The fundamental limitation of the static precision strategy is its inability to adapt to the significant dynamics of power line inspection processes. Power line inspections typically involve multiple phases. In the phase of inspecting a large-scale high-speed line corridor between two towers, the primary goal is to process video streams quickly with high throughput to cover a wider area. In the precision inspection phase, which involves targeted data collection and identification of minor defects near specific towers or substation equipment, the focus is on maximizing identification accuracy. The former emphasizes real-time performance and efficiency, while the latter emphasizes accuracy and reliability. Static accuracy strategies are difficult to adjust with task mode switching and cannot be adaptively optimized based on real-time operating conditions such as ambient light, weather conditions, remaining battery power, and chip temperature.
[0018] Therefore, existing technologies face an inherent contradiction between static, fixed computational optimization strategies and dynamic, variable inspection task requirements when applying high-performance inference models to power grid inspection terminals. In practical applications, it is often difficult to balance sacrificing overall inspection efficiency and endurance to ensure accuracy under worst-case conditions, or prioritizing efficiency at the expense of insufficient accuracy in critical stages, leading to missed detections, misjudgments, and potential safety hazards. This invention provides a technical solution that can perceive the inspection status in real time and adaptively adjust the model's computational accuracy, thereby achieving a dynamic optimal balance between inspection efficiency, endurance, and detection reliability.
[0019] This invention addresses edge intelligent inference models deployed on power line inspection drones and robots. By dynamically adjusting computational precision, it achieves an optimal balance between inference speed, power consumption, and recognition accuracy under various operating conditions, ensuring the reliability of key defect identification while meeting real-time and endurance constraints. A dynamic hybrid precision control mechanism for power line inspection scenarios is constructed, organically combining inspection task status perception, model structure analysis, and computational resource monitoring to form an adaptive precision management solution throughout the entire inference process.
[0020] like Figure 1 As shown, this invention provides an architecture diagram of a graph-text multimodal training and inference system based on dynamic hybrid precision. The system mainly comprises four core components: a state monitoring module 100, a dynamic precision decision-making module 200, a precision execution engine 300, and a closed-loop evaluation and security mechanism 400.
[0021] The status monitoring module 100 is responsible for collecting key information supporting accuracy decisions in real time during model inference or training. First, this module acquires task status information, identifying whether the current inspection is in the high-speed inspection phase of the line corridor or the close-range fine inspection phase, and collects related parameters such as sampling frequency, platform attitude, and scene change characteristics, thereby quantifying the priority weight of task objectives between coverage efficiency and defect identification accuracy. Second, the module analyzes the network topology of the multimodal model, distinguishing different functional modules such as visual encoding, text decoding, and cross-modal fusion, and determining the data modality type currently dominating computation, providing a model-level basis for subsequent accuracy allocation. Simultaneously, the module also monitors the resource status of the edge computing platform in real time, including operating parameters such as memory usage, computing unit load, chip temperature, and remaining power, providing hard constraint boundary conditions for accuracy scheduling.
[0022] The dynamic precision decision module 200, as the core unit of this invention, comprehensively utilizes multi-source information provided by the status monitoring module and generates corresponding precision configuration schemes for different sub-modules and computation paths within the model based on preset strategy rules or a lightweight decision model obtained through offline training. During the inference phase, the decision module determines the computational precision configuration hierarchically based on the current resource status, the type of inspection task, and the precision sensitivity of each network module. When the device is in a resource-constrained state and performing a high-speed inspection task, the feature extraction module, which has high numerical stability and is insensitive to quantization errors, prioritizes the use of low-precision calculation formats such as INT8 or FP16 to improve inference throughput and reduce power consumption. When the device enters a fine-tuning inspection phase near critical equipment or when resources are relatively abundant, the FP32 high-precision calculation format is prioritized for the defect discrimination module and the cross-modal fusion module to maintain overall recognition performance and accuracy. In scenarios involving edge-side fine-tuning or online learning, the decision module also performs phased precision control for the gradient calculation and parameter update stages. In the early stages of training, when gradient magnitudes are large and the model is in a rapid convergence phase, lower precision can be used to accelerate the iteration process. In the later stages of training, when the model enters the fine-tuning phase, the precision of gradient-related calculations is gradually increased to avoid quantization noise drowning out subtle update information and to ensure the model's convergence quality. Simultaneously, the model's principal weights are stored and updated with FP32 precision throughout the entire training process to ensure long-term numerical stability.
[0023] The Precision Execution Engine 300 is responsible for implementing the precision configuration scheme output by the dynamic precision decision module during actual inference and training. This engine parses the configuration results and maps the specified precision to the corresponding operators and network layers within the deep learning runtime framework, ensuring that forward propagation, back propagation, and parameter updates operate collaboratively under the given precision mode. Through fine-grained management of operator-level implementations and runtime scheduling strategies, the engine keeps the switching overhead between different precisions within an acceptable range, avoiding overall performance loss due to frequent switching of numerical formats. Simultaneously, through deep collaboration with the underlying hardware acceleration library and energy management module, it achieves a more refined trade-off and optimization between computing power utilization and energy efficiency.
[0024] To ensure the safety and reliability of dynamic precision adjustment, this invention introduces a closed-loop evaluation and security mechanism 400. During the inspection task execution, the system continuously monitors operational metrics such as inference latency, data throughput, power consumption, and temperature, as well as performance metrics such as target detection recall rate trends and key alarm confidence distributions. When inference performance or detection results significantly deviate from expected thresholds, or when the device's operating status approaches safety boundaries, the system automatically triggers protection strategies, reverting the model precision configuration to a preset conservative scheme, or forcibly enabling high-precision paths in critical scenarios, thereby reducing the risk of missed detections and misjudgments due to excessive precision reduction. Through continuous closed-loop monitoring and strategy correction, adaptive control of the dynamic mixed-precision inference process is achieved.
[0025] This invention achieves synergistic optimization between inference speed, energy consumption control, and detection accuracy in the typical resource-constrained and highly dynamic application scenario of power line inspection. The invention exhibits significant adaptability, dynamically adjusting accuracy configuration based on different inspection stages (high-speed inspection or fine-grained detection), resource states (such as power levels / chip temperature), and model characteristics. In terms of performance improvement, compared to traditional static FP32 solutions, this invention can increase inference throughput by 1.5 to 2 times and reduce power consumption by 30% to 40% during high-speed inspection; during fine-grained inspection, the accuracy of key defect identification remains above 95%, and the false negative rate is reduced by 15%. Furthermore, this invention requires no modification to the hardware platform or model architecture, and can be implemented solely through software-level optimization, resulting in low deployment costs and strong engineering practicality. Compared with traditional static mixed precision solutions, this invention can adaptively adjust to changes in inspection task status, environmental conditions, and equipment operating status. It improves processing throughput and endurance during large-scale high-speed inspections and significantly enhances the stability and reliability of defect identification during fine inspections of critical equipment. This effectively improves the overall intelligence level and security of power inspection services and provides a new technical path for edge intelligent reasoning in resource-constrained scenarios.
[0026] like Figure 2As shown, this embodiment of the invention provides a method for training and inference of image and text multimodal modes based on dynamic mixed precision. The method includes steps A11-A14.
[0027] A11. Real-time acquisition of multi-source status information of inspection equipment during the inspection process.
[0028] In some embodiments, multi-source state information includes task state information, model structure information, and computing resource state information.
[0029] It should be noted that this embodiment of the invention provides a method for adaptively adjusting inference accuracy based on inspection status, applied to an edge intelligent inference system deployed on a power inspection drone. This drone is equipped with an NVIDIA Jetson AGXOrin edge computing platform, configured with 32GB of memory, and has a power consumption budget of 30W. Its multimodal defect detection model is based on the Vision Transformer architecture, including a visual encoder, a text decoder, and a cross-modal fusion module, with approximately 500M model parameters.
[0030] In the offline phase before deployment, an accuracy sensitivity analysis was first performed on the multimodal defect detection model. Inference tests were conducted on each functional module of the model using three accuracy formats: FP32, FP16, and INT8. An accuracy sensitivity index was defined and measured on a standard power line inspection dataset to evaluate the detection accuracy, inference latency, and power consumption levels under different accuracy configurations.
[0031] ; in, For accuracy sensitivity index, Representative model module, This indicates the computational precision, FP32, FP16, or INT8. For module In accuracy The decrease in detection accuracy; To reduce inference latency; and These are preset weighting coefficients, representing the degree of attention paid to accuracy loss and latency, respectively.
[0032] In this embodiment of the invention, based on the calculation results of formula (1), a precision sensitivity spectrum table is established for different modules of the model under different precision levels. The shallow convolutional feature extraction layer of the visual encoder is not sensitive to quantization error. When using INT8 precision, the accuracy only decreases by 0.5%, but the inference speed is increased by 1.8 times and the power consumption is reduced by 35%. The cross-modal fusion module and the defect classification head have higher precision requirements. When using FP16 precision, the accuracy decreases by 2.1%, and when using INT8 precision, the accuracy decreases by 5.3%, which does not meet the business requirements. Based on the above analysis results, a precision sensitivity spectrum table for each module of the model is constructed, and an initial precision configuration strategy library is set, including three preset schemes: high-performance mode (full FP32), balanced mode (encoder FP16 + fusion layer FP32), and energy-saving mode (encoder INT8 + fusion layer FP16).
[0033] During the inspection mission, the status monitoring module 100 collects three types of key information in real time. The task status information acquisition unit analyzes the inspection path planning data and GPS positioning information to identify whether the current inspection is in the high-speed inspection phase of the line corridor or the fine inspection phase of the tower equipment. In the high-speed inspection phase, the drone flies at 10-15 m / s, the camera sampling frequency is 5 fps, and the task priority is coverage efficiency; in the fine inspection phase, the drone hovers or flies at a low speed of less than 2 m / s, the camera sampling frequency increases to 15 fps, and the task priority is defect accuracy identification. The model structure information parsing unit monitors the activation status and computational load distribution of each functional module during the current inference process through a Hook mechanism to determine whether the data modality currently dominating the computation is pure visual input or a combined visual-text input. The resource status information acquisition unit obtains GPU memory usage, SM computing unit utilization, chip temperature, and remaining battery power in real time by calling the system API, with a monitoring frequency of 1 Hz.
[0034] At the 10th minute of a typical inspection task, the real-time information obtained by the status monitoring module was as follows: currently in the high-speed inspection phase, flight speed 12m / s, sampling frequency 5fps; the model is currently processing pure visual input, with the visual encoder accounting for 65% of the total computation; GPU memory usage is 58%, SM utilization is 72%, chip temperature is 61℃, and remaining battery power is 45%.
[0035] A12. Based on multi-source state information, dynamically generate accuracy configuration schemes for each functional module in the large multi-modal model of graphics and text under the current inspection state.
[0036] As one possible implementation, step A12 can be specifically implemented as steps A121-A126.
[0037] A121. Based on the task status information, determine the inspection stage and priority requirements of the current inspection task.
[0038] In some embodiments, the inspection phase includes a high-speed inspection phase or a fine inspection phase; priority requirements include priority requirements for processing speed and recognition accuracy.
[0039] A122. Based on the model structure information, analyze each functional module in the large multimodal model of graphics and text to obtain the analysis results.
[0040] In some embodiments, the functional modules include a visual encoding module, a text understanding module, a cross-modal fusion module, and a classification output module.
[0041] A123. Based on the analysis results and the preset precision sensitivity analysis results, determine the tolerance of each functional module to quantization error.
[0042] A124. Based on computing resource status information, assess the real-time resource stress of inspection equipment.
[0043] In some embodiments, real-time resource scarcity includes chip temperature, remaining battery power, and computing unit utilization.
[0044] A125. Based on the inspection stage, priority requirements, the tolerance of each functional module for quantization error, and the real-time resource shortage, intelligent decision-making is carried out to generate a preliminary accuracy configuration scheme.
[0045] In some embodiments, the accuracy configuration scheme is the calculation accuracy of each functional module; the calculation accuracy includes FP32, FP16 or INT8.
[0046] For example, step A125 can be specifically implemented as steps B1-B5.
[0047] A1. Calculate the resource stress score based on the real-time resource stress level.
[0048] A2. Quantify the inspection phase and priority requirements to generate task priority scores.
[0049] A3. Integrate resource stress score and task priority score, and calculate the comprehensive decision score through a preset weighted decision formula.
[0050] A4. Input the comprehensive decision score into the preset precision configuration strategy library for matching and obtain the matching result.
[0051] In some embodiments, the precision configuration strategy library stores combinations of computational precision for each functional module within different decision intervals.
[0052] A5. Based on the matching results, determine the preliminary accuracy configuration scheme.
[0053] A126. Based on the preliminary accuracy configuration scheme and the preset security threshold, perform security verification to generate the final accuracy configuration scheme.
[0054] For example, the dynamic precision decision module 200 makes a comprehensive judgment based on the status information. The inference precision decision unit first assesses the resource scarcity level, as shown in formula (2).
[0055] ; in, Score the resource scarcity level; For real-time monitoring of chip temperature; This refers to the remaining battery power of the device. and Preset temperature and power weights; and To map temperature and power values to The normalization function for an interval.
[0056] In this embodiment of the invention, the chip temperature is 61°C, which is close to the preset safety threshold of 65°C, and the remaining power is 45%, which is at a medium level. The resource status is determined to be "moderately strained".
[0057] Subsequently, task priority and resource scarcity are combined to calculate the final decision score.
[0058] ; in, For the final decision; Quantify the priority of the current task; The resource stress score is calculated according to formula (2); and These are hyperparameters used to adjust the importance of tasks and resources. When applying them, according to... The range of values is determined, and the optimal precision configuration scheme is selected from the strategy library.
[0059] For example, considering the task type as high-speed inspection, the task priority as coverage efficiency, and the characteristic that the visual encoder is not sensitive to accuracy, the decision unit selects an energy-saving mode from the policy library and generates the following accuracy configuration scheme: The visual encoder's convolutional feature extraction layers, Conv1 to Conv5, use INT8 precision; The attention mechanism layer uses FP16 precision; The cross-modal fusion module uses FP32 precision; The defect classification head uses FP32 precision.
[0060] This configuration scheme ensures high precision for key modules while using low precision for computationally intensive convolutional layers, which is expected to reduce power consumption by 28% and increase inference speed by 1.6 times.
[0061] A13. Based on the precision configuration scheme, configure each functional module in the large image and text multimodal model to obtain the large image and text multimodal model with real-time configuration.
[0062] As one possible implementation, step A13 can be specifically implemented as steps A131-A135.
[0063] A131. Resolution precision configuration scheme to determine the execution precision of each functional module.
[0064] A132. Based on the execution accuracy of each functional module, identify and mark the target network layer that needs to be applied with the corresponding accuracy level in the network structure of the large multimodal graph model.
[0065] A133. Configure the precision of the target network layer, apply the numerical format of each precision level to the forward computation process of each target network layer, and obtain multiple computational flows with different precisions.
[0066] A134. By applying a delay conversion strategy and a hardware coordination mechanism, the scheduling and optimization of multiple configured computational flows of different precisions are performed to obtain the optimized target network layer and computational flow.
[0067] A135. Based on the optimized target network layer and computation flow, generate a large-scale multimodal model with real-time configuration of graphics and text.
[0068] A14. Based on the real-time configured multimodal model of images and text, and the real-time inspection data of the inspection equipment, the inspection recognition result is obtained through model recognition and reasoning.
[0069] As one possible implementation, step A104 can be specifically implemented as steps A141-A145.
[0070] A141. Acquire multimodal inspection data collected in real time by the inspection equipment.
[0071] In some embodiments, multimodal inspection data includes visible light images, infrared thermal images, and textual description information.
[0072] A142. Input the multimodal inspection data into the real-time configured image and text multimodal large model, perform forward calculation and model inference, and obtain the original recognition results output by the model.
[0073] For example, step A142 can be specifically implemented as steps C1-C4.
[0074] C1. Monitor system performance metrics in real time during model inference.
[0075] In some embodiments, system performance metrics include inference latency, data throughput, device power consumption, chip temperature, and confidence level of model output results.
[0076] C2. Based on system performance indicators, conduct a closed-loop evaluation to determine whether the system is in an abnormal state.
[0077] In some embodiments, an abnormal state is defined as an inference delay exceeding a preset threshold, a chip temperature exceeding a safety threshold, or a confidence level below a confidence lower limit.
[0078] C3. If the condition is determined to be normal, the accuracy configuration scheme remains unchanged.
[0079] C4. If an abnormal state is determined, a safety protection strategy will be automatically triggered to achieve adaptive dynamic adjustment of the model accuracy.
[0080] In some embodiments, the security protection strategy includes: switching the accuracy configuration scheme to a preset high-precision conservative scheme, and / or reducing the data input frequency. The high-precision conservative scheme sets the calculation accuracy of each functional module to the highest accuracy.
[0081] A143. Based on the original recognition results output by the model, generate inspection recognition results.
[0082] For example, after receiving the precision configuration scheme, the precision execution engine 300 parses the configuration results, mapping INT8 precision to the five convolutional layers of the visual encoder, FP16 precision to the twelve Transformer Blocks, and FP32 precision to the cross-modal fusion module and the classification head. The precision switching control unit, within the PyTorch framework, uses the torch.cuda.amp automatic precision mixing module and quantization tool to manage operators of different precisions. During model forward propagation, data is automatically converted to INT8 format for computation when passing through convolutional layers, to FP16 format when passing through Transformer layers, and to FP32 format when entering the fusion layer. To reduce format conversion overhead, the switching control unit employs a delayed conversion strategy, maintaining the data format unchanged between consecutive operators of the same precision. The hardware co-optimization unit calls the TensorRT acceleration library to perform kernel fusion optimization on the INT8 convolutional layers and manages the concurrent execution of different precision computation streams through the CUDA event synchronization mechanism. In actual testing, the format conversion overhead accounts for less than 5% of the total inference time.
[0083] For example, after the accuracy execution, the performance monitoring unit continuously monitors system operating metrics. Within 5 minutes of adopting the energy-saving mode configuration, the inference latency decreased from 120ms to 75ms, the data throughput increased from 8.3fps to 13.3fps, the GPU power consumption decreased from 18W to 12W, the chip temperature decreased from 61℃ to 56℃, and the remaining power consumption rate decreased from 0.8% per minute to 0.5%. Meanwhile, the detection accuracy remained at 94.2%, only 0.9 percentage points lower than the 95.1% in full FP32 mode, meeting the operational requirements of the high-speed inspection phase.
[0084] At the 25-minute mark of the inspection mission, the drone reached the vicinity of a 35kV transmission tower and began performing a detailed inspection. The status monitoring module detected a task phase transition: the flight speed decreased to 1.5m / s, the sampling frequency increased to 15fps, and the task priority changed to defect accuracy recognition. Simultaneously, it detected that the current scene involved close-up imaging of insulator strings, and historical data showed a high defect miss rate in this scenario. Based on the task phase change and scene characteristics, the dynamic accuracy decision module determined that improved detection accuracy was needed and adjusted the configuration to a balanced mode, using FP16 accuracy for the visual encoder and FP32 accuracy for the cross-modal fusion module and classification head. The accuracy execution engine responded quickly to the configuration change, switching to the new accuracy mode during the next frame's image inference. The performance monitoring unit detected that after the configuration adjustment, the inference latency increased back to 95ms, but the detection accuracy improved to 96.8%, and the recall rate for insulator self-explosion defects increased from 89.3% to 94.7%, meeting the high accuracy requirements of the detailed inspection phase.
[0085] This invention provides a dynamic hybrid precision-based image and text multimodal training and inference method. By acquiring and fusing three types of multi-source state information—task state, model structure, and computing resources—in real time, it perceives and understands the current inspection scenario and resource boundaries. Based on this, it dynamically generates and executes differentiated precision configuration schemes for each functional module, avoiding the drawbacks of a single static precision configuration. This allows the computational precision of the large image and text multimodal model to accurately match real-time task requirements and equipment status. While ensuring the reliability of key scenario recognition, it intelligently balances processing speed and computational precision for different inspection stages, improving the efficiency of inspection operations and equipment endurance, and achieving synergistic optimization of recognition reliability, operational efficiency, and energy utilization.
[0086] Optionally, the image-text multimodal training and inference method based on dynamic mixed precision provided in this embodiment of the invention further includes steps A21-S23 after step A14.
[0087] A21. Based on the inspection identification results, generate a structured inspection report.
[0088] In some embodiments, the inspection report includes the defect type, location, confidence level, and original data frame index.
[0089] A22. When the inspection and identification results contain key defects or anomalies of a preset category, a real-time alarm message is automatically generated and sent to the monitoring center.
[0090] A23. Based on the type and severity of defects identified in the inspection results, plan a re-inspection path for the inspection equipment.
[0091] Thus, this embodiment of the invention adds closed-loop processing after identification is completed. By automatically generating structured reports, real-time alarms, and intelligent path planning, the intelligent identification results are directly transformed into executable operation and maintenance decisions and actions. This significantly improves the automation level and emergency response speed of power inspection, realizing intelligent processing of the entire process from defect detection to reporting, alarming, and handling, reducing delays and errors caused by manual intervention, and improving the efficiency and security of inspection management.
[0092] Optionally, the image-text multimodal training and inference method based on dynamic mixed precision provided in this embodiment of the invention further includes steps A31-A33 after step A14.
[0093] A31. Record the multi-source state information, precision configuration scheme, system performance indicators, and recognition effect indicators used in this inference task to obtain the task record data.
[0094] A32. Add the data recorded in this task to the historical experience library to obtain the updated historical experience library.
[0095] A33. Decision model for optimizing precision configuration scheme based on updated historical experience database.
[0096] Thus, by systematically recording task data for each inference and updating the experience base, the implementation of this invention enables the decision-making model to continuously learn and fine-tune parameters using historical data. This mechanism endows the system with long-term adaptive evolution capabilities, and the accuracy and scenario adaptability of decisions can be continuously enhanced over time, thereby optimizing the long-term performance and resource utilization efficiency of the overall system and demonstrating the system's intelligent growth.
[0097] Optionally, the image-text multimodal training and inference method based on dynamic mixed precision provided in this embodiment of the invention further includes steps A41-A42.
[0098] A41. Based on the path planning information of the inspection equipment, predict the trend of task phase changes in the future.
[0099] A42. When a transition in a task phase is predicted, generate a transitional mixed-precision configuration scheme.
[0100] In some embodiments, a transitional hybrid precision configuration scheme is used to reduce the precision reconfiguration overhead during task switching.
[0101] Thus, embodiments of the present invention can proactively manage the accuracy switching process by predicting changes in task phases and pre-generating transitional accuracy schemes. This feature effectively avoids performance fluctuations or additional overhead caused by hasty configuration adjustments at the switching critical point during the inspection phase, ensuring the smoothness and real-time performance of system state switching. This further optimizes the fluidity of the dynamic adjustment process and the overall stability of the system, achieving a better seamless transition between changing task requirements.
[0102] Furthermore, to cope with extreme operating conditions, the system has a built-in safety fallback mechanism, the triggering conditions of which are as follows: ; in, A Boolean trigger flag for a safety fallback mechanism; Delay for current inference; The preset maximum delay threshold; Real-time chip temperature; The preset maximum safe temperature threshold; The average detection confidence of recent inference results; This is a preset minimum confidence threshold. When any condition is met, this mechanism is triggered, and the system will enforce the preset, most secure, and most accurate configuration strategy. This is to ensure the stability of the system and the reliability of the results.
[0103] For example, at the 40-minute mark of the inspection mission, the drone encountered a strong backlight environment, resulting in a significant decrease in image quality, while the GPU load rapidly increased due to complex scene processing. The performance monitoring unit detected the following anomalies: The inference latency suddenly increased to 150ms, exceeding the preset threshold of 120ms; The detection confidence levels were generally low, with the average confidence level dropping from 0.92 to 0.68. The GPU SM utilization reached 95%, and the chip temperature rose rapidly to 68°C, exceeding the safe threshold of 65°C.
[0104] The safety protection unit immediately triggered the protection strategy, forcibly reverting the accuracy configuration to full FP32 "high-performance mode" while reducing the data sampling frequency from 15fps to 10fps to alleviate computational burden. After the configuration rollback, although the inference latency increased to 110ms, the detection confidence level rebounded to 0.85, and the chip temperature dropped to 63°C within 3 minutes, restoring the system to safe and stable operation. This safety fallback mechanism effectively avoided the risk of missing critical defects due to excessive accuracy reduction.
[0105] This invention enables adaptive accuracy adjustment of power line inspection drones under different operating conditions. Compared with the traditional static FP32 solution, this method extends the flight time by 35% and expands the data coverage by 40% during high-speed inspection; during fine inspection, it ensures a detection accuracy rate of over 95% and improves the recall rate of critical defects by 5 percentage points, comprehensively improving the efficiency and reliability of power line inspection operations.
[0106] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0107] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.
[0108] Figure 3 A schematic diagram of a graph-text multimodal training inference device based on dynamic mixed precision provided in an embodiment of the present invention is shown. The inference device 50 includes a communication module 51 and a processing module 52.
[0109] The communication module 51 is used to acquire multi-source status information of the inspection equipment in real time during the inspection process. The multi-source status information includes task status information, model structure information and computing resource status information.
[0110] The processing module 52 is used to dynamically generate a precision configuration scheme for each functional module in the large image and text multimodal model under the current inspection state based on multi-source state information; based on the precision configuration scheme, configure each functional module in the large image and text multimodal model to obtain a large image and text multimodal model with real-time configuration; based on the large image and text multimodal model with real-time configuration and the real-time inspection data of the inspection equipment, obtain the inspection recognition result through model recognition reasoning.
[0111] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 60 includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the processor 61 executes the computer program 63, it implements the steps in the above-described method embodiments. Alternatively, when the processor 61 executes the computer program 63, it implements the functions of each module / unit in the above-described device embodiments.
[0112] For example, the computer program 63 may be divided into one or more modules / units, which are stored in the memory 62 and executed by the processor 61 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 63 in the electronic device 60.
[0113] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0114] The memory 62 can be an internal storage unit of the electronic device 60, such as a hard disk or memory of the electronic device 60. The memory 62 can also be an external storage device of the electronic device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 60. Furthermore, the memory 62 can include both internal and external storage units of the electronic device 60. The memory 62 is used to store the computer program and other programs and data required by the terminal. The memory 62 can also be used to temporarily store data that has been output or will be output.
[0115] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A multimodal training and inference method for text and image based on dynamic mixed precision, characterized in that, include: Real-time acquisition of multi-source status information of inspection equipment during the inspection process, including task status information, model structure information and computing resource status information; Based on the multi-source state information, a precision configuration scheme for each functional module in the large image and text multimodal model under the current inspection state is dynamically generated. Based on the aforementioned precision configuration scheme, each functional module in the large image and text multimodal model is configured to obtain a large image and text multimodal model with real-time configuration. Based on the real-time configured multimodal model of images and text, and the real-time inspection data of the inspection equipment, the inspection identification result is obtained through model recognition and reasoning.
2. The image-text multimodal training and inference method based on dynamic mixed precision according to claim 1, characterized in that, The process of dynamically generating accuracy configuration schemes for each functional module in the current inspection state of the large multi-modal image and text model based on the multi-source state information includes: Based on the task status information, the inspection stage and priority requirements of the current inspection task are determined; the inspection stage includes a high-speed inspection stage or a fine inspection stage; the priority requirements include priority requirements for processing speed and recognition accuracy. Based on the model structure information, the functional modules in the large-scale image-text multimodal model are analyzed to obtain the analysis results; the functional modules include a visual encoding module, a text understanding module, a cross-modal fusion module, and a classification output module; Based on the analysis results and the preset precision sensitivity analysis results, the tolerance of each functional module to quantization error is determined. Based on the computing resource status information, the real-time resource stress of the inspection equipment is assessed. The real-time resource stress includes chip temperature, remaining power, and computing unit utilization. Based on the inspection stage, priority requirements, tolerance of each functional module for quantization errors, and real-time resource constraints, intelligent decision-making is performed to generate a preliminary accuracy configuration scheme. The accuracy configuration scheme is the calculation accuracy of each functional module; the calculation accuracy includes FP32, FP16, or INT8. Based on the preliminary accuracy configuration scheme and the preset security threshold, a security check is performed to generate the final accuracy configuration scheme.
3. The image-text multimodal training and inference method based on dynamic mixed precision according to claim 2, characterized in that, The system makes intelligent decisions based on the inspection stage, priority requirements, the tolerance of each functional module for quantization errors, and the real-time resource constraints, generating a preliminary accuracy configuration scheme, including: Based on the real-time resource stress level, a resource stress score is calculated; The inspection stages and priority requirements are quantified to generate task priority scores; The resource stress score and task priority score are combined, and a comprehensive decision score is calculated using a preset weighted decision formula. The comprehensive decision score is input into a preset precision configuration strategy library for matching to obtain the matching result. The precision configuration strategy library stores the calculation precision combination of each functional module in different decision score intervals. Based on the matching results, a preliminary accuracy configuration scheme is determined.
4. The image-text multimodal training and inference method based on dynamic mixed precision according to claim 1, characterized in that, Based on the aforementioned precision configuration scheme, each functional module in the large-scale image-text multimodal model is configured to obtain a real-time configured large-scale image-text multimodal model, including: The accuracy configuration scheme is analyzed to determine the execution accuracy of each functional module; Based on the execution accuracy of each functional module, the target network layer that needs to be applied with the corresponding accuracy level is determined and marked in the network structure of the large graph-text multimodal model. The precision of the target network layer is configured, and the numerical format of each precision level is applied to the forward computation process of each target network layer to obtain multiple computational flows with different precisions. By applying a delay conversion strategy and a hardware coordination mechanism, multiple configured computational flows of different precisions are scheduled and optimized to obtain the optimized target network layer and computational flow. Based on the optimized target network layer and computation flow, a large-scale multimodal model with real-time configuration of graphics and text is generated.
5. The image-text multimodal training and inference method based on dynamic mixed precision according to claim 1, characterized in that, The large-scale multimodal model based on real-time configuration of images and text, and the real-time inspection data from the inspection equipment, are used to obtain inspection recognition results through model recognition and reasoning, including: The inspection equipment acquires multimodal inspection data in real time, including visible light images, infrared thermal images, and text description information. The multimodal inspection data is input into the real-time configured image and text multimodal large model for forward computation and model inference to obtain the original recognition result output by the model. The inspection identification result is generated based on the original identification result output by the model.
6. The image-text multimodal training and inference method based on dynamic mixed precision according to claim 5, characterized in that, The process of inputting the multimodal inspection data into the real-time configured image-text multimodal large model for forward computation and model inference to obtain the original recognition result output by the model includes: During model inference, system performance metrics are monitored in real time, including inference latency, data throughput, device power consumption, chip temperature, and confidence level of model output results. Based on the system performance indicators, a closed-loop evaluation is performed to determine whether the system is in an abnormal state; the abnormal state is when the inference delay exceeds a preset threshold, or the chip temperature exceeds a safety threshold, or the confidence level is lower than the lower confidence limit. If the condition is determined to be normal, the accuracy configuration scheme will remain unchanged. If an abnormal state is detected, a safety protection strategy is automatically triggered to achieve adaptive dynamic adjustment of the model's accuracy. This safety protection strategy includes: switching the accuracy configuration scheme to a preset high-precision conservative scheme, and / or reducing the data input frequency. The high-precision conservative scheme sets the calculation accuracy of each functional module to the highest possible level.
7. The image-text multimodal training and inference method based on dynamic mixed precision according to any one of claims 1 to 6, characterized in that, The process of obtaining the inspection recognition result through model recognition and inference, based on the real-time configured multimodal model of images and text and the real-time inspection data of the inspection equipment, further includes: Based on the inspection identification results, a structured inspection report is generated, which includes defect type, location, confidence level, and original data frame index. When the inspection and identification results contain key defects or anomalies of a preset category, a real-time alarm message is automatically generated and sent to the monitoring center. Based on the defect type and severity in the inspection identification results, a re-inspection path is planned for the inspection equipment.
8. The image-text multimodal training and inference method based on dynamic mixed precision according to any one of claims 1 to 6, characterized in that, The process of obtaining the inspection recognition result through model recognition and inference, based on the real-time configured multimodal model of images and text and the real-time inspection data of the inspection equipment, further includes: Record the multi-source state information, accuracy configuration scheme, system performance indicators, and recognition effect indicators used in this inference task to obtain the task record data; Add the data recorded in this task to the historical experience library to obtain the updated historical experience library; The decision model for optimizing the precision configuration scheme is optimized based on the updated historical experience database.
9. The image-text multimodal training and inference method based on dynamic mixed precision according to any one of claims 1 to 6, characterized in that, The method further includes: Based on the path planning information of the inspection equipment, predict the trend of task stage changes in the future period. When a task phase switch is predicted, a transitional hybrid precision configuration scheme is generated to reduce the precision reconfiguration overhead during task switching.
10. A graph-text multimodal training and inference system based on dynamic mixed precision, characterized in that, The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to invoke and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 9.