Industrial inspection intelligent decision-making method and system based on large and small model collaboration
Through the collaborative architecture of large cloud models and small edge models, the problems of insufficient multimodal information fusion and scarce fault samples in extreme environments are solved, high-precision and low-latency equipment fault detection is achieved, and inspection needs in complex scenarios are met.
Patent Information
- Application Number
- CN202510772377.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-23
AI Technical Summary
Existing intelligent inspection technology lacks multimodal information fusion in extreme environments, fault samples are scarce, and the decision-making mechanism is single and lacks coordination, making it difficult to balance the accuracy and real-time performance of equipment fault detection.
It adopts an architecture that collaborates large cloud models with small edge models, and achieves comprehensive perception and intelligent decision-making of multimodal data through dynamic task planning, cross-modal feature fusion and conditional diffusion model generation technology.
It achieves high-precision and low-latency equipment fault detection in extreme environments, improves the robustness and real-time performance of inspections, and meets the needs of multi-modal fault detection in complex scenarios.
Smart Images

Figure CN120688891A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent inspection technology, and in particular to an industrial inspection intelligent decision-making method and system based on collaboration of large and small models. Background Art
[0002] In the field of industrial production, industries such as mining, oil and gas, and chemicals have long been high-incidence areas for production safety accidents due to their complex process flows, high-risk working environments, and the use of a large number of special equipment. Inspections, as a key link in preventing accidents, ensuring the normal operation of equipment, and ensuring production safety, have a direct and significant impact on the industry's safety situation and operational efficiency.
[0003] From a broader technological perspective, the industry is accelerating towards digitalization and intelligentization. Emerging technologies such as the Internet of Things (IoT), big data, artificial intelligence (AI), and cloud computing are constantly emerging and maturing, injecting a powerful impetus into technological innovation across various industries. Within this macro-trend, the traditional inspection model, which relies solely on manual experience and simple tools, can no longer meet the demands of modern industry for refined and efficient management of production safety. Enterprises are eager to leverage new technologies to achieve intelligent transformation of inspection processes, thereby enhancing their ability to predict and address equipment failures and safety hazards, and thereby reducing the risk of accidents.
[0004] From a technical perspective, the unique production characteristics of industries like mining, oil and gas, and chemicals place extremely specific demands on inspections. For example, mining environments are plagued by harsh conditions such as high dust levels, high humidity, dim underground lighting, and complex terrain. Oil and gas extraction sites face the constant risk of leaks of flammable, explosive, and toxic gases. Chemical production processes involve high temperatures, high pressures, and corrosive substances. These complex environments not only pose a serious threat to the personal safety of inspectors, but also make it difficult for manual inspections to comprehensively and accurately obtain equipment operating data and information on potential hazards. Furthermore, the diverse and complex equipment within these industries requires that some equipment be inspected while in operation, further reducing the efficiency and accuracy of manual inspections.
[0005] Although a variety of automated and even intelligent inspection methods have emerged, existing intelligent inspection technologies mainly rely on a limited single or a few data sources for monitoring and diagnosis. The limitations caused by the lack of multimodal information fusion and fault samples in extreme environments are particularly obvious. Specific shortcomings include:
[0006] 1. Insufficient multimodal perception
[0007] Traditional manual inspections and sensor-based automated monitoring systems mostly collect only single physical quantities, such as visible light images, temperature, pressure, or vibration. They are unable to simultaneously acquire information from multiple modalities, including infrared thermal imaging, acoustics / ultrasound, LiDAR point clouds, gas / chemical composition sensing, and fiber optic strain. Consequently, relying solely on a single modality can easily create blind spots in high-temperature, humid, dusty, or noisy environments, potentially overlooking internal or underlying hidden dangers.
[0008] Although drone / robot inspections can be equipped with optical cameras, infrared thermal imagers, ultrasonic sensors, and lidar, most systems are based on traditional deep learning or rule engines. They lack the ability to jointly model multimodal data such as audio (such as the decibels and spectrum of mechanical noise), ultrasonic signals, and force tactile sensing (such as tiny mechanical changes caused by loose equipment connections), making it difficult to accurately identify the correlation characteristics between multi-source information.
[0009] 2. Failure samples in extreme environments are rare
[0010] In extreme environments like petrochemicals, high-voltage power transmission, and mining, the low probability of actual failures and the high cost of data collection result in extremely limited anomaly samples for training. Existing solutions based on big data or traditional data augmentation methods (such as simple dithering and mirror flipping) can only generate limited image or temperature data, making it difficult to cover multimodal fault types in complex scenarios, such as abnormal gas composition associated with extreme gas leaks, high-frequency, sharp acoustic signals associated with structural crack propagation, and abnormal laser reflection point clouds caused by internal equipment corrosion.
[0011] Existing deep learning solutions rely on large amounts of labeled data and are prone to overfitting when faced with small or zero sample situations (especially professional modes such as acoustic and ultrasonic signals, gas concentration fluctuation curves, etc.), making it impossible to reliably detect and diagnose new or rare defects in actual deployment.
[0012] 3. Single decision-making mechanism and lack of coordination
[0013] Traditional sensor networks and IoT big data platforms focus on threshold alarms or rule reasoning, and often only issue alarms based on temperature and vibration thresholds. They lack the ability to conduct global analysis and complex reasoning on multimodal information such as infrared thermal images, laser point clouds, gas composition, sound spectra, and tactile feedback.
[0014] In drone / robot systems, large models have high deployment costs and large response delays; lightweight models on the edge have the advantage of fast response, but it is difficult to perform joint reasoning on multimodal inputs (such as optical images and ultrasonic and mechanical tactile signals). The two are independent of each other and cannot complement each other, making it difficult to balance accuracy and real-time performance. Summary of the Invention
[0015] In view of the deficiencies of the existing technology, the purpose of the present invention is to provide an industrial inspection intelligent decision-making method and system based on the collaboration of large and small models.
[0016] To achieve the aforementioned object of the invention, the technical solutions adopted by the present invention include:
[0017] In a first aspect, the present invention provides an intelligent decision-making method for industrial inspection based on collaboration of large and small models, which includes:
[0018] Using a large cloud model to divide inspection tasks into multiple subtasks, and perform dynamic task planning based on the subtasks;
[0019] Calling the edge model to execute the subtask and obtain multimodal raw data;
[0020] Using a large cloud model to perform cross-modal feature fusion on the multimodal raw data to obtain cross-modal fusion features;
[0021] A large cloud-based model is used to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
[0022] In a second aspect, the present invention further provides an industrial inspection intelligent decision-making system based on collaboration between large and small models, which includes:
[0023] An analysis agent, which uses a large cloud model to divide the inspection task into multiple subtasks and perform dynamic task planning based on the subtasks;
[0024] A decision agent, configured to call the edge model to execute the subtask and obtain multimodal raw data;
[0025] A perception agent, configured to perform cross-modal feature fusion on the multimodal raw data using a large cloud model to obtain cross-modal fusion features;
[0026] The result reasoning agent is used to use the cloud-based large model to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
[0027] Based on the above technical solution, compared with the prior art, the beneficial effects of the present invention include:
[0028] This invention effectively solves the problems of scarcity of fault samples and perception bias in extreme environments through a comprehensive approach of comprehensive multimodal perception, conditional diffusion model enhancement and cloud-edge collaborative decision-making, taking into account high precision and low latency, bringing new breakthroughs in intelligent inspection.
[0029] The above description is only an overview of the technical solution of the present invention. In order to enable those skilled in the art to more clearly understand the technical means of this application and implement them according to the contents of the specification, the following is an explanation of the preferred embodiments of the present invention with detailed drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a schematic diagram of the overall architecture of an industrial inspection intelligent decision-making method and system provided by a typical implementation case of the present invention;
[0031] Figure 2 This is a flowchart of an intelligent decision-making method for industrial inspection provided by a typical implementation case of the present invention;
[0032] Figure 3 This is a schematic diagram of the process of cross-modal feature fusion in the industrial inspection intelligent decision-making method provided by a typical implementation case of the present invention;
[0033] Figure 4 This is a schematic diagram of the process of generating a training set using a diffusion generation model in an industrial inspection intelligent decision-making method provided by a typical implementation case of the present invention;
[0034] Figure 5 This is a code diagram of a cross-modal attention mechanism in an industrial inspection intelligent decision-making method provided by a typical implementation case of the present invention;
[0035] Figure 6 This is a gated fusion code diagram in an industrial inspection intelligent decision-making method provided by a typical implementation case of the present invention;
[0036] Figure 7 It is a multimodal semantic alignment code diagram in an industrial inspection intelligent decision-making method provided by a typical implementation case of the present invention. DETAILED DESCRIPTION
[0037] Currently, the implementation solutions that are most similar to intelligent inspection in the inspection industry mainly include the following:
[0038] (1) Traditional manual inspection: This is the most basic inspection method. Inspectors carry simple inspection tools (such as thermometers, vibration meters, multimeters, etc.) according to a predetermined route and time period to conduct visual inspections, parameter measurements, and record them. This method is highly dependent on the experience and sense of responsibility of the inspectors, and has problems such as low efficiency, strong subjectivity, irregular data recording, and difficulty in discovering potential hidden dangers. Especially when facing high altitude, narrow, or dangerous environments, the limitations of manual inspections are more significant. Zhang Jun et al. pointed out in "A Review of Research on Intelligent Inspection Robots for Power Plants in China" that traditional manual inspections often cannot guarantee fast and accurate feedback on equipment status in all-weather, high-reliability scenarios, and are prone to missing early fault signals. In addition, data is difficult to standardize, manage, and trace. Li Chengwei and Yue Xiang also discussed the bottlenecks of manual inspection models in data integrity and inspection efficiency in "Research Status and Development Trends of Intelligent Inspection Robots for Power Plants" in "Research Status and Development Trends of Intelligent Inspection Robots for Power Plants".
[0039] (2) Sensor-based automated monitoring system: Temperature sensors, pressure sensors, vibration sensors, gas concentration sensors and other types of sensors are installed at key parts of the equipment (such as transformers, switch cabinets, pipelines, etc.) to collect operating parameters in real time. The data is transmitted to a centralized monitoring platform via wired or wireless networks. The monitoring system and background algorithms analyze the data, diagnose faults and issue warnings. Feng Xiong et al. detailed the application status of various sensors in substation online monitoring and intelligent inspection in "Application status and development trend of intelligent sensor technology in the electric power industry", including infrared temperature sensors, ultrasonic sensors, SF6 gas monitoring sensors, etc., as well as the deep integration of online monitoring data and big data platforms brought about by this. Internationally, Becker et al. proposed a distributed sensor data mining method based on federated learning to ensure data privacy while achieving cross-organizational knowledge sharing. In addition, patent CN119917955A [Keyless multi-platform collection of equipment operation information and generation of maintenance recommendations] proposed a solution for synchronously collecting equipment operation data from multiple management platforms, which can assist monitoring personnel in timely discovering potential problems and formulating maintenance plans.
[0040] (3) UAV or robot inspection: In some high-altitude, dangerous or inaccessible areas, UAVs or inspection robots are being used for operations. In “Automated optical inspection of FAST's reflector surface using drones and computer vision”, Xu Tingfa et al. proposed a drone automatic inspection system integrating high-definition cameras, thermal imagers and deep learning algorithms for the inspection of the reflector panel of the 500-meter Aperture Spherical Radio Telescope (FAST). The system can detect defects on the reflective surface with an accuracy of centimeters, significantly improving the inspection efficiency and accuracy. In “Research on the Application of Intelligent Inspection Robots in Substation Operation and Maintenance”, Yang Zengrong introduced an inspection robot system equipped with multiple sensors (including infrared imaging, visible light cameras and laser radar, etc.). The system can move autonomously in a small space, complete equipment status inspections and generate real-time reports, greatly reducing personnel risks and operation and maintenance costs. In addition, patents CN120017801A and CN119975611A respectively proposed a composite robot system that combines drones and robots to assist intelligent inspections in high-altitude and narrow environments.
[0041] (4) Inspection management system based on the Internet of Things and big data: This system uses the Internet of Things (IoT) technology to connect and synchronize data from multiple sources, such as equipment sensors, inspection personnel's handheld terminals, drones / robots, etc., in real time, to achieve the collection, transmission and sharing of inspection data. By using big data and cloud computing platforms to mine and analyze massive amounts of historical and real-time data, equipment operation models can be established to predict faults in advance and plan the optimal inspection route. In "Base Station Inspection Application Practice Based on Artificial Intelligence", Zhou Feng and Wang Bing elaborated in detail on how to combine drones with edge computing in communication base station scenarios to achieve visual inspection and intelligent alarms for base station equipment, and conduct deep learning and big data analysis on the collected data to provide decision support for intelligent operation and maintenance. In the power industry, patent CN117200060A "An Inspection System and Method Based on the Internet of Things" proposes a method of combining the Internet of Things system with inspection robots, dynamically planning inspection routes according to the status of power equipment, and realizing real-time data collection and analysis of the entire process in the scenario of new energy power plants, providing a complete framework and specification for intelligent inspection.
[0042] Internationally, intelligent inspection technology is developing towards multi-disciplinary and multi-platform integration:
[0043] 1. Integration of sensing and edge computing: The distributed sensor data mining method based on federated learning proposed by Becker et al. provides a new approach for ensuring data privacy while achieving cross-organizational knowledge sharing. At the SPIE conference, Smarsly et al. discussed the feasibility of using quadruped robots for structural health monitoring, demonstrating the trend of extending intelligent inspection to diversified robot platforms.
[0044] 2. Computer Vision Applications: Emerald Insight's research shows that there is significant potential for using computer vision to improve the efficiency of traditional industrial inspections. For example, in oil leak detection scenarios, convolutional neural network (CNN) models have achieved over 99% accuracy in detecting leak locations and can be applied to real-time decision-making and early warning in complex environments.
[0045] In summary, intelligent inspection technology is currently experiencing rapid development both domestically and internationally, with different solutions offering their own advantages and challenges. When selecting a specific solution, it's important to consider the complexity, security requirements, and cost budget of the application scenario (such as power, communications, transportation, and manufacturing). A hybrid application model combining human labor, sensors, robotics, and big data should be considered to achieve more efficient and accurate inspection and maintenance.
[0046] In view of the shortcomings of the prior art, the inventors of this case, after long-term research and extensive practice, have proposed the technical solution of the present invention. The following will further explain this technical solution, its implementation process and principles.
[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0048] The objects of the present invention are:
[0049] Build a multimodal perception system for extreme environments, including visible light images, infrared thermal imaging, acoustic / ultrasonic signals, LiDAR point clouds, gas / chemical composition data, force tactile sensing, and fiber optic strain, to achieve comprehensive perception of equipment status and fault detection in complex environments;
[0050] We propose an extreme data generation technology based on the conditional diffusion model to supplement multimodal fault samples for scenarios with few samples. For example, we generate multimodal synthetic data of extreme gas leaks and acoustic and LiDAR anomaly data simulating structural cracks, improving the model's generalization ability for rare fault types.
[0051] A collaborative decision-making architecture is designed that integrates large cloud models and small edge models. The large cloud model is used to process high-dimensional multimodal information and perform global reasoning, while the small edge model quickly responds to on-site multimodal input, achieving closed-loop collaboration from multimodal perception to intelligent decision-making to meet the inspection requirements for high robustness, high reliability, and low latency in extreme scenarios.
[0052] For the above purpose, see Figure 1 and Figure 2 As shown, an embodiment of the present invention provides an industrial inspection intelligent decision-making method based on large and small model collaboration, which includes the following steps:
[0053] Using a large cloud model to divide inspection tasks into multiple subtasks, and perform dynamic task planning based on the subtasks;
[0054] Calling the edge model to execute the subtask and obtain multimodal raw data;
[0055] Using a large cloud model to perform cross-modal feature fusion on the multimodal raw data to obtain cross-modal fusion features;
[0056] A large cloud-based model is used to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
[0057] Traditional manual inspections are limited by subjective experience, complex environments, and high workloads, making them prone to misjudgments and missed inspections. With the continuous expansion of power grids and the increasing shortage of operation and maintenance personnel, a single manual approach can no longer meet the requirements for safe, timely, and comprehensive inspections. Advanced technologies are urgently needed to achieve intelligent decomposition, dynamic planning, and decision-making diagnosis of inspection tasks, thereby improving inspection efficiency and ensuring safe and stable equipment operation. Complex and long-term inspection tasks place extremely high demands on intelligent decision-making and diagnosis. Inspection tasks not only involve external anomaly detection, but also require the timely detection of hidden dangers such as unqualified insulation inside equipment and overheating of special parts. This technology, supported by large models and multi-agent collaboration, constructs an embodied large-scale model-driven intelligent decision-making and diagnosis system, aiming to achieve task understanding, reasonable decomposition, scheduling and allocation, dynamic planning, and safe alignment of inspection tasks, thereby ensuring efficient, stable, and safe operation of the system.
[0058] In the engineering implementation of the edge computing platform, this invention constructs an innovative architecture that collaborates between a large cloud-based model and a small edge-based model. During local computing, the traditional model of relying solely on cloud-based processing is completely abandoned. The lightweight and responsive nature of the small edge-based model is fully leveraged, and highly real-time computing tasks such as sensor data acquisition, preliminary analysis, and action command generation are deployed on local edge computing nodes. Through the efficient scheduling of computing resources by the small edge-based model, the latency from data processing to command generation is compressed to milliseconds, perfectly adapting to the stringent low-latency requirements of dynamic scenarios such as real-time obstacle avoidance and emergency steering, achieving accurate and rapid path planning. Meanwhile, the large cloud-based model focuses on in-depth analysis of complex data and global strategy optimization. The edge computing unit independently runs the small reinforcement learning model, effectively sharing the computational burden of the central processing unit. Even under complex multi-tasking conditions, the system maintains stable performance. This division of labor and collaboration between the large cloud-based model and the small edge-based model ensures efficient processing of real-time tasks while enabling in-depth mining of complex data, significantly improving the overall utilization of computing resources and the overall operational efficiency of the system.
[0059] In some embodiments, the multimodal raw data includes at least three of visible light features, non-visible light features, acoustic features, mechanical features, radar features, and gas features.
[0060] In some embodiments, the process of acquiring the fusion feature may include:
[0061] A cross-modal attention mechanism is used to align the dimensions of multimodal raw data to obtain attention fusion features;
[0062] Performing gated fusion on the attention fusion features to obtain gated fusion features;
[0063] The gated fusion features are subjected to multimodal semantic alignment, and the obtained aligned features are used as the cross-modal fusion features.
[0064] In some embodiments, the cross-modal attention mechanism may specifically include:
[0065] The original data of each modality is first mapped to a unified feature space through an independent linear transformation layer, and a cross-modal attention matrix is established through dot product similarity calculation;
[0066] The gated fusion specifically includes:
[0067] The attention fusion features are spliced along the channel dimension, input into the two-layer MLP network, and the normalized modal weights are output through the Softmax function;
[0068] fusing the attention fusion features of different channels according to the modality weights;
[0069] The multimodal semantic alignment is a three-layer serial architecture of modality encoder, semantic mapping space and association decoder;
[0070] Among them, the modal encoder adopts a hybrid structure of image processors and non-image processors to convert different modal data into high-dimensional feature vectors; the semantic mapping space is trained through contrastive learning to make the distances between abnormal state-related features close to each other while maintaining the distance from the normal state-related features, forming semantic space features; the association decoder generates a cross-modal explanatory graph based on the semantic space features.
[0071] In this invention, cross-modal fusion is one of the key points. In complex industrial environments, a single sensing mode is difficult to fully capture the status of the equipment. This invention adopts a multi-modal collaborative acquisition strategy, deploying multiple types of sensors to synchronously obtain multi-dimensional information about the equipment operation. The multi-modal sensing configuration mainly includes:
[0072] (1) Optical imaging: Industrial cameras with wide dynamic range technology can still produce stable images in strong and weak light environments. Infrared thermal imagers can also capture infrared images, which are then aligned with visible light images at the pixel level.
[0073] (2) Acoustic sensing: 40kHz ultrasonic sensor detects internal defects (cracks, etc.) of machinery.
[0074] (3) Gas analysis: Multi-channel electrochemical sensor array, detecting 8 gases including H2S / CH4 / O2.
[0075] Of course, the specific data modality is not limited to this. The preferred embodiment of the present invention adopts a three-level fusion strategy. The specific process can be found in Figure 3 shown.
[0076] (1) Cross-modal attention mechanism: The cross-modal attention mechanism is the core component of the present invention to achieve multimodal information fusion. Its design is inspired by the attention allocation mechanism of the human perception system. This technology realizes intelligent feature interaction between different sensor data by establishing a dynamic association model between modalities. This technology adopts the query-key-value (QKV) triple projection architecture to achieve deep feature fusion between modalities. Specifically, each modal feature is first mapped to a unified feature space through an independent linear transformation layer, and a cross-modal attention matrix is established through dot product similarity calculation, and a scaling factor is introduced to stabilize gradient propagation. The innovative use of a multi-head parallel computing mechanism enables the model to capture diverse feature interaction patterns from different subspaces. In actual deployment, the module also integrates residual connection and layer normalization technology, which effectively alleviates the gradient disappearance problem in deep network training. Industrial tests show that this mechanism increases the complementary utilization rate of multimodal features by 62% in the task of equipment surface defect detection, which is significantly better than the traditional serial fusion method.
[0077] The specific code of this mechanism can be found in Figure 5 As shown, but not limited to this, the code diagram is only an example.
[0078] (2) Gated fusion: As the core component of decision-level fusion, the network adopts a three-level processing flow of feature splicing-weight generation-dynamic weighting. In terms of technical implementation, the multimodal features after dimension alignment are first spliced along the channel dimension and input into a two-layer MLP network with 128-dimensional hidden layers, and the normalized modal weights are output through the Softmax function. The dynamic Dropout mechanism (probability 0.2) is innovatively introduced to randomly block some modal connections during training, which significantly improves the robustness of the model. Practical applications show that the network can automatically reduce the weight of the visible light modality interfered by dust (an average reduction of 43%), while improving the decision contribution of infrared and acoustic modalities, reducing the false alarm rate of the system to 3.2% in harsh environments.
[0079] The specific code of this mechanism can be found in Figure 6 As shown, but not limited to this, the code diagram is only an example.
[0080] (3) Multimodal semantic alignment: To address the semantic gap problem of cross-modal features, the network constructs a three-layer architecture of "modal encoder-semantic mapping space-association decoder". The modal encoder uses a hybrid structure of 3D convolution (for image processing) and CNN-LSTM (for time-series acoustic signal processing) to convert different modal data into 256-dimensional feature vectors; the semantic mapping space is trained through contrastive learning to make the visible light edge features, infrared temperature gradient features, and ultrasonic spectrum features of "equipment cracks" have a distance of less than 0.3 (Euclidean distance) in this space, while the distance with normal state features is greater than 1.5; the association decoder generates a cross-modal explanatory map based on the semantic space features.
[0081] The specific code of this mechanism can be found in Figure 7 As shown, but not limited to this, the code diagram is only an example.
[0082] In some embodiments, during the training of the cloud-based large model, some modality connections in the gated fusion are randomly masked.
[0083] In some embodiments, the industrial inspection intelligent decision-making method may further include the following steps:
[0084] According to the abnormal factors, the priorities and / or task parameters of the multiple subtasks divided according to the inspection task in the next round are adjusted.
[0085] In some embodiments, the embodiments of the present invention can also use a diffusion generation model to generate a training set for the collaborative training of the cloud-based large model and the edge small model; wherein, in the image encoding, diffusion modeling, reverse sampling and image decoding stages of the diffusion generation model, the image encoding converts the original input into a discrete latent vector, the forward diffusion perturbs the latent vector to a Gaussian noise state, the reverse sampling traces back to the initial state through the ODE integral path, and the image decoding maps the latent variable of the Gaussian noise state of the VQGAN decoder back to the pixel space to synthesize the scene image.
[0086] In some embodiments, during the generation of the training set, multiple extreme scenario simulations are achieved through diffusion condition control and style transfer: style disturbances of different weather conditions and detailed disturbances about equipment aging are added, and multiple different situations are superimposed to obtain multiple scene images.
[0087] In some embodiments, a diversity adjustment loss is used to adjust the generation process of the training set, and the calculation method of the diversity adjustment loss is expressed as:
[0088] L=(1-m PA )·Sim+m PA (1-Sim)
[0089] Wherein, L represents the diversity adjustment loss, m PA represents the perceptual attention weight, which represents the high-level visual perceptual consistency between the scene image and the original input; Sim represents the sample similarity, which represents the consistency between the scene image and the style template.
[0090] As a typical example of the aforementioned solution, to address the current challenges of scarce and uniformly distributed power inspection image data, as well as the difficulty in acquiring samples for extreme scenarios, we propose an intelligent power scene enhancement technology based on diffusion models. This technology, with a controllable diffusion generation process as its core mechanism, combines key techniques such as image latent variable modeling, style transfer, and topology preservation to synthesize realistic and diverse images of extreme weather conditions and equipment changes, significantly expanding the coverage of existing power image data.
[0091] like Figure 4 As shown (for ease of observation, the present invention Figure 4 The present invention constructs an image generation architecture based on the diffusion probability modeling theory. The core process includes four stages: image encoding, diffusion modeling, reverse sampling, and image decoding.
[0092] Step 1: Image Encoding (VQGAN Encoder)
[0093] The original input, which can be a hand-drawn sketch, map topology, or structural line drawing, is converted into a discrete latent vector using the VQGAN (Vector Quantized Generative Adversarial Network) encoder. VQGAN combines the perceptual loss of GAN with the latent space discretization advantages of VAE, making the generated structure more stable, the encoding expression clearer, and facilitating the diffusion process.
[0094] Step 2: Forward SDE
[0095] By introducing the stochastic differential equation (SDE), the latent variable is gradually disturbed to the Gaussian noise state x T :
[0096] dx=[f(x t )+g 2 (t)h (x t ,t,y,T)]dt+g(t)dW t
[0097] Where f is the drift term, g(t) is the control disturbance intensity, h(x t,t,y,T) introduces time and scene semantics (such as rain, snow, and night) as control variables, making the diffusion process conditionally generative. This mechanism ensures that the model not only generates realistic images, but also generates specific types of extreme power scenarios.
[0098] Step 3: Probability Flow ODE
[0099] In order to improve sampling efficiency and stability, the inverse diffusion path in the form of ODE is adopted:
[0100]
[0101] This process does not require the addition of noise; instead, it simply traces back to the initial state through the ODE integral path, achieving high-quality sampling from noise to clear structure. Compared to traditional diffusion sampling, this method is more efficient and has more controllable image quality.
[0102] Step 4: Decoding and Reconstruction (VQGAN Decoder)
[0103] Finally, the VQGAN decoder maps the processed latent variables back to pixel space, synthesizing high-definition, structurally consistent, and stylishly distinctive power scene images. This decoder inherits the discrete vocabulary information from the encoding stage, which helps reproduce the structural layout of the input.
[0104] On the basis of ensuring a reasonable structure, the present invention further simulates and optimizes the appearance, style, texture, etc. of the image to achieve true restoration of extreme situations.
[0105] (1) Extreme weather and equipment degradation simulation
[0106] Through diffusion condition control + style transfer, a variety of extreme situation simulations can be achieved:
[0107] Severe weather simulation: adding local or global style disturbances such as rain and snow textures, low light conditions (nighttime), and fuzzy wind fields (sand and dust / strong wind);
[0108] Equipment aging and failure: simulate details such as rust, cracks, missing parts, overheating and deformation, and perform visual simulation through texture reconstruction and color mapping;
[0109] Composite scenario modeling: Multiple conditions are superimposed (such as nighttime + heavy rain + aging) to test the model's generalization ability for extremely rare scenarios.
[0110] (2) Balance between style consistency and diversity
[0111] In order to generate images with realistic style without losing the original topological and semantic information, a style consistency optimization strategy is adopted, including:
[0112] Perceptual similarity: maintaining high-level visual perceptual consistency between the source and target images;
[0113] Style loss: Texture transfer is achieved by referring to a specific style template or target atlas;
[0114] Diversity Adjustment Loss:
[0115] L=(1-m PA )·Sim+m PA (1-Sim)
[0116] where m PA To perceive the attention weight, Sim represents the sample similarity, guiding the model to achieve a dynamic balance between diverse generation and style preservation.
[0117] A second aspect of an embodiment of the present invention further provides an industrial inspection intelligent decision-making system based on large and small model collaboration, which includes:
[0118] An analysis agent, which uses a large cloud model to divide the inspection task into multiple subtasks and perform dynamic task planning based on the subtasks;
[0119] A decision agent, configured to call the edge model to execute the subtask and obtain multimodal raw data;
[0120] A perception agent, configured to perform cross-modal feature fusion on the multimodal raw data using a large cloud model to obtain cross-modal fusion features;
[0121] The result reasoning agent is used to use the cloud-based large model to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
[0122] In some embodiments, the analysis agent, decision-making agent, perception agent, and result reasoning agent share the agent's status, confidence, and historical feedback.
[0123] The technical solution of the present invention is further described in detail below through several embodiments and in conjunction with the accompanying drawings. However, the selected embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0124] Example 1
[0125] The present invention proposes a multi-agent collaborative work architecture with clear hierarchy. The carrier of the multi-agent system is a large-parameter language model deployed on a private cloud, which uses its huge Internet-level training data to realize the behavioral reasoning of the entire inspection system. s ,Sense Agent), Analysis Agent (A a,AnalystAgent),Decision Agent (A d ,Decision Agent) and Result Reasoning Agent (A r ,ResultsInference Agent) are composed of four core agents. Each agent has a clear division of labor and a collaborative working mechanism. The specific responsibilities are as follows:
[0126] Perceptual Agent A s :It undertakes the analysis function of multi-source data. Based on the strong semantic understanding ability of the large model, it performs in-depth processing on multi-source heterogeneous data such as input images, location information, sensor data, etc. m}, where x i Represents different types of data. Through different feature extraction networks, multi-source data is multimodally encoded to obtain the feature set F = {f1, f2, ..., f m For example, in time series modeling, for sensor time series data S = {s1, s2, ..., s k} , the data is processed through the time series Transformer to achieve abnormal trend prediction and identify nonlinear disturbances; in terms of images, for image I, after being divided into several blocks (patches), it is sent to the visual large model for analysis.
[0127] Analytical Agent A a :Receive perception agent A s The generated perception understanding is integrated with the power domain ontology knowledge O, and the reasoning accuracy of the large model in professional scenarios is optimized through prompt learning Prompt(X, O), and finally the task description T is formed. desc In the task decomposition process, the analysis agent is responsible for task decomposition, decomposing the original task T into multiple subtasks (t1, t2, ..., t n ), the decomposition process is shown in the formula: (t1, t2,…, t n )=AnalystAgent(T|I,L). i is the i-th subtask, I represents the input image, and L represents the input language prompt.
[0128] Decision-making agent A d :Receive analysis agent A a Provided task information T desc, use the fine-tuned large model tool calling capability to call the small model modules deployed on the end side (such as MobileNet, YOLOv7-tiny, EfficientNet and other lightweight network structures) as needed and in a timely manner to perform specific detection operations. By using model pruning, quantization, distillation and other optimization techniques on the small model, the number of model parameters and computational complexity are reduced to make it suitable for edge deployment. In response to typical problems in power inspection, target detection and semantic segmentation modules are designed, and loss functions (such as cross entropy loss function) are used to train and optimize the model. It is worth noting that these subtasks (t1, t2,…, t n ) is not fixed, but will be dynamically adjusted according to the execution results of the previous task. The adjustment process is shown in formula (6): i =Update(t i ,r1,r2,…,r i-1 ). Combined with the reinforcement learning strategy, define the state space S, action space A, and reward function R, and use the strategy network π(a|s) to score and sort each subtask to form the inspection task queue Q Task , and generate its own sub-agent based on each sub-task, and independently formulate inspection action paths and strategies.
[0129] Result Reasoning Agent A r :For decision agent A d The returned diagnostic results are further reasoned and verified. Based on the abnormal evolution chain model, a directed acyclic graph G = (V, E) is constructed, where the node V represents the abnormal state and the edge E represents the causal relationship between the abnormal states, to generate a possible causal graph. Using a tree-type tracing mechanism, starting from the result node, through the reverse reasoning path P i , realize decision transparency and output the final reasoning result R final .
[0130] Throughout the inference process, an additional agent memory mechanism (Memory) is implemented. Through a state sharing and feedback mechanism, agents share task status, confidence information, and historical feedback. This information sharing forms a circular information flow, supporting the system's closed-loop control and enabling continuous optimization and adjustment of the entire multi-agent collaborative architecture to better complete power inspection tasks.
[0131] The specific steps of the system's operation are:
[0132] Step 1: Multi-source data feature extraction
[0133] Input multi-source heterogeneous data such as images and sensors, use corresponding algorithms such as time series Transformer and visual big model to extract and encode features of different types of data respectively, form a unified feature set, and lay the foundation for subsequent processing.
[0134] Step 2: Task analysis and decomposition
[0135] By integrating ontological knowledge in the power field, the reasoning accuracy of large models is optimized through prompt learning, the original task is decomposed into multiple subtasks, and a task description containing task details and execution requirements is generated.
[0136] Step 3: Dynamic task planning and execution
[0137] A lightweight small model on the call side is used to execute detection tasks, and the model performance is optimized in combination with the loss function. Subtasks are dynamically adjusted according to the execution results, and tasks are sorted using reinforcement learning strategies to form an inspection task queue. Each subtask generates a task agent separately to complete the corresponding task trajectory autonomously.
[0138] Step 4: Result reasoning and verification
[0139] A causal graph is constructed based on the abnormal evolution chain model, and a tree-type tracing mechanism is used to reversely reason the detection results to achieve decision transparency and output the final reasoning conclusion.
[0140] Step 5: Closed-loop collaborative optimization
[0141] Through the "memory dashboard" mechanism, the intelligent physical examination status, confidence level and historical feedback are shared, forming a circular information flow, driving the system closed-loop control, and continuously optimizing the execution effect of inspection tasks.
[0142] Compared with many existing inspection solutions, the key measures implemented by the embodiments of the present invention include at least:
[0143] 1. Multimodal Data Fusion and Perception Framework
[0144] Synchronously collect multiple modal information such as visible light images, infrared thermal imaging, acoustic / ultrasonic signals, lidar point clouds, gas / chemical composition data, force touch and fiber strain; design cross-modal feature alignment algorithms, including time synchronization, spatial calibration and feature mapping, to achieve seamless fusion of multimodal information in the space-time dimension, and provide complete perception input of the environment and device status.
[0145] 2. Diffusion model-assisted multimodal data enhancement technology in extreme environments
[0146] For extreme scenario failures with small or no sample sizes, a multimodal synthetic data generation scheme based on Conditional Diffusion Models is proposed. This scheme constructs environmental condition vectors (such as high temperature, high pressure, corrosion level, severe weather, etc.) and fault feature embedding vectors (such as crack morphology, acoustic spectrum characteristics, gas composition anomalies, etc.), generating a comprehensive dataset in the cloud that matches real extreme environments and covers visible light, infrared, acoustic, LiDAR, gas, and other modalities, providing rich and high-fidelity fault samples for subsequent training.
[0147] Multi-agent decision-making architecture for cloud-edge collaboration
[0148] Construct a multi-agent system consisting of an "analysis agent," a "decision-making agent," and a "reasoning agent," where:
[0149] Analytical Agent (Cloud): Deploys a large-scale multimodal Transformer variant model to perform global joint feature learning and reasoning on the multimodal data generated by the diffusion model and the real-time data collected on site, identifying potential failure modes and generating diagnostic recommendations.
[0150] Decision-making agent (edge): Deploys a small model with a lightweight network structure, focusing on rapid detection and early warning based on multimodal inputs (such as on-site infrared thermal images, acoustic / ultrasonic signals, laser point clouds, etc.) to achieve real-time response and local alarms.
[0151] Reasoning agent (cloud-based posterior): The preliminary results of the edge agent are post-processed and recalibrated on the cloud, and the final comprehensive diagnosis and recommendations are given by combining historical multimodal data with the global environment model to achieve cloud-edge closed-loop collaborative decision-making.
[0152] 4. Dynamic task decomposition and adaptive process optimization
[0153] At the start of an inspection, the analytical agent breaks down the task into several subtasks (such as environmental assessment, key equipment location, and multimodal perception mode selection). Based on the results of previous subtasks (such as on-site gas concentration changes, acoustic anomaly spectra, and LiDAR point cloud occlusion), it dynamically adjusts subsequent subtask strategies (such as prioritizing force and tactile sensing or adjusting LiDAR detection angles), enabling online adaptive optimization for dynamic environments and complex scenarios, improving the automation and robustness of the entire inspection process.
[0154] Edge small model lightweight and accelerated deployment technology
[0155] Optimization technologies such as model pruning, network quantization, and knowledge distillation are used on small edge models to ensure that multi-modal rapid inference can be completed on embedded edge devices (such as mobile robots, drones, and embedded sensor terminals) with only limited computing power, while taking into account high precision and low latency, so that the system can still operate efficiently in scenarios with unstable network or limited computing power.
[0156] Therefore, the embodiments of the present invention have at least the following outstanding advantages:
[0157] (1) Comprehensive multi-modal fault perception capabilities
[0158] The present invention is not only based on visible light and infrared imaging, but also integrates multiple modal information such as acoustic / ultrasonic waves, LiDAR point clouds, gas / chemical sensing, force touch and fiber optic strain. It can maintain stable fault perception performance in extreme environments such as low illumination, strong interference or dust obstruction, which is significantly better than traditional solutions that only rely on single modal data such as optical, temperature or vibration.
[0159] (2) Advantages of few-sample multimodal enhancement in extreme environments
[0160] The conditional diffusion model proposed in this invention can generate synthetic fault samples covering multimodal features in the cloud, simulating acoustic, infrared, LiDAR, gas and other multimodal data in various extreme scenarios such as high temperature, high pressure, corrosion, and air leakage. It greatly compensates for the constraints of the scarcity of measured data on model training, making the small edge model and the large cloud model have stronger generalization ability and robustness when facing rare faults.
[0161] (3) Collaborative decision-making between large cloud models and small multimodal models at the edge
[0162] Compared to deploying only large models (which have high response latency and strong bandwidth dependence) or relying solely on lightweight edge models (which have limited perception range and reasoning capabilities), this invention, through a multi-agent collaborative architecture, organically combines deep inference of multimodal global information in the cloud with rapid detection of multimodal real-time input at the edge. When conducting deep multimodal joint analysis in the cloud, small edge models can capture and warn in advance, feeding back key audio, infrared, LiDAR, gas, and other data to the cloud in real time for fine-tuning, forming a "local warning + cloud recalibration" closed loop, significantly improving overall decision-making accuracy and response speed.
[0163] (4) Adaptive dynamic task processing and system scalability
[0164] This invention features task decomposition and dynamic subtask adjustment capabilities, enabling online optimization of inspection paths and strategies based on on-site conditions (such as acoustic noise levels, sudden changes in gas concentration, and LiDAR point cloud density). This avoids the lag and blind spot issues inherent in traditional static inspection processes in complex scenarios. Furthermore, the modular cloud-edge multimodal architecture can be easily expanded to other industries or new sensing modalities, meeting the needs of subsequent functional upgrades and multi-scenario deployment.
[0165] It should be understood that the above embodiments are merely illustrative of the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent variations or modifications made in accordance with the spirit and substance of the present invention are intended to be encompassed within the scope of protection of the present invention.
Claims
1. An intelligent decision-making method for industrial inspection based on collaboration of large and small models, characterized by: include: Using a large cloud model to divide inspection tasks into multiple subtasks, and perform dynamic task planning based on the subtasks; Calling the edge model to execute the subtask and obtain multimodal raw data; Using a large cloud model to perform cross-modal feature fusion on the multimodal raw data to obtain cross-modal fusion features; A large cloud-based model is used to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
2. The industrial inspection intelligent decision-making method according to claim 1 is characterized in that: The process of obtaining the fusion features includes: A cross-modal attention mechanism is used to align the dimensions of multimodal raw data to obtain attention fusion features; Performing gated fusion on the attention fusion features to obtain gated fusion features; The gated fusion features are subjected to multimodal semantic alignment, and the obtained aligned features are used as the cross-modal fusion features.
3. The industrial inspection intelligent decision-making method according to claim 2 is characterized in that: The cross-modal attention mechanism specifically includes: The original data of each modality is first mapped to a unified feature space through an independent linear transformation layer, and a cross-modal attention matrix is established through dot product similarity calculation; The gated fusion specifically includes: The attention fusion features are spliced along the channel dimension, input into the two-layer MLP network, and the normalized modal weights are output through the Softmax function; fusing the attention fusion features of different channels according to the modality weights; The multimodal semantic alignment is a three-layer serial architecture of modality encoder, semantic mapping space and association decoder; Among them, the modal encoder adopts a hybrid structure of image processors and non-image processors to convert different modal data into high-dimensional feature vectors; the semantic mapping space is trained through contrastive learning to make the distances between abnormal state-related features close to each other while maintaining the distance from the normal state-related features, forming semantic space features; the association decoder generates a cross-modal explanatory graph based on the semantic space features.
4. The industrial inspection intelligent decision-making method according to claim 3 is characterized in that: During the training process of the large cloud model, some modal connections in the gated fusion are randomly shielded.
5. The industrial inspection intelligent decision-making method according to claim 1 is characterized in that: Also includes: According to the abnormal factors, the priorities and / or task parameters of the multiple subtasks divided according to the inspection task in the next round are adjusted.
6. The industrial inspection intelligent decision-making method according to claim 1 is characterized in that: A diffusion generation model is used to generate a training set for collaborative training of the large cloud model and the small edge model; Among them, the diffusion generation model includes image encoding, diffusion modeling, reverse sampling and image decoding stages. The image encoding converts the original input into a discrete latent vector, the forward diffusion perturbs the latent vector to a Gaussian noise state, the reverse sampling traces back to the initial state through the ODE integral path, and the image decoding maps the latent variable of the Gaussian noise state of the VQGAN decoder back to the pixel space to synthesize the scene image.
7. The industrial inspection intelligent decision-making method according to claim 6 is characterized in that: During the generation of the training set, a variety of extreme scenario simulations are achieved through diffusion condition control and style transfer: style disturbances of different weather conditions and detailed disturbances about equipment aging are added, and multiple scene images are obtained by superimposing multiple different situations.
8. The industrial inspection intelligent decision-making method according to claim 7 is characterized in that: A diversity adjustment loss is used to adjust the generation process of the training set. The calculation method of the diversity adjustment loss is expressed as: L=(1-m PA )·Sim+m PA ·(1-Sim) Wherein, L represents the diversity adjustment loss, m PA represents the perceptual attention weight, which represents the high-level visual perceptual consistency between the scene image and the original input; Sim represents the sample similarity, which represents the consistency between the scene image and the style template.
9. An industrial inspection intelligent decision-making system based on collaboration of large and small models, characterized by: include: An analysis agent, which uses a large cloud model to divide the inspection task into multiple subtasks and perform dynamic task planning based on the subtasks; A decision agent, configured to call the edge model to execute the subtask and obtain multimodal raw data; A perception agent, configured to perform cross-modal feature fusion on the multimodal raw data using a large cloud model to obtain cross-modal fusion features; The result reasoning agent is used to use the cloud-based large model to perform reasoning based on the cross-modal fusion features, extract abnormal factors, and formulate new inspection tasks based on the analysis of the abnormal factors.
10. The industrial inspection intelligent decision-making system according to claim 9, characterized in that: The analysis agent, decision-making agent, perception agent, and result reasoning agent share the state, confidence, and historical feedback of the agent.
Citation Information
Patent Citations
Intelligent inspection method and system based on new energy power plant
CN117200060A
Intelligent inspection method and system based on multi-dimensional data fusion analysis
CN119917955A
Unmanned aerial vehicle and mobile robot collaborative inspection system and method for intelligent construction site
CN119975611A
Dynamic comprehensive management method and platform for intelligent security and protection
CN120017801A
Cited By
Industrial process soft measurement modeling method and system for incomplete data
CN121072793A
Crack propagation prediction method based on masking condition diffusion model
CN121211861A
Control capability detection system and method based on deep sea closed environment simulation
CN121300635A
Intelligent inspection vehicle dynamic cooperation method based on multi-mode perception fusion and related equipment
CN121349104A
Dual-mode linkage control method and system for intelligent inspection of power equipment
CN121584549A