Device anomaly identification method, device, medium and product
By combining multimodal sensing data and image segmentation models, accurate anomaly identification of power equipment is achieved, generating detailed anomaly reports. This solves the problem of difficulty in detecting hidden defects in traditional inspection methods, and improves the stability and reliability of equipment operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional manual inspection methods are difficult to detect hidden defects in power equipment. Existing intelligent inspection systems lack multimodal information fusion capabilities, resulting in low detection rates and high false alarm rates for early hidden defects, making it difficult to ensure the stable operation of equipment.
The system acquires device images using multimodal sensing data, segments components using a preset image segmentation model, and generates a device anomaly identification report by analyzing the difference between multimodal positive sample images and the target area images, including anomaly identification results and repair strategies.
It improves the accuracy of equipment anomaly identification, reduces data collection and labeling costs, helps maintenance personnel quickly locate problems and take effective measures, and improves equipment operation stability and reliability.
Smart Images

Figure CN121280852B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of equipment inspection technology, and in particular to methods, equipment, media and products for identifying equipment anomalies. Background Technology
[0002] With the rapid development of ultra-high voltage power grids and the large-scale deployment of smart substations, the safe operation of power equipment has placed unprecedented demands on the accuracy of defect detection. In critical facilities such as substations and DC converter stations, the presence of millimeter-level micro-defects in components of core equipment like main transformers, circuit breakers, disconnectors, and surge arresters (such as insulator strings, operating mechanism contacts, and conductive joints) can trigger partial discharge, overheating, or even equipment failure, threatening the stable operation of the power grid.
[0003] Traditional manual inspection methods rely on visual observation, which makes it difficult to detect hidden defects and is inefficient. In recent years, intelligent inspection systems based on the collaboration of multiple devices such as drones, fixed cameras, and wheeled robots have gradually become popular. However, their defect identification capabilities depend on a massive number of defect samples. In addition, existing methods usually analyze single-modal data in isolation and lack the ability to integrate multimodal information for cross-validation and comprehensive diagnosis, resulting in a low detection rate and a high false alarm rate for early hidden defects. Summary of the Invention
[0004] This application provides a method, device, medium, and product for identifying equipment anomalies, which reduces the cost of data collection and labeling, improves the accuracy of equipment anomaly identification, helps maintenance personnel quickly locate problems and take appropriate measures, reduces the impact of equipment anomalies, and improves the overall operational stability and reliability of equipment.
[0005] In a first aspect, embodiments of this application provide a device anomaly identification method, comprising: acquiring multimodal sensing data corresponding to a device to be detected, the multimodal sensing data including image data; performing component segmentation on the image data based on a preset image segmentation model to obtain a target region image corresponding to the device to be detected; determining a multimodal positive sample image corresponding to the target region image, and determining a difference map between the multimodal positive sample image and the target region image; and determining an anomaly identification report of the device to be detected based on the difference map, the component identification report including anomaly identification results of the device to be detected and a repair strategy corresponding to the anomaly identification results.
[0006] In one possible implementation, determining the difference map between the multimodal positive sample image and the target region image to be detected includes: determining the multimodal differences between the multimodal positive sample image and the target region image to be detected, wherein the multimodal differences include structural differences, texture differences, and depth differences; and determining the difference map between the multimodal positive sample image and the target region image to be detected based on the structural differences, texture differences, and depth differences.
[0007] In one possible implementation, determining the anomaly identification report of the device under test based on the difference map includes: performing weighted calculations on the structural differences, texture differences, and depth differences based on a preset weighting strategy to obtain a comprehensive difference score value corresponding to the device under test; determining a target difference threshold value corresponding to the device under test; and determining the anomaly identification report corresponding to the device under test based on the target difference threshold value and the comprehensive difference score value.
[0008] In one possible implementation, the method further includes: determining the device characteristics corresponding to the device to be tested, and determining a basic difference threshold corresponding to the device to be tested based on the device characteristics; acquiring real-time environmental information corresponding to the device to be tested, and determining an environmental correction factor for the basic difference threshold based on the real-time environmental information; correcting the basic difference threshold based on the environmental correction factor, and using the corrected basic difference threshold as a target difference threshold.
[0009] In one possible implementation, the step of segmenting the image data based on a preset image segmentation model includes: determining segmentation prompt information corresponding to the device to be detected, and guiding the image segmentation model to generate multiple candidate masks corresponding to the image data based on the segmentation prompt information, wherein the segmentation prompt information is used to locate the spatial coordinates and / or bounding boxes of the components of the device to be detected; determining the confidence score value and edge optimization evaluation value corresponding to the multiple candidate masks; determining the target mask from the multiple candidate masks based on the confidence score value and the edge optimization evaluation value, and performing component segmentation on the image data based on the target mask.
[0010] In one possible implementation, determining the segmentation prompt information corresponding to the device to be detected includes: determining the shooting angle corresponding to the image data and the device type corresponding to the device to be detected; determining the segmentation prompt information corresponding to the device to be detected from a preset prompt rule database based on the shooting angle and the device type, wherein the preset prompt rule database is constructed based on the three-dimensional model of the device to be detected and the historical inspection image set corresponding to the device to be detected.
[0011] In one possible implementation, the method further includes: if the confidence scores and / or edge optimization evaluation values corresponding to the plurality of candidate masks do not reach the corresponding preset evaluation thresholds, then a component segmentation quality failure result is obtained; based on the component segmentation quality failure result, a plurality of candidate masks corresponding to the image data are regenerated until the target mask is determined.
[0012] Secondly, embodiments of this application provide a device for identifying device anomalies. The device includes: an acquisition module for acquiring multimodal sensing data corresponding to a device to be detected, the multimodal sensing data including image data; a segmentation module for segmenting the image data into components based on a preset image segmentation model to obtain a target region image corresponding to the device to be detected; a determination module for determining a multimodal positive sample image corresponding to the target region image and determining a difference map between the multimodal positive sample image and the target region image; and an identification module for determining an anomaly identification report of the device to be detected based on the difference map, the component identification report including anomaly identification results of the device to be detected and a repair strategy corresponding to the anomaly identification results.
[0013] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0014] The memory stores computer-executed instructions;
[0015] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0017] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0018] The device anomaly identification method, device, medium, and product provided in this application can comprehensively capture device status information by acquiring multimodal sensing data, including image data, of the device under test. Then, a preset image segmentation model is used to segment the image data into components, accurately locating the target area image to be detected, avoiding interference from irrelevant areas, and improving the targeting of the detection. This allows for the determination of the corresponding multimodal positive sample image of the target area image and the generation of a difference map, which can intuitively present the subtle differences between the component under test and its normal state, effectively identifying potential defects. It also effectively overcomes the dependence of traditional methods on massive defect samples, significantly reducing data acquisition and annotation costs. Finally, an anomaly identification report is generated based on the difference map, providing not only the anomaly identification results but also corresponding repair strategies. This helps maintenance personnel quickly locate problems and take remedial measures, reducing the impact of device anomalies and improving the overall operational stability and reliability of the equipment. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] Figure 1 A schematic diagram illustrating a scenario for device anomaly identification provided in this application;
[0021] Figure 2 Flowchart of the device anomaly identification method provided in this application Figure 1 ;
[0022] Figure 3 Flowchart of the device anomaly identification method provided in this application Figure 2 ;
[0023] Figure 4 A schematic diagram of the device anomaly identification device provided in this application;
[0024] Figure 5 A schematic diagram of the structure of the electronic device provided in this application.
[0025] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0027] First, let me explain the terms used in this application:
[0028] The Segment Anything Model (SAM) is a revolutionary, fundamental image segmentation model. Its core capability lies in its ability to perform high-precision, pixel-level segmentation of any object in an image based on a simple "clue" (such as a point, a bounding box, or a piece of text). Trained on a massive dataset containing billions of masks, the SAM model possesses powerful zero-shot generalization capabilities. This means it can quickly and accurately segment various objects that have never been seen during training without requiring additional training for specific targets, thus significantly reducing the technical barriers and application costs of image segmentation.
[0029] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating the application scenario of device anomaly identification provided in this application, such as... Figure 1 As shown, terminal device 110 acquires multimodal sensing data corresponding to the device under test, including image data. Then, server 120 performs component segmentation on the image data based on a preset image segmentation model, thereby obtaining the image of the target region corresponding to the device under test. Next, server 120 determines the multimodal sample image corresponding to the target region image and determines the difference map between the multimodal positive sample image and the target region image. Finally, server 120 can determine an anomaly identification report for the device under test based on the difference map. This anomaly identification report includes the anomaly identification result of the device under test and the corresponding repair strategy. This achieves accurate identification of anomalies in the device under test.
[0030] in, Figure 1 The terminal device 110 shown can be any terminal device that supports the collection of sensing data, such as a drone, a fixed camera, a wheeled robot, or a wearable device, but is not limited to these. Figure 1The server 120 shown can be, for example, a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. No restrictions are placed on this. The terminal device 110 can communicate with the server 120 via wireless networks such as 3G (third-generation mobile information technology), 4G (fourth-generation mobile information technology), and 5G (fifth-generation mobile information technology). No restrictions are placed on this as well.
[0031] The existing technologies have significant limitations in power equipment defect identification. For example, while image processing methods based on dual-light fusion can combine visible light and infrared data, they are limited by image quality, environmental interference (such as rain, fog, and changes in lighting), and manual parameter adjustments, resulting in high false negative and false positive rates. Commonly used target detection algorithms, such as YOLO and Faster R-CNN, rely on a large number of defect samples, while novel defect samples are scarce, and the models have weak generalization ability, making it difficult to cover unknown defects in equipment. Furthermore, existing technologies lack multimodal data fusion mechanisms, failing to comprehensively characterize equipment status. This leads to insufficient ability to identify subtle anomalies (such as slight contamination of insulators or minor deformation of joints), and significantly reduced detection stability and reliability under complex operating conditions (such as low light, high temperature, and high humidity environments).
[0032] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0033] Figure 2 Flowchart of the device anomaly identification method provided in this application Figure 1 ,like Figure 2 As shown, the process of this device anomaly identification method includes at least steps S201 to S204, which are described in detail below:
[0034] Step S201: Obtain multimodal sensing data corresponding to the device to be detected. The multimodal sensing data includes image data.
[0035] For example, when acquiring multimodal sensing data corresponding to the device under test, a collaborative acquisition system integrating multiple types of sensors, such as visible light, infrared, and ultraviolet, can be deployed. Specifically, the visible light sensor captures surface details, structural contours, and marking information of the device through high-resolution imaging to form image data; the infrared sensor acquires surface temperature distribution data of the device through thermal imaging technology to reflect thermal anomalies and generate infrared thermal images (which belong to image data within multimodal sensing data); and the ultraviolet sensor generates ultraviolet discharge images by detecting ultraviolet light signals generated by partial discharge of the device (also belonging to image data within multimodal sensing data). In addition, other modal data such as acoustic fingerprint sensors to collect operating noise characteristics and vibration sensors to monitor mechanical vibration parameters can be integrated as needed. All sensors must cover key areas of the equipment in terms of spatial layout and avoid mutual interference. During the acquisition process, sensor parameters (such as exposure time and aperture size of visible light sensors, temperature measurement range and sensitivity of infrared sensors, and detection threshold of ultraviolet sensors) need to be dynamically adjusted according to the real-time operating conditions of the equipment (such as load level, ambient temperature, light intensity, etc.). Then, synchronous acquisition technology is used to ensure that the multimodal data are strictly aligned in the time dimension, and finally a multimodal sensing data set of power equipment operating status, including visible light images, infrared thermal images, ultraviolet discharge images and other modal data, is formed.
[0036] Step S202: Based on a preset image segmentation model, the image data is segmented into components to obtain the image of the target region corresponding to the device to be detected.
[0037] For example, the acquired device image is input into a pre-trained image segmentation model. This model uses a deep convolutional neural network structure to extract multi-level features from the input image, parsing layer by layer from low-level edge textures to high-level semantic features, accurately capturing the visual features of each component of the device. Subsequently, the image segmentation model uses the extracted feature maps to perform pixel-level classification and region clustering. It determines the affiliation of each pixel using a sliding window or fully convolutional operation, grouping pixels with similar features (such as color, texture, and shape) into independent regions. Simultaneously, it incorporates prior knowledge of the device (such as typical component sizes and positional relationships) to constrain and optimize the segmentation results, avoiding over-segmentation or incorrect merging. Furthermore, during the segmentation process, the image segmentation model dynamically adjusts the receptive field size to adapt to the recognition needs of components of different scales. Small receptive fields are used for small components to enhance detail capture, while larger receptive fields are used for large components to ensure overall structural integrity. Finally, post-processing steps such as morphological operations (e.g., dilation, erosion) are used to optimize boundary smoothness, remove isolated noise areas, and merge adjacent similar regions based on connected component analysis. This ensures that each image of the target region to be detected accurately corresponds to the actual physical component of the device and has clear boundary contours and a reasonable size range.
[0038] Optionally, during the process of segmenting image data into components based on a preset image segmentation model to obtain the corresponding target region image of the device to be detected, the acquired device image can be input into a preset SAM (SegmentAnything Model) segmentation model. This model achieves efficient and accurate segmentation through its three-module collaborative architecture (image encoder, cue encoder, and mask decoder). Specifically, the image encoder first converts the input image into a high-dimensional feature representation, extracting detailed information such as the device's appearance, structure, and texture. Then, combined with the requirements of the device detection scenario, interactive cues (such as points, boxes, and text) guide the model to focus on key component regions, and the cue encoder maps these interactive information into vectors aligned with image features. Subsequently, the mask decoder integrates image features and cue vectors, generates multiple candidate masks in a short time, evaluates their confidence, and finally outputs the optimal segmentation result, accurately outlining the boundary contours of each component of the device to be detected, such as the coils of a transformer, the contacts of a circuit breaker, and the skirts of an insulator. These segmented regions are the target region images to be detected.
[0039] Step S203: Determine the multimodal positive sample image corresponding to the image of the target region to be detected, and determine the difference map between the multimodal positive sample image and the image of the target region to be detected.
[0040] First, it's important to clarify that constructing a multimodal positive sample database is fundamental to defect identification. Its core is establishing a comprehensive, structurally standardized, and efficiently searchable database of equipment's normal operating status. Specifically, this involves integrating multi-source acquisition devices such as drones, fixed cameras, and wheeled robots to systematically acquire multimodal data covering visible light, infrared, ultraviolet light, and acoustic signatures. The system ensures sample coverage of key equipment types, different viewing angles, seasons, and load conditions, while employing standardized protocols to guarantee data quality and format consistency. Subsequently, the format-consistent data undergoes preprocessing and quality control, including denoising, registration, and standardization, to establish a three-dimensional index structure based on equipment identification, shooting angle, and environmental parameters. This is then combined with hash mapping and index trees for rapid retrieval. Finally, data fusion technology is used to spatiotemporally align and correlate features of the multispectral data, forming a unified cross-modal positive sample benchmark system.
[0041] For example, based on key information such as the device identifier corresponding to the device capturing the image, the shooting angle, and environmental parameters (e.g., temperature, humidity, load level), normal-state multimodal data matching the image of the target area to be detected can be matched from a pre-built multimodal positive sample library. This normal-state multimodal data covers multiple modalities, including visible light images, infrared thermal imaging, ultraviolet discharge detection, and acoustic vibration signals. Subsequently, real-time multimodal data of the target area image to be detected can be compared and analyzed in multiple dimensions with the corresponding multimodal positive sample data. Specifically, in the visible light modality, appearance anomalies are identified through pixel-level difference calculations (e.g., grayscale value differences, texture feature differences) or semantic feature matching (e.g., edge contours, shape structures). In the infrared modality, temperature distribution differences are compared to locate thermal anomaly areas. In the acoustic vibration modality, frequency and amplitude variation characteristics are analyzed to detect mechanical faults. In the ultraviolet modality, abnormal discharge signals are detected. Finally, cross-modal feature fusion technology is used to perform spatiotemporal alignment and correlation analysis on the difference information of each modality, generating a comprehensive difference map containing information such as spatial location, difference type, and difference degree. This map presents the deviation of the image of the target area to be detected from the normal state in an intuitive graphical way.
[0042] Step S204: Determine the anomaly identification report of the device under test based on the difference map. The component identification report includes the anomaly identification result of the device under test and the corresponding repair strategy.
[0043] For example, in-depth analysis of the difference maps is required, and the difference information of each modality is quantitatively evaluated using a preset anomaly threshold model. For instance, in the visible light mode, if the texture difference value of a component in the image exceeds a set threshold, the area is determined to have an appearance anomaly. In the infrared mode, if the temperature difference exceeds the normal operating range of the equipment, it is marked as a thermal anomaly. In the acoustic vibration mode, abnormal fluctuations in frequency or amplitude are identified as mechanical fault signals. In the ultraviolet mode, if the discharge signal intensity exceeds the standard value, it is determined to be a corona or arc fault. Simultaneously, multimodal data fusion analysis is combined to cross-validate the preliminary judgment of a single modality. For example, the appearance anomalies detected by visible light are spatially correlated with the thermal anomaly areas detected by infrared light; if their locations overlap, the confidence of the anomaly judgment is increased. Otherwise, further analysis is needed to determine if there is a possibility of misjudgment. Finally, a list of anomaly identification results is generated by integrating the analysis results of each modality. Furthermore, the list of anomaly identification results clearly indicates the anomaly type (such as appearance damage, overheating, mechanical loosening, partial discharge, etc.), anomaly location (accurate to the equipment component level), and confidence level (the higher the value, the more reliable the anomaly judgment).
[0044] Furthermore, corresponding repair solutions can be matched from a pre-defined repair strategy library based on the anomaly type. For example, for anomalies involving cosmetic damage, repair strategies might include component replacement or surface repair. For thermal anomalies, it's necessary to determine whether operating parameters need adjustment or heat dissipation system maintenance based on equipment load conditions. For mechanical faults, further diagnosis of the specific fault point and development of a maintenance plan are required. For partial discharge anomalies, insulation testing or component replacement needs to be arranged. Finally, the anomaly identification results and corresponding repair strategies are integrated into a complete anomaly identification report. This report covers key information such as basic equipment information, anomaly identification results (type, location, confidence level), repair strategy recommendations, and recommended implementation time, thus providing maintenance personnel with clear and actionable decision-making support.
[0045] In the embodiments provided in this application, component-level precise segmentation of device images is achieved by fusing multimodal sensing data, and difference maps are generated by positive sample comparison. This not only resists environmental interference and modality loss through multi-source data fusion and robust model design, ensuring the stability of anomaly detection, but also achieves accurate identification of minute defects by leveraging high-precision segmentation and difference quantification analysis, and finally outputs a reliable report containing anomaly location and repair strategies.
[0046] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of determining the difference map between the multimodal positive sample image and the target region image to be detected may further include steps S301 and S302, which are described in detail below:
[0047] Step S301: Determine the multimodal differences between the multimodal positive sample image and the image of the target region to be detected. The multimodal differences include structural differences, texture differences, and depth differences.
[0048] Step S302: Determine the difference map between the multimodal positive sample image and the target region image to be detected based on structural differences, texture differences, and depth differences.
[0049] For example, structural features of the target region image can be quantitatively analyzed. Specifically, for structural differences, edge detection algorithms can be used to extract the contour boundaries between positive samples and the target region, and structural parameters such as contour matching degree, shape similarity, and size deviation can be calculated. Simultaneously, spatial geometric information of the target region can be derived using 3D reconstruction techniques or depth estimation models, and the depth data of the positive samples can be compared to identify structural anomalies, positional shifts, or size variations. For texture differences, texture feature extraction methods such as Local Binary Pattern (LBP) or Gray-Level Co-occurrence Matrix (GLCM) can be used to analyze the surface texture distribution, roughness, color gradient, and pattern repeatability of the positive samples and the target region. Then, the degree of texture difference can be quantified by calculating texture similarity indices (such as Structural Similarity Index SSIM and Mean Square Error MSE), with particular attention paid to texture changes caused by surface wear, corrosion, and stains. For depth differences, monocular or multi-view depth estimation models can be used to calculate the 3D depth information of the target region image, and a point-to-point comparison can be performed with the depth data of the positive samples to identify anomalies such as abrupt depth changes, surface unevenness, or spatial positional shifts. Finally, the difference data in terms of structure, texture and depth are spatiotemporally aligned and feature fused, and a comprehensive difference map can be generated by heatmap, difference distribution map or 3D visualization model. The difference map can intuitively present the spatial distribution, type and degree of each modality difference in the form of color gradient, contour line or 3D deformation.
[0050] In the embodiments provided in this application, by quantifying the feature differences between multimodal positive samples and the image to be detected in the structural, texture and depth dimensions and fusing them to generate a difference map, it is only necessary to rely on normal sample data to achieve accurate positioning and visual representation of anomalies such as deformation of device components, material degradation and spatial offset.
[0051] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of determining the anomaly identification report of the device to be detected based on the difference map may further include steps S401 and S402, which are described in detail below:
[0052] Step S401: Based on a preset weighting strategy, the structural differences, texture differences, and depth differences are weighted and calculated to obtain the comprehensive difference score value corresponding to the device to be tested.
[0053] Step S402: Determine the target difference threshold corresponding to the device to be tested, and determine the anomaly identification report corresponding to the device to be tested based on the target difference threshold and the comprehensive difference score.
[0054] For example, a weighting strategy can be determined based on the equipment type, component importance, and historical fault data of the device under test. Then, weight coefficients are set for structural differences, texture differences, and depth differences according to this strategy. For instance, structural differences, which directly affect the mechanical integrity of the equipment, may be given higher weight; texture differences, which reflect changes in surface condition, may be given moderate weight; and depth differences, which involve spatial deformation, may be given appropriate weight. Subsequently, the quantified values of the three differences—for example, the contour matching deviation of structural differences, the structural similarity index of texture differences, and the depth abrupt change amplitude of depth differences—are fused according to preset weights using linear weighting or a nonlinear combination method to calculate the corresponding comprehensive difference score. This comprehensive difference score comprehensively reflects the overall deviation of the multimodal differences of the device under test.
[0055] Optionally, a target difference threshold can be set through statistical analysis of historical positive sample data (such as the upper limit of the 95% confidence interval under a normal distribution), expert experience calibration, or dynamic adaptive adjustment (such as real-time correction based on equipment operating conditions and environmental changes). Then, when the comprehensive difference score exceeds the target threshold, the equipment is determined to be abnormal. Based on the comparison between the comprehensive difference score and the threshold, and combined with the spatial location and difference type information in the difference map, an anomaly identification report corresponding to the device under test is generated. The report includes the anomaly type (such as structural deformation, texture wear, depth unevenness), anomaly location (accurate to the device component level), confidence level (calculated based on multimodal difference consistency and historical data verification), severity assessment (classified according to the difference between the score and the threshold), and corresponding repair strategy suggestions (such as structural repair, surface treatment, depth correction, etc.), providing maintenance personnel with accurate anomaly location and repair guidance.
[0056] In the embodiments provided in this application, a comprehensive score value is generated by fusing structure, texture and depth differences through a preset weighted strategy, and combined with a dynamic threshold determination mechanism, which can adapt to different device characteristics to achieve quantitative evaluation of the anomaly level, thereby accurately outputting an intelligent identification report that includes anomaly location and graded early warning.
[0057] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of the above-mentioned device anomaly identification method may further include steps S501 to S503, which are described in detail below:
[0058] Step S501: Determine the equipment characteristics corresponding to the device to be tested, and determine the basic difference threshold corresponding to the device to be tested based on the equipment characteristics.
[0059] Step S502: Obtain the real-time environmental information corresponding to the device to be tested, and determine the environmental correction factor for the basic difference threshold based on the real-time environmental information.
[0060] Step S503: Correct the basic difference threshold based on the environmental correction factor, and use the corrected basic difference threshold as the target difference threshold.
[0061] For example, an equipment importance classification system can be pre-established. This system can comprehensively consider dimensions such as the functional criticality of the equipment (e.g., whether it directly affects production safety or core processes), the scope of its failure impact (e.g., whether a single equipment failure leads to the shutdown of the entire production line), maintenance costs (including spare parts costs, downtime losses, and maintenance man-hours), and historical failure frequency. Equipment can then be classified into high, medium, and low levels (e.g., core production equipment is high-level, auxiliary equipment is medium-level, and spare or non-critical equipment is low-level). Subsequently, differentiated basic difference threshold strategies can be set for different levels of equipment. For example, for high-level equipment, due to the severe consequences of its failure, stricter threshold standards are required. For instance, the contour matching deviation threshold for structural differences could be set to within 0.5mm, the structural similarity index threshold for texture differences to be above 0.95, and the abrupt change threshold for depth differences to be within 2mm, in order to detect potential anomalies as early as possible. For medium-level equipment, the threshold range can be appropriately relaxed (e.g., the structural deviation threshold can be relaxed to 1mm, the texture structural similarity index threshold to 0.9, and the depth abrupt change threshold to 3mm), to balance the false alarm rate and the false negative rate. For low-end devices, the thresholds are further relaxed (e.g., the structural deviation threshold is set to 2mm, the texture structure similarity index threshold is set to 0.85, and the depth mutation threshold is set to 5mm) to reduce unnecessary maintenance interventions.
[0062] Simultaneously, the classification thresholds can be dynamically calibrated by combining historical operation and maintenance data of the equipment (such as multimodal difference distribution under normal conditions and difference characteristics during typical faults). For example, by statistically analyzing the fluctuation range of structural, texture, and depth differences of advanced equipment during long-term operation, the threshold settings can be optimized. Finally, the classification thresholds are associated with equipment identifiers (such as equipment number and location information) and stored in a threshold library. In the anomaly identification process, the corresponding basic difference threshold is automatically called according to the classification of the equipment to be detected, and the corresponding target difference threshold is generated by combining real-time environmental information (such as temperature, humidity, and load) through a preset environmental correction model (for example, for every 10°C increase in temperature, the structural difference threshold of advanced equipment is relaxed by 0.1mm). This ensures the accuracy and adaptability of the threshold determination and provides differentiated anomaly identification and protection for equipment of different importance.
[0063] Optionally, environmental information can be collected in real time through a sensor network deployed at the equipment site or environmental monitoring station. This environmental information includes ambient temperature, humidity, light intensity, atmospheric pressure, wind speed, rainfall, and equipment load level, combined with equipment operating conditions, including current load rate and operating duration. A comprehensive analysis is then performed based on the environmental information and equipment operating conditions. Specifically, an environmental correction factor can be calculated based on a preset mapping model between the environment and thresholds (such as a linear correction function between temperature and thresholds, or a nonlinear adjustment curve between humidity and thresholds). This environmental correction factor quantifies the impact of environmental changes on the differences in normal equipment status. For example, high-temperature environments may cause thermal expansion of equipment, requiring an appropriate increase in the temperature-related difference threshold to avoid false alarms. Finally, the environmental correction factor and the basic difference threshold are dynamically fused. Dynamic fusion includes weighted superposition, multiplicative adjustment, or piecewise function mapping to generate a target difference threshold that adapts to the real-time environmental state. This target difference threshold retains the core constraints of equipment characteristics while dynamically responding to normal state fluctuations caused by environmental changes, ensuring the accuracy and adaptability of threshold determination during anomaly identification.
[0064] Optionally, in some feasible embodiments, maintenance personnel can manually review reports via mobile terminals. For example, if a structural basis difference value in a circuit breaker contact area exceeds a threshold triggers an alarm, and manual inspection confirms it as a texture anomaly caused by slight oxidation of the contact surface, it is marked as a "true anomaly." If the difference is a false alarm caused by sensor noise, it is marked as a "false alarm." The manual confirmation results are then associated with corresponding difference maps, equipment characteristics, environmental parameters, and other data and stored in a feedback database to form a labeled training sample set. Subsequently, an incremental learning algorithm is used to dynamically optimize the difference threshold model. For example, logistic regression or random forest models can be used to analyze the feature distribution of "true anomaly" samples. For example, the contour matching deviation range of structural differences, the structural similarity index distribution of texture differences, and the abrupt change amplitude of depth differences can be analyzed. Then, the mapping relationship between equipment importance classification and difference threshold is refitted. For example, the texture structure similarity index threshold for advanced equipment in high-temperature environments is dynamically adjusted from 0.95 to 0.93 to reduce the false alarm rate. Simultaneously, the weighting coefficients are adjusted in reverse using false positive samples. For example, the weight of the depth difference dimension, which frequently causes false positives, is reduced. The validation process can also be run automatically according to a preset cycle. For example, 10% of historical samples are selected for blind testing, and the accuracy, recall, and overall performance metrics of the optimized model are calculated. If the overall performance metric improves by more than 5%, the entire difference threshold library is updated.
[0065] In the embodiments provided in this application, by combining the basic difference threshold with the device characteristics and introducing real-time environmental information to dynamically adjust the correction factor, adaptive optimization of the threshold setting is achieved, which effectively improves the accuracy and environmental adaptability of anomaly detection under complex working conditions.
[0066] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of segmenting image data based on a preset image segmentation model may further include steps S601 to S603, which are described in detail below:
[0067] Step S601: Determine the segmentation prompt information corresponding to the device to be detected, and guide the image segmentation model to generate multiple candidate masks corresponding to the image data based on the segmentation prompt information. The segmentation prompt information is used to locate the spatial coordinates and / or bounding boxes of the components of the device to be detected.
[0068] For example, the specific category of the device to be detected (such as transformer, circuit breaker, insulator, etc.) is determined by device type recognition (e.g., image classifier or device identifier resolution), and corresponding standardized positioning rules are retrieved from a pre-built device prior knowledge base. These standardized positioning rules may originate from historical data mining, expert experience summaries, and 3D model analysis, and also include typical spatial coordinate ranges of key components, bounding box size ratios (e.g., the width-to-height ratio range), key point positional relationships (e.g., the offset of the contact center point relative to the overall device), and spatial constraints between components (e.g., a component must be located above another component or within a specific angle range). Subsequently, combined with the preprocessing results of the current image, specifically, the overall device area is bounded by edge detection, and candidate component boxes are labeled by target detection, dynamically calibrating parameters such as coordinate ranges and size ratios in the prior rules. For example, the absolute size of the bounding box is adjusted according to the actual imaging size of the device in the image, or the projection position of the coordinate points is corrected according to the device's tilt angle, generating segmentation prompt information adapted to the current scene. This segmentation prompt information may include data such as coordinate point sets, bounding box parameters, and spatial relationship constraints.
[0069] Optionally, these prompts are encoded into guiding signals that the SAM model can understand. Specifically, coordinate points can be converted into heatmap peaks, where the peak position reflects the component center and the diffusion range reflects positioning uncertainty. The bounding box is then converted into a binary mask region, where the region inside the bounding box is given high weight, and the region outside the bounding box is suppressed. Alternatively, an attention weight matrix can be generated through relational constraints. Subsequently, the encoded prompts and the original image data are input into the SAM model. The SAM model uses a multi-head attention mechanism to focus on the target region in the feature map, suppressing irrelevant background interference, and generates multiple candidate masks based on global semantic consistency (e.g., the hierarchical logic between the component and the device as a whole) and local detail precision (e.g., the balance of edge sharpness).
[0070] Step S602: Determine the confidence score and edge optimization evaluation value corresponding to multiple candidate masks;
[0071] Step S603: Determine the target mask from multiple candidate masks based on the confidence score and edge optimization evaluation value, and perform component segmentation on the image data based on the target mask.
[0072] For example, the accuracy of localization is quantified by comparing the fit between multiple candidate masks and the spatial coordinate point set and bounding box parameters in the prompt information. Specifically, the Euclidean distance between the mask center point and the prompt coordinates, and the overlap ratio between the mask bounding box and the prompt box are used. Then, the inherent reliability of the mask is evaluated by combining the probability map or uncertainty estimate output by the SAM model, such as the average probability value of the mask region and the probability gradient of edge pixels. Simultaneously, similarity analysis between candidate masks eliminates duplicate or redundant results, generating a confidence score. A higher confidence score indicates better mask performance in terms of localization accuracy and model reliability. Subsequently, the mask boundaries are extracted using an edge detection algorithm, and their continuity, sharpness, and fit with the actual boundaries of device components are evaluated to determine the edge-trimming evaluation value corresponding to the multiple candidate masks. For example, the average gradient intensity of edge pixels is calculated to measure sharpness, the number of edge breakpoints is counted to assess continuity, and the fit is quantified by the Hausdorff distance to the device's 3D model or historical standard boundaries to generate an edge optimization evaluation value. The higher the edge optimization evaluation value, the clearer the mask edge is and the more it conforms to the actual outline of the part.
[0073] Then, the confidence score and edge optimization evaluation value are weighted and fused. For example, higher weight can be given to localization accuracy to prioritize segmentation accuracy, or higher weight can be given to edge quality to improve visual effect, thereby generating a comprehensive evaluation score. Candidate masks are then ranked based on this score. The candidate mask with the highest comprehensive evaluation score is selected as the target mask. Part segmentation is then performed on the image data based on the target mask to generate the corresponding binary segmentation result map, where the area inside the mask is marked as the part region and the area outside is marked as the background.
[0074] In the embodiments provided in this application, the spatial coordinates and bounding boxes of device components are accurately located by segmentation prompt information to guide the generation of multiple candidate masks. The optimal target mask is selected by combining confidence scores and edge optimization evaluation values, thereby achieving high-precision positioning and edge detail optimization of component segmentation and significantly improving the reliability and practicality of segmentation results in complex scenarios.
[0075] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of determining the segmentation prompt information corresponding to the device to be detected may further include steps S701 and S702, which are described in detail below:
[0076] Step S701: Determine the shooting angle corresponding to the image data and the device type corresponding to the device to be detected.
[0077] Step S702: Based on the shooting angle and device type, determine the segmentation prompt information corresponding to the device to be detected from the preset prompt rule database. The preset prompt rule database is constructed based on the 3D model of the device to be detected and the historical inspection image set corresponding to the device to be detected.
[0078] For example, edge detection algorithms can be used to extract device contour features. Combined with the spatial projection relationship of the device's 3D model, the angular deviation between the contour edge and the horizontal / vertical direction can be calculated. Alternatively, feature point matching technology can be used to compare the similarity of key points in the image with multiple viewpoint templates of the device's 3D model, thereby determining the shooting viewpoint category and specific angle parameters. Simultaneously, semantic analysis is performed on the image data. An image classifier or a template matching method based on device structural features is used to identify the device type corresponding to the device to be detected in the image, such as a transformer, circuit breaker, or insulator, and to verify the visibility and imaging integrity of the device's key components. Subsequently, based on the shooting viewpoint and device type, corresponding segmentation prompt information is retrieved from a pre-set prompt rule database. This pre-set prompt rule database can be constructed by fusing the spatial coordinate information of the device's 3D model with the actual imaging features of historical inspection image sets, containing standardized positioning rules for different device types under various viewpoints. For example, typical spatial coordinate ranges of key components, bounding box size ratios, key point positional relationships, and spatial constraints between components. For example, for a circuit breaker viewed from the side, the database stores the bounding box size range of the contact area, the offset of the center point relative to the whole device, and the relative positional constraints of the contacts and insulator components.
[0079] During the retrieval and segmentation hint process, the preprocessing results of the current image can be combined to dynamically calibrate parameters such as coordinate range and size ratio in the rules, generating segmentation hint information adapted to the current scene. This segmentation hint information includes the spatial coordinate point set of each key component, bounding box parameters, and spatial relationship constraints between components, ultimately providing accurate guidance signals for subsequent image segmentation models.
[0080] In the embodiments provided in this application, by combining the image shooting perspective and device type, segmentation prompt information is dynamically matched from a rule base constructed based on 3D models and historical inspection data. This achieves accurate adaptation of component positioning and intelligent generation of prompt information in different scenarios, effectively improving the adaptability and accuracy of segmentation of complex device structures.
[0081] Based on the above embodiments, in one exemplary embodiment provided in this application, the specific implementation process of the above-mentioned device anomaly identification method may further include steps S801 and S802, which are described in detail below:
[0082] Step S801: If the confidence scores and / or edge optimization evaluation values corresponding to multiple candidate masks do not reach the corresponding preset evaluation thresholds, then the component segmentation quality is deemed unqualified.
[0083] Step S802: Based on the unqualified component segmentation quality results, regenerate multiple candidate masks corresponding to the image data until the target mask is determined.
[0084] For example, following the embodiments described above, a confidence score and an edge optimization evaluation value can be calculated for each candidate mask. The confidence score comprehensively reflects the degree of matching between the mask and the segmentation prompt information, the inherent reliability of the model, and the differences between candidate masks. The edge optimization evaluation value focuses on the clarity and accuracy of the mask edges; after extracting the boundaries using an edge detection algorithm, its continuity, sharpness, and conformity to the actual contour of the device are calculated. If multiple candidate masks generated based on the initial segmentation prompt information fail to reach their respective preset thresholds after the confidence score and edge optimization evaluation, the component segmentation quality can be determined to be unqualified, and the mask regeneration process can be automatically triggered.
[0085] Specifically, the prompting strategy can be dynamically adjusted based on feedback information such as edge blurring and spatial deviation of unqualified masks. This strategy includes, but is not limited to, increasing the density of spatial prompt points, expanding the bounding box range, or switching to text-based prompt descriptions. Furthermore, a prompt optimizer trained on historical high-quality segmentation data can be used to iteratively refine the initial prompts, generating enhanced segmentation prompts. The optimized prompts are then re-input into the image segmentation model to generate a new round of candidate masks, which are then evaluated for quality again. This process is repeated until the confidence score and edge smoothness of the output mask both meet preset standards, ultimately outputting a qualified target part mask to ensure the reliability of subsequent comparative analysis.
[0086] In the embodiments provided in this application, candidate masks are dynamically screened by setting dual thresholds of confidence and edge optimization, and an iterative optimization mechanism is triggered to regenerate candidate masks when the first segmentation fails. This achieves closed-loop control of component segmentation quality and effectively ensures the reliability and stability of segmentation results in complex scenarios.
[0087] Please see Figure 3 , Figure 3 Flowchart of the device anomaly identification method provided in this application Figure 2 ,like Figure 3As shown, multimodal sensing data corresponding to the device under test is acquired, including image data. Segmentation prompts for the device under test are determined, and the image segmentation model is guided to generate multiple candidate masks corresponding to the image data based on these prompts. The segmentation prompts are used to locate the spatial coordinates and / or bounding boxes of the components of the device under test. Confidence scores and edge optimization evaluation values corresponding to the multiple candidate masks are determined. The target mask is determined from the multiple candidate masks based on the confidence scores and edge optimization evaluation values. If the confidence scores and / or edge optimization evaluation values corresponding to the multiple candidate masks do not reach the corresponding preset evaluation thresholds, a component segmentation quality failure result is obtained. Multiple candidate masks corresponding to the image data are regenerated based on the component segmentation quality failure result until the target mask is determined. Component segmentation is then performed on the image data based on the target mask to obtain the image of the target region corresponding to the device under test.
[0088] Then, the multimodal positive sample image corresponding to the target region image to be detected is determined, and the multimodal differences between the multimodal positive sample image and the target region image to be detected are determined. The multimodal differences include structural differences, texture differences, and depth differences. Based on the structural differences, texture differences, and depth differences, a difference map between the multimodal positive sample image and the target region image to be detected is determined. The structural differences, texture differences, and depth differences are weighted and calculated based on a preset weighting strategy to obtain the comprehensive difference score value corresponding to the device to be detected. The device characteristics corresponding to the device to be detected are determined, and the basic difference threshold corresponding to the device to be detected is determined based on the device characteristics. The real-time environmental information corresponding to the device to be detected is obtained, and the environmental correction factor of the basic difference threshold is determined based on the real-time environmental information. The basic difference threshold is corrected based on the environmental correction factor, and the corrected basic difference threshold is used as the target difference threshold. The anomaly identification report corresponding to the device to be detected is determined based on the target difference threshold and the comprehensive difference score value. For detailed implementation process, please refer to the descriptions in the aforementioned embodiments, which will not be repeated here.
[0089] Figure 4 This is a schematic diagram of the structure of the device anomaly identification device provided in this application, such as... Figure 4As shown, the device anomaly identification device 40 provided in this embodiment includes: an acquisition module 410, used to acquire multimodal sensing data corresponding to the device to be detected, the multimodal sensing data including image data; a segmentation module 420, used to perform component segmentation on the image data based on a preset image segmentation model to obtain a target region image corresponding to the device to be detected; a determination module 430, used to determine the multimodal positive sample image corresponding to the target region image to be detected, and determine the difference map between the multimodal positive sample image and the target region image to be detected; and an identification module 440, used to determine an anomaly identification report of the device to be detected based on the difference map, the component identification report including the anomaly identification result of the device to be detected and the repair strategy corresponding to the anomaly identification result.
[0090] In one possible implementation, the determining module 430 is further configured to determine the multimodal differences between the multimodal positive sample image and the target region image to be detected, the multimodal differences including structural differences, texture differences and depth differences; and to determine the difference map between the multimodal positive sample image and the target region image to be detected based on the structural differences, texture differences and depth differences.
[0091] In one possible implementation, the determining module 430 is further configured to: perform weighted calculations on structural differences, texture differences, and depth differences based on a preset weighting strategy to obtain a comprehensive difference score corresponding to the device to be detected; determine a target difference threshold corresponding to the device to be detected; and determine an anomaly identification report corresponding to the device to be detected based on the target difference threshold and the comprehensive difference score.
[0092] In one possible implementation, the determining module 430 is further configured to: determine the device characteristics corresponding to the device to be tested, and determine the basic difference threshold corresponding to the device to be tested based on the device characteristics; obtain the real-time environmental information corresponding to the device to be tested, and determine the environmental correction factor of the basic difference threshold based on the real-time environmental information; correct the basic difference threshold based on the environmental correction factor, and use the corrected basic difference threshold as the target difference threshold.
[0093] In one possible implementation, the segmentation module 420 is further configured to: determine segmentation prompt information corresponding to the device to be detected; guide the image segmentation model to generate multiple candidate masks corresponding to the image data based on the segmentation prompt information; the segmentation prompt information is used to locate the spatial coordinates and / or bounding boxes of the components of the device to be detected; determine the confidence score and edge optimization evaluation value corresponding to the multiple candidate masks; determine the target mask from the multiple candidate masks based on the confidence score and edge optimization evaluation value; and perform component segmentation on the image data based on the target mask.
[0094] In one possible implementation, the segmentation module 420 is further configured to: determine the shooting angle corresponding to the image data and the device type corresponding to the device to be detected; and determine the segmentation prompt information corresponding to the device to be detected from the preset prompt rule database based on the shooting angle and device type. The preset prompt rule database is constructed based on the three-dimensional model of the device to be detected and the historical inspection image set corresponding to the device to be detected.
[0095] In one possible implementation, the segmentation module 420 is further configured to: if the confidence scores and / or edge optimization evaluation values corresponding to multiple candidate masks do not reach the corresponding preset evaluation thresholds, then obtain a component segmentation quality failure result; and regenerate multiple candidate masks corresponding to the image data based on the component segmentation quality failure result until the target mask is determined.
[0096] The device anomaly identification device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0097] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes at least one processor 510 and a memory 520. Optionally, the device 50 further includes a communication component 530. The processor 510, memory 520, and communication component 530 are connected via a bus 540.
[0098] In a specific implementation, at least one processor 510 executes computer execution instructions stored in memory 520, causing at least one processor 510 to perform the above-described method.
[0099] The specific implementation process of processor 510 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0100] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0101] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0102] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0103] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0104] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0105] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0106] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0107] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0108] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0109] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0110] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0112] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for identifying equipment malfunctions, characterized in that, include: Acquire multimodal sensing data corresponding to the device under test, wherein the multimodal sensing data includes image data; The image data is segmented into components based on a preset image segmentation model to obtain the image of the target region corresponding to the device to be detected. Determine the multimodal differences between the multimodal positive sample image and the image of the target region to be detected, wherein the multimodal differences include structural differences, texture differences, and depth differences; Based on the structural differences, texture differences, and depth differences, a difference map is determined between the multimodal positive sample image and the image of the target region to be detected; Based on the structural differences, texture differences, and depth differences, a comprehensive difference score is obtained for the device under test. Determine the target difference threshold corresponding to the device to be detected, make anomaly judgment based on the target difference threshold and the comprehensive difference score, and combine the spatial location and difference type information in the difference map to determine the anomaly identification report corresponding to the device to be detected. The identification report includes the anomaly identification results of the device under test and the corresponding repair strategies for the anomaly identification results; The method further includes: Determine the device characteristics corresponding to the device to be tested, and determine the basic difference threshold corresponding to the device to be tested based on the device characteristics; Obtain the real-time environmental information corresponding to the device under test, and determine the environmental correction factor for the basic scoring threshold based on the real-time environmental information; The basic difference threshold is corrected based on the environmental correction factor, and the corrected basic score threshold is used as the target difference threshold.
2. The method as described in claim 1, characterized in that, The step of obtaining a comprehensive difference score for the device under test based on the structural differences, texture differences, and depth differences includes: The structural differences, texture differences, and depth differences are weighted and calculated based on a preset weighting strategy to obtain the comprehensive difference score value corresponding to the device under test.
3. The method according to any one of claims 1 to 2, characterized in that, The component segmentation of the image data based on the preset image segmentation model includes: The segmentation prompt information corresponding to the device to be detected is determined, and the image segmentation model is guided to generate multiple candidate masks corresponding to the image data based on the segmentation prompt information. The segmentation prompt information is used to locate the spatial coordinates and / or bounding boxes of the components of the device to be detected. Determine the confidence score and edge optimization evaluation value corresponding to the multiple candidate masks; The target mask is determined from the plurality of candidate masks based on the confidence score and the edge optimization evaluation value, and the image data is segmented based on the target mask.
4. The method as described in claim 3, characterized in that, The step of determining the segmentation prompt information corresponding to the device to be detected includes: Determine the shooting angle corresponding to the image data and the device type corresponding to the device to be detected; Based on the shooting angle and the device type, segmentation prompt information corresponding to the device to be detected is determined from a preset prompt rule database. The preset prompt rule database is constructed based on the three-dimensional model of the device to be detected and the historical inspection image set corresponding to the device to be detected.
5. The method as described in claim 3, characterized in that, The method further includes: If the confidence scores and / or edge optimization evaluation values corresponding to the multiple candidate masks do not reach the corresponding preset evaluation thresholds, the result of unqualified component segmentation quality is obtained. Based on the unqualified segmentation results of the components, multiple candidate masks corresponding to the image data are regenerated until the target mask is determined.
6. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 5.
8. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Power transformation equipment defect detection method based on multi-modal input in unmanned aerial vehicle scene
CN119251587A
Multi-mode monitoring and anomaly detection method, system and equipment of belt conveying system and medium
CN120558607A