Method, device, equipment and medium for autonomous defect decision-making in industrial inspection

Through the combination of multi-spectral fusion terminals and cross-modal deep analysis models, full-process closed-loop autonomous decision-making for industrial equipment defect detection and maintenance strategies is achieved, solving the problems of personnel dependence and multimodal data fragmentation in existing technologies, and improving information utilization and the real-time and accuracy of decision-making.

CN120355270BActive Publication Date: 2025-09-19INSPUR GENERSOFT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510848045.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-19
Estimated Expiration
2045-06-24

Smart Images

  • Figure CN120355270B_ABST
    Figure CN120355270B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, equipment and medium for autonomous defect decision-making for industrial inspections, which belongs to the field of intelligent manufacturing technology. The method for autonomous decision-making on equipment defects includes: using a multi-spectral fusion terminal to collect equipment operation status information of industrial equipment and upload it to a cloud server; using the cross-modal deep analysis model built into the cloud server to perform cross-modal learning on the equipment operation status information according to a dual-stream heterogeneous deep network architecture, learn the joint distribution of visual-text features of industrial equipment, and obtain the defect feature vector of the industrial equipment; perform multi-dimensional matching of the defect feature vector with similar cases in the maintenance knowledge base and screen the disposal plan corresponding to the defect feature vector to obtain the maintenance strategy; send the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel. The present application can solve the problem of multi-modal data fragmentation and the lack of feature-level fusion in the existing technology, resulting in low utilization of cross-modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of intelligent manufacturing technology, and specifically relates to a method, device, equipment and medium for autonomous defect decision-making for industrial inspection. Background Art

[0002] As intelligent manufacturing develops from automation to intelligence, predictive maintenance of industrial equipment has become a core part of intelligent manufacturing. Predictive maintenance of industrial equipment mainly relies on industrial inspection. Currently, industrial inspection mainly adopts a collaborative mode of manual inspection and sensor network (including temperature and humidity / smoke sensors and monitoring equipment, etc.). This mode includes two implementation paths: First, the equipment image is collected on site and sent to the maintenance personnel, relying on manual experience to judge the defect type and formulate a disposal plan; Second, the abnormal area of ​​the equipment is detected through the image model, and maintenance personnel with maintenance experience are notified to carry out repairs. The above method has the following problems: (1) Severe dependence on personnel: Experienced maintenance personnel are required to participate in the entire decision-making process, which has high labor costs and subjective judgment bias, and delayed response time. (2) Broken decision chain: The defect identification results lack intelligent association with the maintenance knowledge base, the reuse rate of historical cases is low, and the "detection-decision" information island is formed.

[0003] The core reason for the above problems lies in the lack of multimodal data fusion. Specifically: First, defect identification and disposal strategies are separated: in existing technologies, image recognition models usually run independently of the maintenance knowledge base, resulting in a separation between defect identification results and disposal strategy generation. For example, although the anomaly detection system based on a large visual model can locate surface defects of equipment, it lacks the ability to mine the causal relationships implicit in historical maintenance cases (such as the mapping relationship between specific defect types and material aging cycles), and needs to rely on additional human experience to complete strategy matching. In addition, data is easily contaminated: at the data level, the detection results of single-modal sensors (such as infrared thermal imaging) are easily affected by environmental interference (such as reflections on the equipment surface and steam obstructions), and multimodal data only uses a simple weighted fusion method, and no cross-modal feature alignment mechanism is established (such as the spatiotemporal correlation between hot spot distribution and vibration spectrum), resulting in an increased model false detection rate.

[0004] The above defects are mainly due to the fragmentation of multimodal data and the lack of feature-level fusion, which leads to low utilization of cross-modal information. This makes it difficult for existing solutions to meet the minute-level response and high-precision decision-making requirements in complex industrial scenarios. Summary of the Invention

[0005] This application provides a defect autonomous decision-making solution for industrial inspections, which realizes autonomous decision-making on equipment defects based on the "end-cloud-library" collaborative architecture, builds a multimodal feature fusion engine and a dynamic knowledge evolution system, and can integrate dynamic knowledge evolution and multimodal deep reasoning to make autonomous decisions. Specifically, by integrating multispectral imaging data such as visible light, infrared, and ultraviolet with equipment operating status parameters, combined with remote large model reasoning and dynamic knowledge graph update mechanism, the full process closed-loop autonomy of equipment defect detection, maintenance strategy matching, and disposal effect evaluation is realized. Through the above method, multimodal data can be fused at the feature level to avoid the fragmentation of multimodal data, improve the utilization rate of cross-modal information, and provide a real-time and precise decision support system for predictive maintenance of industrial equipment.

[0006] According to the first aspect of the present application, an embodiment of the present application provides an autonomous defect decision-making method for industrial inspection, comprising:

[0007] Use multi-spectral fusion terminals to collect equipment operating status information of industrial equipment and upload it to the cloud server;

[0008] Using the cross-modal deep analysis model built into the cloud server, a dual-stream heterogeneous deep network architecture is used to perform cross-modal learning on equipment operating status information. This model learns the joint distribution of visual and textual features of industrial equipment and obtains defect feature vectors for the industrial equipment.

[0009] Based on the hybrid dual-engine reasoning model, multi-dimensional matching is performed between the defect feature vector and similar cases in the maintenance knowledge base, and the corresponding treatment plans of the defect feature vector are screened to obtain the maintenance strategy;

[0010] The defect information and maintenance strategy corresponding to the defect feature vector are sent to the communication terminal of the maintenance personnel.

[0011] Preferably, the above-mentioned method for autonomous decision-making on equipment defects further comprises, after the step of sending the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel:

[0012] Obtain maintenance score signals for industrial equipment;

[0013] Using knowledge distillation technology, we use maintenance score signals to adjust the node weights of the maintenance knowledge base corresponding to the dynamic knowledge graph and establish a score-weight mapping relationship.

[0014] The maintenance knowledge base is optimized using the score-weight mapping relationship.

[0015] Preferably, in the above-mentioned method for autonomous decision-making on equipment defects, the step of using a multispectral fusion terminal to collect equipment operating status information of industrial equipment and uploading it to a cloud server includes:

[0016] Use multispectral fusion terminals to collect multispectral status images of industrial equipment;

[0017] Use 3D profile calibration algorithms to compare the spatial coordinates of industrial equipment in multispectral status images with the corresponding spatial features in the preset CAD model in real time;

[0018] Combined with the gyroscope attitude compensation mechanism, the spatial coordinates of the industrial equipment are calibrated;

[0019] The multispectral status image after spatial coordinate calibration is filtered and denoised to obtain the equipment operation status information of the industrial equipment.

[0020] Preferably, in the above-mentioned method for autonomous decision-making on equipment defects, the steps of using a cross-modal deep analysis model built into a cloud server to perform cross-modal learning on equipment operating status information based on a dual-stream heterogeneous deep network architecture, learning the joint distribution of visual-text features of industrial equipment, and obtaining a defect feature vector of the industrial equipment include:

[0021] A cross-modal deep analysis model that uploads the equipment operating status information of industrial equipment to a cloud server. The cross-modal deep analysis module has a dual-stream heterogeneous deep network architecture. The equipment operating status information includes multispectral status images and equipment status data.

[0022] Control the visual processing flow of the dual-stream heterogeneous deep network architecture, use the target detection algorithm combined with the multi-scale attention mechanism to establish feature associations of multispectral state images at multiple scales, and detect the visual features of multispectral state images;

[0023] Control the text processing flow of the dual-stream heterogeneous deep network architecture, use the text processing large model to build a domain semantic parsing dictionary, and use the domain semantic parsing dictionary to parse the text features of the device status image;

[0024] Use a cross-modal adversarial distillation pipeline to adversarially collect visual and text features;

[0025] The generator network constructed using the cross-modal adversarial distillation pipeline learns the joint distribution of visual features and text features according to the alignment loss function to obtain the defect feature vector.

[0026] Preferably, in the above-mentioned device defect autonomous decision-making method, the steps of controlling the visual processing stream of the dual-stream heterogeneous deep network architecture, using the target detection algorithm combined with the multi-scale attention mechanism, establishing feature associations of the multispectral state image in the dimensions of multiple scales, and detecting and obtaining the visual features of the multispectral state image include:

[0027] The improved ViT-Transformer algorithm is used to control the visual processing flow. A deformable convolution layer is introduced to dynamically perceive the convolution kernel offset of the multispectral state image and capture device defects in the multispectral state image.

[0028] Introducing a multi-scale attention mechanism into the visual processing flow, establishing feature associations of multi-spectral state images at multiple scales to identify fuzzy defects in industrial equipment;

[0029] Obtain visual features corresponding to image defects and blur defects.

[0030] Preferably, in the above-mentioned autonomous decision-making method for equipment defects, the steps of performing multi-dimensional matching of defect feature vectors with similar cases in a maintenance knowledge base and screening treatment plans corresponding to the defect feature vectors to obtain a maintenance strategy according to a hybrid dual-engine reasoning mode include:

[0031] Call the dynamic knowledge graph corresponding to the maintenance knowledge base to perform multi-dimensional matching between the defect feature vector and similar cases, and calculate the cosine similarity between the defect feature and similar cases based on the multi-dimensional matching degree;

[0032] extracting a predetermined number of maintenance cases whose cosine similarity is above a similarity threshold;

[0033] as well as,

[0034] Based on fuzzy logic, an expert experience decision tree is constructed. The expert experience decision tree is used to process the uncertain conditions corresponding to the maintenance case through the fuzzy membership function to obtain the disposal plan.

[0035] Integrate maintenance cases and disposal plans to obtain maintenance strategies.

[0036] Preferably, the above-mentioned method for autonomous decision-making on equipment defects further comprises, after the step of sending the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel:

[0037] After industrial equipment is repaired, the operating parameters of the industrial equipment are continuously collected;

[0038] Input the operating parameters into the health index model to obtain the improvement rate of industrial equipment after maintenance;

[0039] Use improvement rate to evaluate the maintenance staff's maintenance effect;

[0040] When the improvement rate is less than or equal to a predetermined improvement threshold, a parameter update mechanism of the cross-modal deep analysis model is triggered;

[0041] Obtaining a maintenance scoring signal uploaded by a communication terminal;

[0042] Using maintenance scoring signals, the node weights of the dynamic knowledge graph corresponding to the maintenance knowledge base are iteratively updated through knowledge distillation technology.

[0043] According to the second aspect of the present application, the present application also provides a defect autonomous decision-making device for industrial inspection, comprising:

[0044] The information collection module is used to collect the equipment operation status information of industrial equipment using a multi-spectral fusion terminal and upload it to the cloud server;

[0045] a cross-modal learning module, configured to use a cross-modal deep analysis model built into a cloud server to perform cross-modal learning on the equipment operating status information based on a dual-stream heterogeneous deep network architecture, learn the joint distribution of visual and textual features of the industrial equipment, and obtain a defect feature vector for the industrial equipment;

[0046] A strategy learning module is used to perform multi-dimensional matching between the defect feature vector and similar cases in the maintenance knowledge base based on a hybrid dual-engine reasoning mode, and to screen the treatment plan corresponding to the defect feature vector to obtain a maintenance strategy;

[0047] The information sending module is used to send the defect information corresponding to the defect feature vector and the maintenance strategy to the communication terminal of the maintenance personnel.

[0048] Preferably, the above-mentioned equipment defect autonomous decision-making device also includes: a scoring optimization module for obtaining maintenance scoring signals of industrial equipment; using knowledge distillation technology, using the maintenance scoring signal to adjust the node weights of the maintenance knowledge base corresponding to the dynamic knowledge graph, and establishing a scoring-weight mapping relationship; using the scoring-weight mapping relationship to optimize the maintenance knowledge base.

[0049] Preferably, in the above-mentioned equipment defect autonomous decision-making device, the information acquisition module is specifically used to use a multispectral fusion terminal to collect multispectral status images of industrial equipment; use a 3D contour calibration algorithm to compare in real time the spatial coordinates of the industrial equipment in the multispectral status image and the corresponding spatial features in the preset CAD model; combine the gyroscope attitude compensation mechanism to perform spatial coordinate calibration on the spatial features of the industrial equipment; filter and denoise the multispectral status image after spatial coordinate calibration to obtain equipment operation status information of the industrial equipment.

[0050] Preferably, in the above-mentioned equipment defect autonomous decision-making device, the cross-modal learning module is specifically used to upload the equipment operation status information of the industrial equipment to the cross-modal deep analysis model of the cloud server, wherein the cross-modal deep analysis module has a dual-stream heterogeneous deep network architecture, and the equipment operation status information includes a multi-spectral status image and equipment status data; the visual processing flow of the dual-stream heterogeneous deep network architecture is controlled, and the target detection algorithm is combined with the multi-scale attention mechanism to establish feature associations of the multi-spectral status image in the dimensions of multiple scales, and the visual features of the multi-spectral status image are detected; the text processing flow of the dual-stream heterogeneous deep network architecture is controlled, and a large text processing model is used to construct a domain semantic parsing dictionary, and the domain semantic parsing dictionary is used to parse the text features corresponding to the equipment status image; a cross-modal adversarial distillation pipeline is used to collect visual features and text features; a generator network constructed using the cross-modal adversarial distillation pipeline is used to learn the joint distribution of visual features and text features according to the alignment loss function to obtain a defect feature vector.

[0051] Preferably, the above-mentioned cross-modal learning module is specifically used to control the visual processing flow to adopt an improved ViT-Transformer algorithm, introduce a deformable convolution layer to dynamically perceive the convolution kernel offset of the multispectral state image, and capture the equipment defects in the multispectral state image; introduce a multi-scale attention mechanism in the visual processing flow, establish feature associations of the multispectral state image in dimensions of multiple scales, and identify fuzzy defects of industrial equipment; obtain visual features corresponding to image defects and fuzzy defects.

[0052] Preferably, the above-mentioned strategy learning module is specifically used to call the dynamic knowledge graph corresponding to the maintenance knowledge base, perform multi-dimensional matching of defect feature vectors and similar cases, calculate the cosine similarity of the defect features and similar cases according to the multi-dimensional matching degree; extract a predetermined number of maintenance cases whose cosine similarity is above the similarity threshold; and construct an expert experience decision tree based on fuzzy logic, use the expert experience decision tree to process the uncertain conditions corresponding to the maintenance cases through the fuzzy membership function, and obtain a disposal plan; and obtain a maintenance strategy by integrating the maintenance cases and the disposal plan.

[0053] Preferably, the above-mentioned device also includes: a negative feedback regulation module, which is used to continuously collect the operating parameters of the industrial equipment after the industrial equipment is repaired; input the operating parameters into the health index model to obtain the improvement rate of the industrial equipment after repair; use the improvement rate to evaluate the maintenance and treatment effect of the maintenance personnel; when the improvement rate is less than or equal to the predetermined improvement threshold, trigger the parameter update mechanism of the cross-modal deep analysis model.

[0054] According to the third aspect of the present application, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the autonomous defect decision-making method for industrial inspection provided by any of the above technical solutions.

[0055] According to the fourth aspect of the present application, the present application also provides a computer storage medium on which computer executable instructions are stored. When the computer program is executed by the processor, the autonomous defect decision-making method for industrial inspection provided by any of the above technical solutions is implemented.

[0056] The technical solution of this application has at least the following technical effects:

[0057] The technical solution for autonomous decision-making on defects for industrial inspections provided by this application, by using a multi-spectral fusion terminal, can collect equipment operation status information of industrial equipment of multiple modes through multiple spectral image signals, and then use the cross-modal deep analysis model built into the cloud server to establish the features and channel fusion of the equipment operation status information of the above multiple modes, specifically using a dual-stream heterogeneous deep network foot bone to perform cross-modal learning on the above equipment operation status information, learn the joint distribution of the visual-text features of the industrial equipment, and then obtain the defect feature vector of the industrial equipment, thereby realizing cross-modal and cross-dimensional information utilization; then, according to the hybrid dual-engine reasoning mode, the defect feature vector is matched with similar cases in the maintenance knowledge base in multiple dimensions, and the disposal plan is screened. By combining the matched cases and disposal plans, the maintenance strategy of the defect feature of the industrial equipment can be automatically obtained, and the decision-making process does not need to rely on technical personnel. Finally, the defect information and maintenance strategy corresponding to the defect feature vector are sent to the communication terminal of the maintenance personnel, so that the maintenance of industrial equipment defects can be realized. The above solution can solve the problem of multi-modal data fragmentation in the prior art, the failure to effectively establish feature-level fusion, and the low utilization rate of cross-modal information. This will enable closed-loop automation of the entire process of equipment defect detection, autonomous decision-making on maintenance strategies, and equipment maintenance disposal. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0059] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0060] Figure 2 A flowchart of the first method for autonomous defect decision-making for industrial inspections provided in an embodiment of the present application;

[0061] Figure 3for Figure 1 A flow chart of a method for obtaining device operating status information provided by the illustrated embodiment;

[0062] Figure 4 for Figure 1 A schematic flow chart of a cross-modal learning method for device operating status information provided by the illustrated embodiment;

[0063] Figure 5 for Figure 4 A schematic flow chart of a method for detecting visual features provided by the illustrated embodiment;

[0064] Figure 6 for Figure 1 A schematic flow chart of a maintenance strategy matching method provided by the illustrated embodiment;

[0065] Figure 7 for Figure 1 A flowchart of a method for issuing defect information and repair strategies provided by the illustrated embodiment;

[0066] Figure 8 A flowchart of a method for regulating a maintenance knowledge base using negative feedback of a scoring signal provided in an embodiment of the present application;

[0067] Figure 9 A flowchart of a second method for autonomous defect decision-making for industrial inspections provided in an embodiment of the present application;

[0068] Figure 10 A schematic diagram of the structure of an autonomous defect decision-making device for industrial inspection provided in an embodiment of the present application;

[0069] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0070] Figure 12 A schematic diagram of the overall architecture of an autonomous defect decision-making system for industrial inspections provided in an embodiment of the present application;

[0071] Figure 13 A schematic diagram of the architecture of a cross-modal adversarial distillation pipeline provided in an embodiment of the present application;

[0072] Figure 14 A schematic diagram of the architecture of maintaining a scoring signal combined with knowledge distillation technology provided in an embodiment of the present application. DETAILED DESCRIPTION

[0073] In order to more clearly illustrate the overall concept of the present application, a detailed description is given below in an illustrative manner in conjunction with the accompanying drawings.

[0074] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application may also be implemented in other ways than those described herein, and therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below. It should be noted that the embodiments of the present application and the features of each embodiment may be combined with each other unless there is a conflict.

[0075] In this application, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more embodiments or examples.

[0076] The existing technology has the following defects:

[0077] 1. Severe dependence on personnel: Experienced maintenance personnel are required to participate in the entire decision-making process, resulting in high labor costs, subjective judgment bias, and delayed response time;

[0078] 2. Broken decision-making chain: There is a lack of intelligent connection between defect identification results and the maintenance knowledge base, and the reuse rate of historical cases is low, forming an information island of "detection-decision-making".

[0079] The above defects are caused by the triple division of traditional technical architecture:

[0080] (1) Feature-level channel fusion is not established for multimodal data (images / sensor data / maintenance records), resulting in low utilization of cross-dimensional information;

[0081] (2) the decision-making process is overly dependent on experienced technical personnel;

[0082] (3) Knowledge base updates rely on manual input and lack a self-evolution mechanism based on on-site feedback.

[0083] To solve the above technical problems, see Figure 1, the following embodiment of the present application provides an "end-cloud-library" defect autonomous decision-making method for industrial inspections, including a handheld multispectral fusion terminal 1, a cloud server 2 and an operator's communication terminal 4; wherein, the cloud server 2 includes a maintenance knowledge base 3; it can realize the integration of terminal collection, cloud retrieval, knowledge base matching and knowledge recommendation, thereby effectively solving the problems existing in the above-mentioned method. Specifically, by fusing the visible light, infrared and ultraviolet multispectral imaging data of the multispectral fusion terminal 1 with the operating status parameters of the industrial equipment 5, combined with the large model reasoning and dynamic knowledge graph update mechanism of the cloud server 2, a maintenance strategy is formed and sent to the maintenance personnel's communication terminal 4. The above-mentioned method can realize the full process closed-loop autonomy of equipment defect detection, maintenance strategy matching and disposal effect evaluation. The present application can focus on solving the defects of multimodal data fragmentation processing and response delay and lack of standardization of disposal solutions caused by reliance on manual experience in traditional inspections, and provide a real-time and accurate decision support system for predictive maintenance of industrial equipment.

[0084] To achieve the above purpose, see Figure 2 , Figure 2 This is a flow chart of a defect autonomous decision-making method for industrial inspection provided by an embodiment of the present application. Figure 2 As shown in FIG, the defect autonomous decision-making method for industrial inspection includes:

[0085] S110: Use a multispectral fusion terminal to collect equipment operating status information of industrial equipment and upload it to the cloud server.

[0086] Combine Figure 1 As shown in the application scenario diagram, inspectors use a handheld multispectral fusion terminal (capable of integrating visible light, infrared, and ultraviolet imaging modules) to capture multispectral status images of industrial equipment. These images reflect the operational status of the equipment. This terminal incorporates a basic anti-shake algorithm based on gyroscope attitude compensation and a device contour matching function. This device contour matching function utilizes the scale-invariant feature transform (SIFT) feature point detection algorithm to perform rotation correction and median filtering on the captured images. The preprocessed multispectral status images are then linked to the device ID and acquisition timestamp before being uploaded to the cloud. The multispectral fusion terminal captures these images, providing multimodal operational status data for industrial equipment.

[0087] Specifically, as a preferred embodiment, Figure 3 As shown, step S110: using a multispectral fusion terminal to collect equipment operating status information of industrial equipment and uploading it to a cloud server includes:

[0088] S111: Use a multispectral fusion terminal to collect multispectral status images of industrial equipment;

[0089] S112: Using a 3D profile calibration algorithm, the spatial coordinates of the industrial equipment in the multispectral status image and the corresponding spatial features in the preset CAD model are compared in real time;

[0090] S113: Combine the gyroscope attitude compensation mechanism to calibrate the spatial coordinates of the industrial equipment;

[0091] S114: performing filtering and denoising processing on the multispectral status image after spatial coordinate calibration to obtain equipment operation status information of the industrial equipment.

[0092] The technical solution provided in the embodiments of the present application can realize multi-modal image data acquisition. In a specific embodiment, the inspection personnel use a dedicated three-spectrum fusion handheld terminal (integrated with visible light, infrared, and ultraviolet three-channel imaging modules) to collect equipment status data. When a specific defect mode is selected (such as local overheating), the multi-spectral fusion terminal automatically switches to the optimal imaging combination (such as infrared main channel + ultraviolet auxiliary channel) based on a preset strategy. In order to solve the problem of imaging distortion under complex working conditions, a 3D contour calibration algorithm based on the equipment CAD model is introduced: by comparing the spatial features of the collected image with the preset CAD model in real time, the spatial coordinate calibration is performed in combination with the gyroscope attitude compensation module to reduce the distortion rate of the imaging. The collected data will be processed in the spatial domain to eliminate the shooting angle deviation and denoising, which will effectively improve the availability of the original data in a variety of environments.

[0093] The existing technology has three structural defects: first, the fragmentation of multimodal perception; second, the black box nature of the decision-making process; and third, the black box nature of the decision-making process. In terms of the fragmentation of multimodal perception, the visible light, infrared, and ultraviolet imaging data and sensor parameters collected by traditional technologies are analyzed independently, lacking a cross-modal feature association model. To solve this problem, this application builds a cross-modal deep analysis model in the cloud to achieve cross-modal fusion of multiple imaging data and sensor parameters. Specifically:

[0094] Figure 2 The technical solution provided by the illustrated embodiment, after step S110: using a multi-spectral fusion terminal to collect equipment operating status information of industrial equipment, further includes the following steps:

[0095] S120: Using the cross-modal deep analysis model built into the cloud server, we conduct cross-modal learning of the equipment operating status information based on the dual-stream heterogeneous deep network architecture, learn the joint distribution of the visual and textual features of industrial equipment, and obtain the defect feature vector of the industrial equipment.

[0096] In the embodiment of the present application, the extraction and classification of defect features can be achieved through a cross-modal deep analysis model. Specifically, an improved YOLOv5 detection model is deployed on a cloud server. The improvement of the improved YOLOv5 detection model is that an attention mechanism is introduced in its classification layer. Specifically, multi-scale defect features can be extracted through a feature pyramid network (FPN). The attention mechanism (SE module) is introduced into the detection model. The attention mechanism SE module can play the role of calibration, attention and aggregation, thereby improving the recognition accuracy of minor defects and extracting defect features that are of focus. The output results of the defect features detected include defect location, type and confidence index. Specifically, the spatial attention mechanism SE increases attention in the channel dimension, and after placing the SE module in the spatial pyramid pooling layer SPPF (Spatial Pyramid Pooling Fast), the network is recalibrated.

[0097] Specifically, as a preferred embodiment, Figure 4 As shown, step S120: using the cross-modal deep analysis model built into the cloud server, performing cross-modal learning on the equipment operating status information based on the dual-stream heterogeneous deep network architecture, learning the joint distribution of visual-text features of industrial equipment, and obtaining the defect feature vector of the industrial equipment, includes:

[0098] S121: A cross-modal deep analysis model uploads the operating status information of industrial equipment to a cloud server. The cross-modal deep analysis module incorporates a dual-stream heterogeneous deep network architecture. The equipment operating status information includes multispectral status images and equipment status data. The large model deployed on the cloud, the cross-modal deep analysis model, has a dual-stream heterogeneous deep network architecture consisting of a visual processing stream (i.e., an image encoder) and a text processing stream (i.e., a text encoder). These two streams learn visual and text features related to the equipment operating status, and ultimately combine these features to generate a defect vector.

[0099] S122: Control the visual processing stream of the dual-stream heterogeneous deep network architecture, use the target detection algorithm combined with the multi-scale attention mechanism, establish feature associations for the multispectral state image at multiple scales, and detect and obtain visual features of the multispectral state image. The target detection algorithm provided in the embodiment of this application can use the improved YOLOv5 detection model. The improvement of the detection model lies in the introduction of an attention mechanism (SE module) in its classification layer and the extraction of multi-scale defect features through a feature pyramid network (FPN).

[0100] S123: Controls the text processing flow of the dual-stream heterogeneous deep network architecture, uses the large text processing model to build a domain semantic parsing dictionary, and uses this dictionary to parse the text features of the equipment status image. This large text processing model can be used as a knowledge-enhanced GPT-3 model. By incorporating structured knowledge from equipment maintenance records (such as the association rule "bearing noise - insufficient lubrication - grease change"), it builds a domain-specific semantic parsing dictionary, which can then be used to parse the text features corresponding to the equipment status image.

[0101] S124: Use a cross-modal adversarial distillation pipeline to adversarially collect visual and text features. For example, the cross-modal adversarial distillation pipeline has the same dimensionality as the visual and text features. This allows the adversarial collection of the visual and text features by competing for the dimensionality of the distillation pipeline. This enables cross-modal collection and fusion of bimodal, or even multimodal, features.

[0102] S125: A generator network constructed using a cross-modal adversarial distillation pipeline learns the joint distribution of visual and textual features according to the alignment loss function to obtain a defect feature vector. Figure 13 The cross-modal adversarial distillation pipeline consists of a generator network, a discriminator network, and a feature alignment loss function. The generator network mainly includes two functions: learning the joint distribution of cross-modal features and generating fused features.

[0103] The generator network is the core execution engine of the cross-modal adversarial distillation pipeline. Through adversarial learning of visual and textual features, it can cross-modally collect and learn the joint distribution of features from both modalities, thereby resolving the existing problem of multimodal data fragmentation and the inability to perform feature-level fusion. Specifically, both visual and textual features are 128-dimensional, allowing the generator network to compress the 256-dimensional concatenated vector (128+128) to 128 dimensions.

[0104] As a specific embodiment, the present application uploads the device image data to a cloud-based large model, which can detect the defect information corresponding to the device image data. The cloud-based large model is designed with a dual-stream heterogeneous deep network architecture, which includes a visual processing stream and a text processing stream. Figure 12 As can be seen from the architecture shown, the visual processing flow (i.e., image encoder) uses the improved ViT-Transformer model (input size 384×384) as the model of the target detection algorithm. The improvement of this model lies in the introduction of a deformable convolution layer (Deformable DCNN) into the standard module of the original ViT-Transformer model. DCNN can dynamically perceive the convolution kernel offset (offset range ±5 pixels) and effectively capture tiny defects in multispectral images (such as size <0.5mm).2 Cracks); a multi-scale attention mechanism is also introduced to establish feature associations at three scales, 16×16, 32×32, and 64×64, to improve the recognition of ambiguous defects. This visual processing stream generates a visual feature vector. The text processing stream (i.e., the text encoder), based on the knowledge-enhanced GPT-3 model, incorporates structured knowledge from equipment maintenance records (e.g., association rules such as "bearing noise - insufficient lubrication - grease change") to construct a domain-specific semantic parsing dictionary and generate a text feature vector. Bimodal features (visual and textual) are independently encoded to generate 128-dimensional semantic vectors. A cross-modal adversarial distillation pipeline is designed, within which a generator network (containing four fully connected layers) is constructed to learn the joint distribution of visual and textual features, ultimately generating a defect vector. This defect feature vector is then matched against a knowledge base. The adversarial distillation pipeline is 128-dimensional. Therefore, bimodal features with dimensions greater than or equal to 128 compete with each other when input into the pipeline, requiring adversarial competition for the pipeline's dimensions to achieve a joint distribution of visual and textual features. In this way, the generator network can learn the joint distribution of the visual-text features according to the loss function and obtain the defect feature vector. The alignment loss function is defined as:

[0105]

[0106] in, The image encoder is the Image samples The extracted feature vectors, The text encoder (Text Encoder) Text samples Extracted feature vectors.

[0107] As a preferred embodiment, for the above step S122: controlling the visual processing flow of the dual-stream heterogeneous deep network architecture, using the target detection algorithm combined with the multi-scale attention mechanism, establishing feature associations of the multispectral state image at multiple scales, and detecting and obtaining the visual features of the multispectral state image. Figure 5 :

[0108] S1221: The improved ViT-Transformer algorithm is used to control the visual processing flow. A deformable convolution layer is introduced to dynamically perceive the convolution kernel offset of the multispectral state image and capture device defects in the multispectral state image.

[0109] S1222: Introducing a multi-scale attention mechanism into the visual processing flow to establish feature associations of multispectral state images at multiple scales to identify fuzzy defects in industrial equipment;

[0110] S1223: Obtain visual features corresponding to image defects and blur defects.

[0111] The technical solution provided in the embodiment of the present application uses an improved ViT-Transformer model (input size 384×384) as the target detection algorithm. A deformable convolution layer (DCNN) is introduced into the standard module of the model. The DCNN is used to dynamically perceive the convolution kernel offset (offset range ±5 pixels), which can effectively capture tiny defects in multispectral images (such as those with a size <0.5mm). 2 At the same time, a multi-scale attention mechanism is introduced to establish feature associations at three scales: 16×16, 32×32, and 64×64, improving the ability to recognize fuzzy defects.

[0112] Among the three major flaws in existing technologies, the decision-making process is black-boxed. This means that while deep learning models can output defect classification results, they lack an interpretable maintenance strategy derivation process and a knowledge matching process. The following embodiments of this application compare defect feature vectors with similar cases in a maintenance knowledge base, matching the closest maintenance case and generating a solution for the equipment defect.

[0113] Specifically, Figure 2 The technical solution provided by the illustrated embodiment, after using the cross-modal deep analysis model built into the cloud server to perform cross-modal learning on the equipment operating status information based on the dual-stream heterogeneous deep network architecture, learns the joint distribution of visual and textual features of industrial equipment, and obtains the defect feature vector of the industrial equipment, further includes:

[0114] S130: Based on the hybrid dual-engine reasoning mode, multi-dimensional matching is performed between the defect feature vector and similar cases in the maintenance knowledge base, and the treatment plan corresponding to the defect feature vector is screened to obtain a maintenance strategy.

[0115] The embodiment of the present application searches the local maintenance knowledge base, which stores standard cases such as maintenance work orders and maintenance records generated by daily maintenance, and then calculates the cosine similarity between the 128-dimensional current defect feature vector and the historical cases in the maintenance knowledge base. Specifically, if the cosine similarity is greater than or equal to a certain threshold (threshold>0.75), the top 5 candidate cases are returned. If the threshold does not meet the requirements, no case is returned. The decision tree-based rule engine is used to filter the disposal solutions that meet the current working conditions (such as "prioritize shutdown inspection when the temperature is>80°C"). It should be noted that the defect feature vector is a feature that is fused later, and 128 dimensions represent the dimension of the defect feature vector.

[0116] As a preferred embodiment, Figure 6As shown, in the above-mentioned equipment defect autonomous decision-making method, step S130: based on the hybrid dual-engine reasoning mode, multi-dimensional matching is performed between the defect feature vector and similar cases in the maintenance knowledge base, and the treatment plan corresponding to the defect feature vector is screened to obtain the maintenance strategy, which specifically includes:

[0117] S131: Call the dynamic knowledge graph corresponding to the maintenance knowledge base, perform multi-dimensional matching on the defect feature vector and similar cases, and calculate the cosine similarity between the defect feature and the similar case according to the multi-dimensional matching degree. The calculation formula of the cosine similarity is as follows:

[0118] Assume that the n-dimensional defect feature vector A=(a1,a2,…a n ) and B=(b1,b2,…,b n ), in the embodiment of the present application, n=128, then the calculation formula of cosine similarity is as follows

[0119] Consine Similarity(A,B) =

[0120] Among them, the molecular part It represents the inner product of vectors A and B, and the denominator represents the product of the modules of the two vectors.

[0121] S132: Extract a predetermined number of maintenance cases whose cosine similarity is above a similarity threshold.

[0122] as well as,

[0123] S133: Build an expert experience decision tree based on fuzzy logic. This tree uses fuzzy membership functions to process the uncertainties associated with maintenance cases and derive solutions. These solutions are derived from rules in the knowledge base, real-time operating conditions, and other factors. Defect feature vectors (128 dimensions) are retrieved using cosine similarity to retrieve historical cases. The fuzzy logic-based expert experience decision tree then incorporates real-time operating condition constraints to achieve dynamic corrections.

[0124] S134: Integrate maintenance cases and disposal plans to obtain maintenance strategies.

[0125] The technical solution provided in this embodiment of the application retrieves similar cases from a maintenance knowledge base, matches maintenance strategies, and generates a resolution plan. Specifically, this involves invoking the dynamic knowledge graph corresponding to the maintenance knowledge base, performing multi-dimensional matching on the defect feature vectors, calculating the cosine similarity between the current defect feature vector and historical cases, and screening candidate cases. The dimensions required for matching include equipment, environment, and resources.

[0126] The reasoning process for the maintenance strategy in the embodiment of this application consists of two parts. Specifically, a hybrid dual-engine reasoning model is used, with both a case-based reasoning module and a rule-based reasoning module. The case-based reasoning (CBR) module calculates the cosine similarity between the current defect feature vector and historical cases (threshold > 0.85), screens the top 10 cases as candidate cases, and extracts high-frequency action items (e.g., lubricant replacement accounts for 82% of cases involving "bearing overheating"). The rule-based reasoning (RBR) module constructs an expert experience decision tree based on fuzzy logic and uses fuzzy membership functions to handle uncertain conditions (e.g., "high temperature" is defined as a Gaussian distribution in the range of 70-90°C), achieving human-like decision-making. Knowledge distillation technology is used to convert the maintenance score signals uploaded by operators into graph node weight adjustments, establish a score-weight mapping relationship, and achieve progressive optimization of the knowledge base.

[0127] Figure 2 The technical solution provided by the illustrated embodiment further includes the following steps after obtaining the maintenance strategy:

[0128] S140: Send the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel.

[0129] After receiving a maintenance decision and executing it, maintenance personnel continuously collect equipment operating parameters (such as temperature and vibration) to calculate the equipment health index (HI). When the HI improvement rate is less than or equal to a certain level, an analysis of the maintenance cases in the maintenance knowledge base is triggered. Maintenance decisions are manually revised, and the parameters of the cross-modal deep analysis model (such as the convolution kernel parameters) are regularly fine-tuned.

[0130] Specifically, as a preferred embodiment, Figure 7 As shown, after step S140: sending the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel, the method further includes:

[0131] S141: After the industrial equipment is repaired, the operating parameters of the industrial equipment are continuously collected;

[0132] S142: Inputting the operating parameters into the health index model to obtain the improvement rate of the industrial equipment after maintenance;

[0133] S143: Use improvement rate to evaluate the maintenance personnel's maintenance treatment effect;

[0134] S144: When the improvement rate is less than or equal to a predetermined improvement threshold, triggering a parameter update mechanism of the cross-modal deep analysis model;

[0135] S145: Acquire a maintenance score signal uploaded by the communication terminal; the maintenance score signal here is generated in combination with the improvement rate.

[0136] S146: Using the maintenance scoring signal, the node weights of the maintenance knowledge base corresponding to the dynamic knowledge graph are iteratively updated through knowledge distillation technology.

[0137] Knowledge distillation is a model compression and acceleration technology that aims to transfer the knowledge learned by a large model (usually called a teacher model) to a small model (usually called a student model), allowing the small model to achieve performance as close to that of the large model as possible while reducing computing resource consumption and inference time.

[0138] like Figure 14 As shown, in an embodiment of the present application, the maintenance personnel's subjective rating of the maintenance plan (maintenance rating signal) is characterized as a teacher signal, which represents the actual evaluation result of the plan; the decision recommendation currently output by the dynamic knowledge graph is transformed into a student signal, which reflects the reasoning result of the model. The embodiment of the present application constructs a mapping through knowledge distillation technology. After receiving the maintenance rating signal, the system locates the graph node of the dynamic knowledge graph directly associated with the plan, calculates the weight change value through the mapping function, and the high or low rating affects the direction of adjustment. By comparing the difference between the teacher signal and the student signal, the weight distribution of the associated nodes in the knowledge graph is dynamically adjusted. This process enables high-scoring plans to be prioritized in subsequent case matching and low-scoring plans to be suppressed, thereby achieving progressive optimization of the cloud-based AI model (the above-mentioned cross-modal deep analysis model).

[0139] In the solution provided by the embodiments of this application, defect information and repair strategies are communicated to maintenance personnel via cloud messaging. After performing repair operations, maintenance personnel continuously collect equipment operating parameters (vibration values, temperature curves, etc.) and evaluate the effectiveness of the repair using a health index model. Taking the defect information detected by industrial equipment: temperature, vibration values, and current output power as an example, the health index model corresponding to these three indicators is as follows:

[0140]

[0141] If it is used to detect whether the surface temperature of the device is abnormal, the above formula Indicates the temperature change of key parts of the equipment, that is, the absolute value of the temperature difference at the same monitoring point before and after maintenance; Indicates the standard temperature value of the equipment under normal working conditions; The effective value of the current vibration amplitude of the equipment, which represents the stability of the mechanical structure. The vibration signal can be collected by an acceleration sensor (such as a piezoelectric sensor); Indicates the baseline vibration amplitude of the equipment in a healthy state; Indicates the actual measured value of the device's current output power; is the rated power of the equipment; the above weights (0.5, 0.3, and 0.2) respectively reflect the contribution of temperature, vibration, and power to the health status of the equipment. Different equipment can be replaced with other evaluation indicators.

[0142] In the embodiment of the present application, when the HI improvement rate is <30%, a necessary model update is triggered, and the daily maintenance score (1-5 stars) updates the graph node weights through knowledge distillation, forming a closed-loop feedback system to iteratively update the weights.

[0143] The method for autonomous decision-making on defects for industrial inspections in an embodiment of the present application, by using a multispectral fusion terminal, can collect equipment operating status information of industrial equipment of multiple modalities through multiple spectral image signals, and then use the cross-modal deep analysis model built into the cloud server to establish the features and channel fusion of the equipment operating status information of the above multiple modalities, specifically using a dual-stream heterogeneous deep network foot bone to perform cross-modal learning on the above equipment operating status information, learn the joint distribution of the visual-text features of the industrial equipment, and then obtain the defect feature vector of the industrial equipment, thereby realizing cross-modal and cross-dimensional information utilization; then, according to the hybrid dual-engine reasoning mode, the defect feature vector is matched with similar cases in the maintenance knowledge base in multiple dimensions, and the disposal plan is screened. By combining the matched cases and disposal plans, the maintenance strategy of the defect feature of the industrial equipment can be automatically obtained, and the decision-making process does not need to rely on technical personnel. Finally, the defect information and maintenance strategy corresponding to the defect feature vector are sent to the communication terminal of the maintenance personnel, so that the maintenance of industrial equipment defects can be realized. The above scheme can solve the problem of multimodal data fragmentation in the prior art, the failure to effectively establish feature-level fusion, and the low utilization rate of cross-modal information. This will enable full automation of the entire process of equipment defect detection, autonomous decision-making on maintenance strategies, and equipment maintenance and disposal.

[0144] In addition, it should be noted that the existing maintenance knowledge base has a sluggish knowledge evolution, and the maintenance case library update relies on manual input, which makes it impossible to achieve a closed-loop iteration from field feedback to knowledge refinement. To solve the above problems, the embodiment of the present application triggers the necessary model update when the improvement rate of industrial equipment is less than a certain level. The daily maintenance score updates the graph node weights through knowledge distillation, forming a closed-loop feedback, and the system iteratively updates the weights.

[0145] Specifically, as a preferred embodiment, Figure 8 As shown, after the step S140 of sending the defect information and maintenance strategy corresponding to the defect feature vector to the communication terminal of the maintenance personnel, the above-mentioned equipment defect autonomous decision-making method further includes:

[0146] S150: Obtaining a maintenance score signal of the industrial equipment;

[0147] S160: Using knowledge distillation technology, the maintenance score signal is used to adjust the node weights of the maintenance knowledge base corresponding to the dynamic knowledge graph, and establish a score-weight mapping relationship;

[0148] S170: Optimize the maintenance knowledge base using the score-weight mapping relationship.

[0149] Knowledge distillation is a technique that transfers knowledge from a large model (teacher model) to a small model (student model). The goal is to train the student model to achieve or approach the performance of the teacher model at a lower computational cost. Figure 14 As shown, this application first extracts features from the maintenance scoring signal; then constructs a teacher model, whose input is node embedding and multi-dimensional scoring features, and outputs the probability distribution of node weights; and establishes a multi-source scoring fusion loss function, a score-weight consistency loss function and a distillation optimization objective. Finally, based on the teacher model, a score-weight mapping relationship is generated to update the maintenance knowledge base.

[0150] In summary, the technical solutions provided by the embodiments of this application require feedback and model optimization after repairing defects in industrial equipment to achieve a closed-loop feedback loop. For example, after maintenance personnel receive a repair decision and execute the repair, they continuously collect equipment operating parameters (such as temperature and vibration) to calculate the equipment health index (HI). When the HI improvement rate is less than 20%, an analysis of the repair cases in the maintenance knowledge base is triggered, the repair decision is manually revised, and the model parameters of the cross-modal deep analysis model, such as the convolution kernel parameters, are regularly fine-tuned.

[0151] Also, see Figure 9 , Figure 9 A flowchart of a method for autonomous defect decision-making for industrial inspection provided in an embodiment of the present application is shown as follows: Figure 9 As shown, the autonomous decision-making methods for the device defects include:

[0152] S201: Multimodal image acquisition.

[0153] S202: Cross-modal deep analysis model.

[0154] S203: Dynamic knowledge decision system.

[0155] S204: Determine whether there is a device defect; if so, execute step S205; if not, return to execute step S201.

[0156] S205: Matching maintenance knowledge base.

[0157] S206: Issue maintenance instructions.

[0158] In summary, the defects of the existing technology are due to three major architectural contradictions: first, the fragmentation of multimodal perception, where visible light, infrared, and ultraviolet imaging data and sensor parameters are analyzed independently, and there is a lack of cross-modal feature association models; second, the sluggish evolution of knowledge, where the maintenance case library is updated relying on manual entry, and closed-loop iteration from on-site feedback to knowledge extraction cannot be achieved; third, the decision-making process is black-boxed, and although the deep learning model can output defect classification results, it lacks an explainable maintenance strategy derivation link and a knowledge matching process. The autonomous defect decision-making solution for industrial inspections provided in the above embodiments of the present application and the proposed "end-cloud-library" method can realize the integration of terminal collection, cloud retrieval, knowledge base matching, and knowledge recommendation, effectively solving the above three architectural contradictions existing in the existing methods.

[0159] In addition, participate Figure 10 , Figure 10 An autonomous defect decision-making device for industrial inspection provided in an embodiment of the present application includes:

[0160] The information collection module 110 is used to collect the equipment operation status information of the industrial equipment using the multi-spectral fusion terminal and upload it to the cloud server;

[0161] a cross-modal learning module 120 for performing cross-modal learning on the equipment operating status information based on a dual-stream heterogeneous deep network architecture using a cross-modal deep analysis model built into the cloud server, learning the joint distribution of visual and textual features of the industrial equipment, and obtaining a defect feature vector of the industrial equipment;

[0162] A strategy learning module 130 is configured to perform multi-dimensional matching between the defect feature vector and similar cases in a maintenance knowledge base and select treatment plans corresponding to the defect feature vector to obtain a maintenance strategy based on a hybrid dual-engine reasoning mode;

[0163] The information sending module 140 is configured to send the defect information corresponding to the defect feature vector and the maintenance strategy to the communication terminal of the maintenance personnel.

[0164] The embodiment of the present application is implemented through the following technical solutions: inspection personnel collect equipment image data through handheld terminals; equipment image data is uploaded to a large cloud model, and defect information is detected on the cloud; similar cases are retrieved from the maintenance knowledge base, maintenance strategies are matched, and disposal plans are generated; defect information and maintenance strategies are sent to maintenance personnel for repair.

[0165] In summary, the technical solutions provided by the above embodiments of this application address the three major technical difficulties that have long existed in the field of industrial inspections, namely, multimodal data fragmentation analysis, static solidification of maintenance knowledge, and excessive reliance on manual labor in the decision-making process. A method for autonomous decision-making on equipment defects based on the "end-cloud-library" collaborative architecture is proposed. By constructing a multimodal feature fusion engine and a dynamic knowledge evolution system, the performance bottleneck of traditional technologies is broken through, and the full process from defect perception to maintenance decision-making is realized to be autonomous, solving the excessive reliance of traditional methods on manual experience. After actual verification, the time required to generate maintenance plans has been compressed from hours to minutes.

[0166] In addition, it should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0167] In addition, see Figure 11 The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the autonomous defect decision-making method for industrial inspection of any of the above embodiments.

[0168] Reference below Figure 11 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application can include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0169] like Figure 11As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the electronic device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003 including, for example, a magnetic tape or hard disk; and communication device 1009. The communication device 1009 enables the autonomous defect decision-making device for industrial inspections to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a model building device with various systems, it should be understood that implementation or presence of all the illustrated systems is not required. More or fewer systems may be implemented or present instead.

[0170] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0171] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0172] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0173] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the defect autonomous decision-making method for industrial inspection in the above-mentioned embodiment.

[0174] The computer-readable storage medium provided in this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0175] The computer-readable storage medium carries one or more programs that, when executed by the model building device, can be written in one or more programming languages ​​or a combination thereof to implement computer program code for performing the operations of the present application. The programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0177] The modules described in the embodiments of the present application can be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0178] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for autonomous defect decision-making for industrial inspection.

[0179] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0180] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present application.

Claims

1. A defect autonomous decision-making method for industrial inspection, characterized by: include: Use multi-spectral fusion terminals to collect equipment operating status information of industrial equipment and upload it to the cloud server; Using a cross-modal deep analysis model built into a cloud server, cross-modal learning is performed on the equipment operating status information based on a dual-stream heterogeneous deep network architecture to learn the joint distribution of visual and textual features of the industrial equipment and obtain a defect feature vector of the industrial equipment; Based on the hybrid dual-engine reasoning model, the defect feature vector is matched with similar cases in the maintenance knowledge base in multiple dimensions, and the corresponding treatment plans are screened to obtain a maintenance strategy. The hybrid dual-engine reasoning model includes a case-based reasoning module and a rule-based reasoning module. The case-based reasoning module calculates the cosine similarity between the current defect feature vector and historical cases to screen cases, while the rule-based reasoning module constructs an expert experience decision tree based on fuzzy logic and uses a fuzzy membership function to handle uncertain conditions. Sending the defect information corresponding to the defect feature vector and the maintenance strategy to the communication terminal of the maintenance personnel; The step of using the cross-modal deep analysis model built into the cloud server to perform cross-modal learning on the equipment operating status information based on a dual-stream heterogeneous deep network architecture, learning the joint distribution of the visual-text features of the industrial equipment, and obtaining the defect feature vector of the industrial equipment includes: Uploading the equipment operation status information of the industrial equipment to the cross-modal deep analysis model of the cloud server, wherein the cross-modal deep analysis model includes the dual-stream heterogeneous deep network architecture, and the equipment operation status information includes a multispectral status image and equipment status data; wherein the dual-stream heterogeneous deep network architecture includes an improved YOLOv5 detection model, and the improved point of the improved YOLOv5 detection model is to introduce an attention mechanism into its classification layer; Controlling the visual processing stream of the dual-stream heterogeneous deep network architecture, using a target detection algorithm combined with a multi-scale attention mechanism to establish feature associations of the multispectral state image at multiple scales, and detecting and obtaining visual features of the multispectral state image; Controlling the text processing flow of the dual-stream heterogeneous deep network architecture, using the text processing large model to build a domain semantic parsing dictionary, and using the domain semantic parsing dictionary to parse text features corresponding to the device status data; The visual features and text features are collected using a cross-modal adversarial distillation pipeline; wherein the cross-modal adversarial distillation pipeline includes a generator network, a discriminator network and a feature alignment loss function, and the generator network includes learning a joint distribution of cross-modal features and generating fused features; A generator network constructed using the cross-modal adversarial distillation pipeline is used to learn the joint distribution of the visual features and text features according to a feature alignment loss function to obtain the defect feature vector.

2. The method according to claim 1, wherein After the step of sending the defect information corresponding to the defect feature vector and the maintenance strategy to the communication terminal of the maintenance personnel, the method further includes: obtaining a maintenance score signal of the industrial equipment; Adopting knowledge distillation technology, using the maintenance score signal to adjust the node weights of the maintenance knowledge base corresponding to the dynamic knowledge graph, and establishing a score-weight mapping relationship; The maintenance knowledge base is optimized using the score-weight mapping relationship.

3. The method according to claim 1, wherein The step of using a multispectral fusion terminal to collect equipment operating status information of industrial equipment and uploading it to a cloud server includes: Using the multispectral fusion terminal, collecting the multispectral status image of the industrial equipment; Using a 3D profile calibration algorithm, the spatial coordinates of the industrial equipment in the multispectral status image are compared in real time with the corresponding spatial features in the preset CAD model; Performing spatial coordinate calibration on the spatial characteristics of the industrial equipment in combination with a gyroscope attitude compensation mechanism; The multispectral state image after spatial coordinate calibration is subjected to filtering and denoising processing to obtain equipment operation state information of the industrial equipment.

4. The method according to claim 1, wherein The step of controlling the visual processing flow of the dual-stream heterogeneous deep network architecture, using a target detection algorithm combined with a multi-scale attention mechanism, establishing feature associations of the multispectral state image in dimensions of multiple scales, and detecting and obtaining visual features of the multispectral state image includes: Controlling the visual processing flow using an improved ViT-Transformer algorithm, introducing a deformable convolution layer to dynamically perceive the convolution kernel offset of the multispectral state image, and capturing device defects in the multispectral state image; Introducing a multi-scale attention mechanism into the visual processing flow to establish feature associations of the multi-spectral state image at multiple scales to identify fuzzy defects of the industrial equipment; Obtain visual features corresponding to the image defect and the blur defect.

5. The method according to claim 1, wherein The step of performing multi-dimensional matching between the defect feature vector and similar cases in the maintenance knowledge base and screening the treatment plan corresponding to the defect feature vector according to the hybrid dual-engine reasoning mode to obtain the maintenance strategy includes: Calling the dynamic knowledge graph corresponding to the maintenance knowledge base, performing multi-dimensional matching on the defect feature vector and similar cases, and calculating the cosine similarity between the defect feature and the similar cases according to the multi-dimensional matching degree; Extracting maintenance cases whose cosine similarity is greater than a similarity threshold; as well as, Constructing an expert experience decision tree based on fuzzy logic, and using the expert experience decision tree to process the uncertain conditions corresponding to the maintenance case through a fuzzy membership function to obtain a disposal plan; The maintenance cases and treatment plans are combined to obtain the maintenance strategy.

6. The method according to claim 1, wherein After the step of sending the defect information corresponding to the defect feature vector and the maintenance strategy to the communication terminal of the maintenance personnel, the method further includes: After the industrial equipment is repaired, continuously collecting operating parameters of the industrial equipment; Inputting the operating parameters into a health index model to obtain an improvement rate of the industrial equipment after maintenance; Using the improvement rate to evaluate the maintenance effect of the maintenance personnel; When the improvement rate is less than or equal to a predetermined improvement threshold, triggering a parameter update mechanism of the cross-modal deep analysis model; Obtaining a maintenance scoring signal uploaded by the communication terminal; The maintenance scoring signal is used to iteratively update the node weights of the dynamic knowledge graph corresponding to the maintenance knowledge base through knowledge distillation technology.

7. A defect autonomous decision-making device for industrial inspection, characterized by: include: The information collection module is used to collect the equipment operation status information of industrial equipment using a multi-spectral fusion terminal and upload it to the cloud server; a cross-modal learning module, configured to use a cross-modal deep analysis model built into a cloud server to perform cross-modal learning on the equipment operating status information based on a dual-stream heterogeneous deep network architecture, learn the joint distribution of visual and textual features of the industrial equipment, and obtain a defect feature vector for the industrial equipment; A strategy learning module is used to perform multi-dimensional matching between the defect feature vector and similar cases in the maintenance knowledge base based on a hybrid dual-engine reasoning model, and to screen the corresponding treatment plans for the defect feature vector to obtain a maintenance strategy. The hybrid dual-engine reasoning model includes a case-based reasoning module and a rule-based reasoning module. The case-based reasoning module calculates the cosine similarity between the current defect feature vector and historical cases to screen cases, and the rule-based reasoning module constructs an expert experience decision tree based on fuzzy logic and uses a fuzzy membership function to handle uncertain conditions. An information sending module, configured to send the defect information corresponding to the defect feature vector and the maintenance strategy to a communication terminal of a maintenance personnel; Among them, the cross-modal learning module is specifically used to upload the equipment operation status information of industrial equipment to the cross-modal deep analysis model of the cloud server, wherein the cross-modal deep analysis module has a dual-stream heterogeneous deep network architecture, and the equipment operation status information includes multi-spectral status images and equipment status data; the dual-stream heterogeneous deep network architecture includes an improved YOLOv5 detection model, and the improvement of the improved YOLOv5 detection model is that the attention mechanism is introduced into its classification layer; the visual processing flow of the dual-stream heterogeneous deep network architecture is controlled, and the target detection algorithm is combined with the multi-scale attention mechanism to establish feature associations of multi-spectral status images in multiple scale dimensions, and detect The visual features of the multispectral status image are measured; the text processing flow of the dual-stream heterogeneous deep network architecture is controlled, the domain semantic parsing dictionary is constructed using the text processing large model, and the text features corresponding to the device status image are parsed using the domain semantic parsing dictionary; a cross-modal adversarial distillation pipeline is used to collect visual features and text features; the cross-modal adversarial distillation pipeline includes a generator network, a discriminator network and a feature alignment loss function, and the generator network includes learning the joint distribution of cross-modal features and generating fusion features; the generator network constructed using the cross-modal adversarial distillation pipeline learns the joint distribution of visual features and text features according to the feature alignment loss function to obtain a defect feature vector.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the autonomous defect decision-making method for industrial inspection as described in any one of claims 1 to 6 is implemented.

9. A computer storage medium having computer executable instructions stored thereon, characterized in that: When the computer executable instructions are executed by a processor, the autonomous defect decision-making method for industrial inspection as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Medical image segmentation method based on dynamic deformable convolution and sliding window adaptive complementary attention mechanism

    CN116805318A

  • Multi-modal analysis method, system and equipment for industrial inspection scene and medium

    CN119128810A

  • Pumped storage power station construction anomaly detection method and system based on unmanned aerial vehicle image analysis

    CN119888507A

  • Intelligent decision management method and device for state monitoring and fault diagnosis of energy-saving equipment

    CN120067772A