Cross-modal knowledge reasoning method and device for industrial quality inspection and medium

By dynamically allocating the weight of multimodal input data in industrial quality inspection, semantic conflict detection and digestion, combining large models and dynamic hypergraph knowledge networks for knowledge reasoning, the real-time and accuracy problems of traditional cross-modal knowledge in industrial quality inspection are solved, and efficient defect detection and analysis are achieved.

CN120069096AActive Publication Date: 2025-05-30INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202510542904.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The traditional cross-modal knowledge inference method is affected by the fixed modal allocation weight method in the static knowledge base, and ignores the differences in different inference task scenarios. When cross-modal data conflicts, it is easy to cause missing inspection of key defects, which cannot meet the real-time and accurate quality inspection requirements in industrial quality inspection scenarios.

Method used

By obtaining the current inference task information, dynamic weight allocation is performed on the multimodal input data, multimodal semantic weight information is determined, and current fused semantic features are determined. Semantic conflict detection is performed on multimodal input data. When there is a semantic conflict, the weight is adjusted based on the multimodal semantic weight information to determine the conflict-dissolving modal weight information. Then, knowledge reasoning is performed through large models and pre-built dynamic hypergraph knowledge networks to determine the structured decision text.

Benefits of technology

Real-time fusion of information between different modal data is realized, the system's attention to task scenarios is enhanced, semantic conflicts are effectively eliminated, key defects are missed, the accuracy and reliability of inference results are improved, and the real-time and accuracy requirements of industrial quality inspection are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069096A_ABST
    Figure CN120069096A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a cross-modal knowledge reasoning method and device for industrial quality inspection and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining current reasoning task information, carrying out the dynamic weight distribution of multi-modal input data according to a task description text in the current reasoning task information, and carrying out the dynamic weight distribution of the multi-modal input data; determining multi-modal semantic weight information, and determining a current fusion semantic feature corresponding to the multi-modal input data; performing semantic conflict detection on the multi-modal input data, and when the multi-modal input data has a semantic conflict, performing weight adjustment on the multi-modal semantic weight information based on the multi-modal semantic weight information, and determining conflict resolution modal weight information; according to the conflict resolution modal weight information and the current fusion semantic features, knowledge reasoning is carried out through a large model and a pre-constructed dynamic hypergraph knowledge network, a structured decision text is determined, and the structured decision text comprises defect types, root cause analysis, association paths, decision suggestions and early warning signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and particularly to a cross-modal knowledge reasoning method, device, and medium for industrial quality inspection. Background Art

[0002] In recent years, the field of artificial intelligence has made remarkable progress in the direction of cross-modal knowledge reasoning technology, but still faces many technical bottlenecks in industrial quality inspection. Current research focuses on multi-modal pre-training models and cross-modal retrieval technology, whose core goal is to achieve feature space alignment through contrastive learning, but it relies on the assumption of static data distribution and lacks support for dynamic knowledge evolution.

[0003] In the industrial quality inspection scenario, industrial quality inspection needs to process high-dimensional heterogeneous data such as images (e.g., surface defects), sensor time-series data (e.g., temperature / pressure fluctuations), and unstructured text (e.g., process records) simultaneously, and the data distribution changes dynamically with the adjustment of production line parameters, equipment aging, etc. Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, and the fixed weight allocation strategy ignores the differences in task scenarios. In addition, factors such as sensor noise and light changes in the industrial environment often lead to contradictions in cross-modal data. For example, the image shows no scratches but the sensor indicates that the stress exceeds the limit. Using fixed conflict handling rules such as simple threshold filtering, and the static knowledge base cannot dynamically adjust the modal weights, which is prone to over-reliance on unreliable features, resulting in missed detection of key defects.

[0004] Therefore, the traditional cross-modal knowledge reasoning method, affected by the fixed modal allocation weight method in the static knowledge base, ignores the differences in different reasoning task scenarios, and is prone to the problem of missed detection of key defects when cross-modal data conflicts, and cannot meet the real-time and accurate quality inspection requirements in the industrial quality inspection scenario. Summary of the Invention

[0005] One or more embodiments of this specification provide a cross-modal knowledge reasoning method, device, and medium for industrial quality inspection, which are used to solve the following technical problems: The traditional cross-modal knowledge reasoning method, affected by the fixed modal allocation weight method in the static knowledge base, ignores the differences in different reasoning task scenarios, and is prone to the problem of missed detection of key defects when cross-modal data conflicts, and cannot meet the real-time and accurate quality inspection requirements in the industrial quality inspection scenario.

[0006] One or more embodiments of this specification adopt the following technical solutions: One or more embodiments of this specification provide a cross-modal knowledge reasoning method for industrial quality inspection. The method includes: obtaining current reasoning task information, dynamically allocating weights to multi-modal input data according to the task description text in the current reasoning task information to determine multi-modal semantic weight information, and determining the current fused semantic features corresponding to the multi-modal input data; detecting semantic conflicts in the multi-modal input data, and when there are semantic conflicts in the multi-modal input data, based on the multi-modal semantic weight information, adjusting the weights of the multi-modal semantic weight information to determine conflict resolution modal weight information; according to the conflict resolution modal weight information and the current fused semantic features, performing knowledge reasoning through a large model and a pre-constructed dynamic hypergraph knowledge network to determine structured decision-making text, where the structured decision-making text includes any one or more of defect type, root cause analysis, association path, decision suggestion, and warning signal.

[0007] One or more embodiments of this specification provide a cross-modal knowledge reasoning device for industrial quality inspection, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0008] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.

[0009] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Through the embodiments of this specification, with the pre-constructed dynamic hypergraph knowledge network, flexible adjustment can be made according to real-time production data and continuously updated industry rules. When new quality problems or data characteristics appear, the node and hyperedge relationships in the hypergraph can be updated to quickly adapt to changes, effectively solving the problem of lack of support for dynamic knowledge evolution; Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, while industrial quality inspection needs to process high-dimensional heterogeneous data such as images, sensor time-series data, and unstructured text. The embodiments of this specification obtain the current inference task information, perform dynamic weight assignment on multi-modal input data, determine the multi-modal semantic weight information, and determine the current fused semantic features, realizing the real-time fusion of information between different modal data; Traditional fixed weight assignment strategies ignore the differences in task scenarios. This technical solution performs dynamic weight assignment on multi-modal input data according to the task description text in the current inference task information, enabling the system to pay more attention to the modal data related to the task; Traditional methods adopt fixed conflict handling rules such as simple threshold filtering, which are prone to over-rely on unreliable features. The embodiments of this specification perform semantic conflict detection on multi-modal input data. When there is a semantic conflict, the weights are adjusted based on the multi-modal semantic weight information to determine the conflict resolution modal weight information. By reducing the weight of conflict modal data or increasing the weight of reliable modal data, a reasonable balance can be achieved between conflict data, making the inference process more inclined to rely on reliable data, effectively resolving semantic conflicts, avoiding missed detection of key defects, and improving the accuracy and reliability of the inference results. By comprehensively applying technologies such as dynamic hypergraph knowledge networks, large models, and conflict resolution, knowledge reasoning can be performed quickly and accurately, and structured decision-making texts can be determined. The local update mechanism of the dynamic hypergraph knowledge network only adjusts the weights of associated hyperedges based on the temporal graph convolutional network (T-GCN), without global reconstruction, and the single update takes a short time, ensuring the real-time performance of the system; Through dynamic weight assignment, semantic conflict detection and resolution, and the powerful semantic understanding and reasoning capabilities of large models, the accuracy of the inference results is improved. In complex industrial quality inspection scenarios, quality problems of products can be detected in a timely manner, defect types and root causes can be accurately analyzed, and reasonable decision-making suggestions and warning signals can be given, effectively meeting the requirements of industrial quality inspection for real-time performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments described in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 Schematic flowchart of a cross-modal knowledge reasoning method for industrial quality inspection provided by an embodiment of this specification; Figure 2 Schematic flowchart of generating a hybrid reasoning path provided by an embodiment of this specification; Figure 3 Schematic flowchart of multi-modal dynamic knowledge modeling provided by an embodiment of this specification; Figure 4 Schematic structural diagram of a cross-modal knowledge reasoning device for industrial quality inspection provided by an embodiment of this specification. Detailed implementation manners

[0011] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0012] An embodiment of this specification provides a cross-modal knowledge reasoning method for industrial quality inspection. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 Schematic flowchart of a cross-modal knowledge reasoning method for industrial quality inspection provided by an embodiment of this specification, as Figure 1 shown, mainly including the following steps: Step S101: Obtain the current reasoning task information, and according to the task description text in the current reasoning task information, perform dynamic weight allocation on the multi-modal input data to determine the multi-modal semantic weight information and determine the current fusion semantic features corresponding to the multi-modal input data.

[0013] In an embodiment of this specification, the current reasoning task information is obtained. Here, the current reasoning task information includes the task description text corresponding to this reasoning task and the multi-modal input data. The task description text is used to distinguish the task type corresponding to the current reasoning task, such as a defect classification task, a root cause analysis task, etc. The multi-modal input data can be various modalities such as image data, sensor data, and text data.

[0014] According to the task description text in the current inference task information, perform dynamic weight allocation on the multi-modal input data to determine the multi-modal semantic weight information, specifically including: generating a dynamic task query vector corresponding to the current inference task through the task description text; constructing cross-modal key-value pairs for each modality in the multi-modal input data; and performing dynamic weight allocation on multiple modalities according to the dynamic task query vector and the cross-modal key-value pairs to determine the multi-modal semantic weight corresponding to each modality.

[0015] In one embodiment of this specification, a dynamic query vector is generated according to the current inference task type, such as "defect classification task", "root cause analysis task". The semantic features of the task description are extracted through a pre-trained semantic encoder and mapped to a 128-dimensional dynamic task query vector to represent the core requirements of the current task (for example, "root cause analysis" needs to associate material properties with process parameters). The task description text is mapped to a query vector through a Task Encoder: , where, is the semantic embedding extracted from the task description by the large model, is a learnable parameter matrix, is the bias term. For example, in the root cause analysis task, to determine the cause of microcracks, the encoder will encode the task text and output a semantic vector, which is vectorized into , and then after mapping and adding the bias term, a query vector is generated, pointing to semantics such as "material properties" and "process parameters".

[0016] Extract the scratch area features through a convolutional neural network to generate a key vector and a value vector . The sensor modality uses a temporal encoder (LSTM) to process the temperature and pressure fluctuation data to generate a key vector k sensor and a value vector v sensor . Parse the process record text to extract keywords (such as "material batch X has insufficient hardness") to generate a text modality key vector k text and a value vector v text . Taking the process of constructing key-value pairs with image, sensor, and text modality features as an example for illustration. In the image modality, , in the sensor modality, , in the text modality , where, is the feature vector of each modality, and ∈ ℝ^{d×d} is the modality projection matrix. The task query vector q task is respectively matched with the key vectors of each modality, and the initial attention scores are calculated through scaled dot product to obtain the weights of each modality , where \(i\in\{image, sensor, text\}\), and \(d = 64\) is the dimensional scaling factor. The final weighted fusion feature is . For example, in the surface scratch detection task, the system dynamically assigns weights: the image modality = 0.7 (relying on visual features); the sensor modality = 0.2 (monitoring temperature anomalies); the text history record modality = 0.1 (referring to historical quality inspection records).

[0017] Through the above technical solutions, based on the generation of dynamic task query vectors, the importance weights of multi-modal data can be automatically adjusted according to the current task type, such as "defect classification", "root cause analysis", and "process optimization". For example, in the root cause analysis task, the weight of sensor time-series data is increased to 0.65 (default 0.3), and the weight of text process records is increased to 0.25 (default 0.1), effectively capturing deep features strongly related to material properties and process parameters, and improving the accuracy of defect root cause localization to 98.5%; through cross-modal key-value pair construction and dynamic weight assignment, the semantic consistency between different modalities can be quantified.

[0018] Step S102, perform semantic conflict detection on the multi-modal input data. When there is a semantic conflict in the multi-modal input data, based on the multi-modal semantic weight information, adjust the weights of the multi-modal semantic weight information to determine the conflict resolution modal weight information.

[0019] In an embodiment of the present specification, for possible semantic conflicts, such as the image showing "no scratches" while the sensor data indicates "pressure overlimit" or the image showing cracks but the laser point cloud having no deformation features, based on the semantic distance threshold for detection. When the threshold is greater than 1.5, it is determined that a semantic conflict has occurred. When a conflict is detected, it will trigger the backpropagation optimization of the feature alignment module, recalibrate the modal representation, and combine the associated paths in the semantic network to generate correction suggestions, which can reduce the false alarm rate of quality inspection and significantly improve the consistency of multi-modal data.

[0020] In an embodiment of the present specification, the conflict detection trigger conditions are as follows: calculate the cross-modal feature distance and determine the semantic consistency , where is an indicator function that outputs 1 when the condition is met and 0 otherwise, is the embedding vector of image, sensor, or text features, and \(\tau = 1.5\) is the preset distance threshold determined by expert experience data. When there exists any > \(\tau\), it is determined that a semantic conflict has occurred.

[0021] Based on the multimodal semantic weight information, the multimodal semantic weight information is adjusted to determine conflict resolution modal weight information, specifically including: determining at least two conflicting modes among multiple modes, and obtaining the cross-modal distance index between the conflicting modes; determining the weight penalty factor corresponding to the conflicting modes through the cross-modal distance index and a preset distance index threshold; and adjusting the weights of the multiple modes respectively according to the weight penalty factor and the multimodal semantic weight information to determine the conflict resolution modal weight information.

[0022] When there is a semantic conflict in the multimodal input data, two conflicting modes with conflicts are determined among the multiple modes, and the cross-modal distance index between these two conflicting modes obtained during the conflict judgment process is obtained, that is . Through the cross-modal distance index and a preset distance index threshold τ, the weight penalty factor corresponding to the conflicting modes is determined, and the penalty factor is , where = 2.0, = 1.5. The weights of the multiple modes are adjusted respectively using the weight penalty factor and the multimodal semantic weight information of the corresponding modes obtained in step S101 to determine the conflict resolution modal weight information.

[0023] According to the weight penalty factor and the multimodal semantic weight information, the weights of the multiple modes are adjusted respectively to determine the conflict resolution modal weight information, specifically including: imposing a weight penalty on the conflicting modes according to the weight penalty factor to determine the current penalty weight parameter corresponding to each of the conflicting modes; obtaining the multimodal semantic weight information corresponding to each of the conflicting modes, and determining the conflict penalty weight through the multimodal semantic weight information corresponding to the conflicting modes and the current penalty weight parameter; based on the conflict penalty weight, boosting the weights of the reliable modes among the multiple modes to determine the conflict boosting weight corresponding to the reliable modes; and determining the conflict resolution modal weight information through the current penalty weight parameter corresponding to each of the conflicting modes and the conflict boosting weight corresponding to the reliable modes.

[0024] In one embodiment of the present specification, according to the weight penalty factor, weight penalty is imposed on the conflicting modality, and the current penalty weight parameter corresponding to the conflicting modality is determined by the product of the weight penalty factor and the original weight in the multi-modal semantic weight information corresponding to this conflicting modality. The conflict penalty weight corresponding to each conflicting modality is determined by the difference between the multi-modal semantic weight information corresponding to each conflicting modality and the corresponding current penalty weight parameter. Based on the conflict penalty weight, the weight of the reliable modality among the multiple modalities is increased to determine the conflict promotion weight corresponding to the reliable modality, where the promotion amount of the conflict promotion weight should be the same as the penalty amount of the conflict penalty weight. In addition, it should be noted that after summarizing the conflict penalty weights of each conflicting modality for which weight conflict penalty is imposed, the cumulative weight penalty amount corresponding to the conflicting modality is determined. The conflict resolution modality weight information is determined by the current penalty weight parameter corresponding to each such conflicting modality and the conflict promotion weight corresponding to the reliable modality.

[0025] Assume that the initial weight of the image modality is 0.7, the sensor modality is 0.2, and the text modality is 0.1. The detected L2 distance between the image and the sensor is 1.8, and the distance between the image modality and the sensor modality is greater than 1.5, indicating that there is a semantic conflict, and the image weight and the sensor weight need to be adjusted. The image weight is calculated by the formula , that is, it is adjusted to 0.38, which is the current penalty weight parameter corresponding to the image modality. The sensor weight is also adjusted to 0.11 by the above formula. The original weight of the image modality is 0.7, and after being adjusted to 0.38, it is reduced by 0.32, that is, the conflict penalty weight under the image modality is 0.32. The original weight of the sensor modality is 0.2, and after being adjusted to 0.11, it is reduced by 0.09, that is, the conflict penalty weight under the sensor modality is 0.09. The cumulative conflict penalty weight corresponding to the two is 0.32 + 0.09 = 0.41, that is, the cumulative weight penalty amount is 0.41. Since there is no conflict in the text modality, the decision proportion of the reliable modality (text) needs to be enhanced, and the weight promotion amount should be the same as the cumulative weight penalty amount. Therefore, the text modality weight is 0.1 + 0.41 = 0.51, and the conflict promotion weight is obtained. According to the above steps, the current penalty weight parameter corresponding to the conflicting modality and the conflict promotion weight corresponding to the reliable modality are obtained, and the conflict resolution modality weight information corresponding to each modality is determined.

[0026] Through the above technical solutions, semantic conflicts are detected based on the semantic distance threshold. When a conflict is detected, the feature alignment module is triggered to perform backpropagation optimization, recalibrate the modal representation, and generate correction suggestions by combining the association paths in the semantic network, which can effectively avoid misjudgments caused by semantic conflicts, thereby reducing the false alarm rate of quality inspection and improving the accuracy of quality inspection. After detecting semantic conflicts, a series of operations are performed to adjust the multi-modal semantic weight information, including determining the conflicting modality, obtaining the cross-modal distance metric, calculating the weight penalty factor, performing weight penalty and promotion, etc., which significantly improves the semantic consistency between multi-modal data and ensures the reliability and accuracy of the data. According to the situation of semantic conflicts, weight penalty is imposed on the conflicting modalities, while weight promotion is performed on the reliable modalities. This dynamic weight adjustment method can reasonably allocate the proportion of each modality in the decision-making, enhance the role of reliable modalities, and reduce the impact of conflicting modalities, thus making the decision-making of the entire multi-modal system more reasonable and reliable.

[0027] Step S103, according to the conflict resolution modal weight information and the current fused semantic features, perform knowledge reasoning through the large model and the pre-constructed dynamic hypergraph knowledge network to determine the structured decision text.

[0028] Among them, the structured decision text includes any one or more of defect type, root cause analysis, association path, decision suggestion, and warning signal.

[0029] In an embodiment of the present specification, a hybrid inference path generation is performed according to the multi-modal input data in the absence of semantic conflicts. Figure 2 It is a schematic flowchart of a hybrid inference path generation provided by an embodiment of the present specification. As Figure 2 shown, it fuses symbolic logic and neural networks to solve the problems of path redundancy and low efficiency of traditional inference methods in complex multi-modal scenarios. Taking the industrial quality inspection scenario as an example, the input data includes a high-definition image of the product surface (such as the detected scratch length is 1.2 mm), laser sensor point cloud data (the deformation depth is 0.05 mm), and production record text (such as "material batch Y, heat treatment temperature deviation ±10°C"). The symbolic reasoning layer constructs a production rule base based on industry detection standards. For example: "scratch length > 0.5 mm → insufficient material hardness → re-inspection required", which is represented symbolically as: "IF (scratch length > 0.5 mm) ∧ (material hardness < 50 HRC) THEN trigger the re-inspection of the heat treatment process", thereby generating candidate inference paths.

[0030] Retrieve associated paths from a multi-dimensional semantic network through querying with the Semantic Web Query Language (SPARQL Protocol and RDF Query Language, SPARQL). For example, when inputting the defect feature "excessive scratches", retrieve the path "excessive scratch length → abnormal material hardness → trace the heat treatment process → adjust the temperature parameter", and encode it into a natural language description (such as "It is detected that the scratch length is 1.2 mm, the material hardness is lower than the standard value, and it is recommended to check the temperature of the heat treatment furnace"). For instance, input the RDF triples of the candidate path into a large model to generate a natural language description through template filling. Template example: It is detected that [defect feature], [root cause analysis], it is recommended [decision suggestion]. RDF triples are only a temporary expression form of the reasoning result, generated through dynamic hyper-edge weights and the context understanding of the large model, which is essentially different from the static triple storage of traditional knowledge graphs. The neural inference layer uses the large model to perform semantic scoring on the candidate paths. Concatenate the path description with the input question (such as "Analysis of the cause of scratches?") and input it into the model, and calculate the matching degree through a scoring function , where is the input question embedding vector, such as "Possible causes of a scratch length of 1.2 mm", is the candidate path description, is the semantic matching layer of the large model. Screen the Top-3 paths with high confidence, such as s 1 = 0.95, s 2 = 0.88, s 3 = 0.82, to ensure that the results are both accurate and interpretable.

[0031] To further optimize the inference efficiency, introduce a reinforcement learning mechanism to dynamically adjust the strategy. Design a multi-objective reward function , where the detection accuracy is verified by historical quality inspection results, and the response time is in seconds, which needs to meet the production line beat requirements. The policy network parameters are updated by the Proximal Policy Optimization (PPO) algorithm. The policy network is an independently trained reinforcement learning model, and its training goal is to learn how to dynamically select the optimal inference path (such as preferentially performing symbolic inference or neural inference) according to multi-modal input features (such as image scratch length, sensor temperature data, etc.) by interacting with the knowledge base environment. During the training process, the policy network processes state-action pairs (such as the state being a combination of defect features and the action being to select the path of "re-inspection process parameters" or "trigger laser repair"), dynamically balancing the detection accuracy and real-time requirements. This hybrid mechanism reduces the path search space by 60%. In the high-speed production line scenario, the response time for complex defect determination is reduced from 5 seconds to 1.5 seconds, and the detection accuracy is increased to 98.5% (85% for traditional methods). To further understand the role of knowledge reasoning, in combination with the industrial scenario, the output of knowledge reasoning is illustrated by examples as follows. When the system detects a micro-crack (length 0.3 mm) on the surface of an aero-engine blade, the inference engine generates the following results: (1) Defect classification: Micro-crack (confidence 0.95); (2) Root cause analysis: Abnormal coating thermal expansion coefficient (probability 89%); (3) Decision suggestion: Adjust the sintering process temperature to 1250 °C and perform X-ray flaw detection verification; (4) Warning signal: Trigger the automatic shutdown of the production line to avoid batch quality accidents.

[0032] It should be noted that before performing hybrid reasoning, it is necessary to pre-perform multi-modal dynamic knowledge modeling and construct a dynamically evolving multi-modal knowledge representation framework to solve the static problem of traditional knowledge graphs. Existing knowledge graphs rely on binary relations, head entity - relation - tail entity, and it is difficult to express the joint semantics of multi-modal data. For example, "image scratch + sensor data + text record" needs to be split into multiple triples. In the embodiments of this specification, the knowledge representation framework is implemented using a hypergraph structure, and through the multi-nodes of the hypergraph, it supports complex cross-modal reasoning and improves the path retrieval efficiency.

[0033] Figure 3 is a schematic flow diagram of a multi-modal dynamic knowledge modeling provided by the embodiments of this specification, as Figure 3As shown, in the industrial quality inspection scenario, the input data includes sensor time-series data such as product surface images, temperature, and pressure, as well as quality inspection report texts. First, through a multi-modal pre-trained large model, here a domestic open-source multi-modal large model can be used to achieve cross-modal feature extraction. For the multi-modal, image modality, the commonly used Vision Transformer (ViT) model in the field can be used to extract local texture features (such as scratches and depressions) in blocks, and a 224×224 pixel image is divided into 16×16 blocks. Each block extracts a d=768-dimensional feature vector through 12 layers of Transformer encoders; the sensor data encodes the time-series features through an LSTM network, with the time step set to t=50 and the hidden layer dimension h=256; the text modality uses the BERT model to parse the keywords in the quality inspection report (such as "qualified" and "defect type A"), and outputs a 768-dimensional semantic vector of the [CLS] token. The multi-modal features interact in the cross-modal attention layer of UNITER to generate a unified semantic representation. At the same time, to achieve cross-modal semantic alignment, a two-tower contrastive learning architecture is designed and trained through the Adam optimizer, using the contrastive loss function , to align the image, sensor, and text features in a unified embedding space, where B is the number of samples in the current batch, is the image feature, is the corresponding text description feature, =0.05 is the temperature coefficient, which is determined through grid search. The two-tower structure includes an image-sensor tower and a text tower, and the semantic alignment between modalities is constrained by the contrastive loss. Among them, the positive sample pairs are the multi-modal data of the same quality inspection sample (such as image scratches + corresponding sensor temperature anomaly + text report), and the negative sample pairs are randomly sampled different sample combinations.

[0034] Based on feature alignment, a multi-dimensional semantic network is constructed. This network uses a hypergraph structure to dynamically represent cross-modal relationships, serving as the underlying data structure of the multi-dimensional semantic network. A hypergraph allows a hyperedge to connect any number of nodes, breaking through the limitation of traditional knowledge graphs that can only express knowledge through binary relationships (i.e., head entity - relationship - tail entity). For example, in the industrial quality inspection scenario, a hyperedge can simultaneously connect an image feature node (such as "scratch length 1.2mm"), a sensor node (such as "abnormal temperature 85°C"), and a text node (such as "material batch Y") to fully represent the cross-modal association of "scratch - temperature - material". The hypergraph node set includes entities (such as "defect type A", "temperature threshold"), concepts (such as "detection standard"), and modal features (such as "image texture feature"), and the hyperedge set represents cross-modal associations (such as "image scratch → sensor temperature anomaly → quality inspection report conclusion"). In addition, when a new defect pattern is detected, such as an undefined candidate defect description generated by an inference large model, calculate its semantic similarity with existing nodes , calculate the similarity using the cosine function, denote the newly added candidate defect node, u is the set of existing knowledge base nodes, is the candidate node 's feature vector, is the feature vector of the existing node. If the similarity exceeds the threshold of 0.6, it is automatically added to the network; at the same time, a temporal graph convolutional network (T-GCN) is used to dynamically adjust the hyperedge weights , where and are the hidden states of the nodes at both ends of the hyperedge, is the LeakyReLU activation function, is the learnable weight matrix used to map the hidden state of the node to the hyperedge weight space; || is the vector concatenation operation to fuse the semantic information of the two end nodes. Through project the concatenated high-dimensional vector into the weight space; represents the weight value of hyperedge e at time step t + 1, which is used for subsequent inference path retrieval and conflict resolution to quantify the importance of this hyperedge in multi-modal semantic association. T-GCN realizes an increase in the dynamic update frequency and improves the real-time response ability of quality inspection by only making local weight adjustments to associated hyperedges (instead of global graph reconstruction), combined with the incremental iteration and lightweight design of temporal hidden states to compress the time consumption of a single update

[0035] Through the above technical solutions, the hypergraph structure is adopted to break through the binary relation limitation of the traditional knowledge graph, support the joint semantic expression of multi-modal data such as images, sensors, and texts, make the knowledge representation more suitable for complex industrial scenarios, and achieve cross-modal semantic alignment through contrastive learning of the two-tower structure and the contrastive loss function, ensuring effective interaction of different modal features in a unified space and solving the problem of modal fragmentation in traditional methods; the combination of symbolic logic and neural networks not only retains the interpretability of rule reasoning, such as "scratch length > 0.5mm → insufficient material hardness", but also uses the generalization ability of neural reasoning to process unknown defect patterns, and reinforcement learning dynamically selects the optimal reasoning path, reducing the response time of complex defect determination and the path search space, meeting the real-time requirements of high-speed production lines; new defect patterns are automatically added through semantic similarity matching, and the hyperedge weights are locally adjusted in combination with the temporal graph convolutional network, avoiding global graph reconstruction, and compressing the update time-consuming to the millisecond level. In the industrial quality inspection scenario, the detection accuracy is improved, the response time is shortened, and the production line beat requirements are met; the reward function combines the detection accuracy and the response time, and balances the contradiction between the two through reinforcement learning; through the multi-node connection characteristics of the hypergraph structure and the evolution mechanism of the temporal graph convolutional network, the limitation of binary relation expression of the traditional knowledge graph is solved, the joint semantic modeling of multi-modal data such as images, point clouds, and texts is realized, and the real-time dynamic update of the knowledge base is supported; the adaptive ability to new defect patterns is significantly improved, and the problems of missed detection caused by lagging knowledge update and cross-modal association breakage of complex defects in industrial scenarios can be effectively solved.

[0036] It should be noted that the conflict resolution mechanism in the embodiments of this specification is essentially different from knowledge resolution. Conflict resolution refers to the semantic contradiction between multi-modal data during real-time reasoning, such as the conflict between image and sensor data. Through feature calibration and rule reasoning, correction suggestions are generated; knowledge resolution belongs to the entity ambiguity processing in the knowledge base construction stage, such as merging synonymous entities. In the embodiments of this specification, entity ambiguity processing in the knowledge base construction stage is achieved through dynamic pruning and merging of hypergraph nodes. Taking the industrial quality inspection scenario as an example, it is necessary to fuse image defect features, sensor abnormal data, and text report conclusions. The present invention designs a multi-head attention mechanism to dynamically allocate modal weights, generate query vectors according to the task type, calculate the attention scores of the key vectors of each modality, and generate joint features after weighted fusion. For possible semantic conflicts, such as the image shows "no scratch" while the sensor data indicates "pressure overlimit" or the image shows a crack but the laser point cloud has no deformation feature, detection is performed based on the semantic distance threshold. When the threshold is greater than 1.5, it is determined that a semantic conflict has occurred. When a conflict is detected, the method will trigger the backpropagation optimization of the feature alignment module, recalibrate the modal representation, and generate correction suggestions in combination with the associated paths in the semantic network. This mechanism can reduce the false alarm rate of quality inspection and significantly improve the consistency of multi-modal data.

[0037] The conflict resolution algorithm adopts a two-stage correction strategy. In the feature calibration stage, the contrast loss function of the multi-modal dynamic knowledge modeling link is used. In the inference correction stage, correction suggestions are generated based on the knowledge base rules and the large language model. In the calibration mechanism of the contrast loss function in the feature calibration stage, the feature vectors of conflicting modalities (such as images and sensors) are aligned to a unified semantic space, reducing the cross-modal feature distance, so that the cross-modal data of the same object (such as scratch images and corresponding temperature sensor data) are close in the embedding space, and the features of different objects are far away. The multi-modal data is mapped to a unified embedding space through a two-tower structure (image-sensor tower and text tower). The cosine similarity of positive sample pairs is calculated and normalized by Softmax. The contrast loss is minimized to push positive sample pairs closer and negative sample pairs farther away. In the inference correction stage, knowledge-driven semantic correction retrieves the inference path associated with the calibrated features from the dynamic hypergraph knowledge base (such as "pressure overlimit → material fatigue → internal damage"). The path is encoded as a natural language instruction (such as "It is recommended to perform X-ray flaw detection"). The correction suggestions in the inference correction stage are based on the features calibrated in the feature calibration stage (such as the aligned sensor features). If not calibrated in the feature calibration stage, the sensor features may be associated with the "temperature anomaly" node, generating incorrect suggestions; after calibration, the sensor features are associated with the "internal damage" node, generating correct suggestions. The execution result of the correction suggestion (such as X-ray confirmation of internal damage) can be used as a supervision signal to adjust the parameters of the contrast loss function (such as the temperature coefficient τ) through reinforcement learning.

[0038] An example of the process of retrieving the inference path associated with the calibrated features from the dynamic hypergraph knowledge base is as follows. First, perform associated path retrieval to retrieve the Top-3 high-weight paths from the multi-dimensional semantic network through an extended SPARQL query; then, perform large model semantic generation, input the path information into the large model, and generate natural language suggestions through a structured prompt template; finally, perform post-processing of the execution result: use the zero-shot classification ability of the large model to perform keyword verification and confidence annotation on the generated suggestions. For example, it can be verified whether the necessary keywords (such as "retest", "ISO standard") are included in the suggestions to ensure compliance with industry standards and operating procedures. The conflict resolution in the inference correction stage is described in detail below.

[0039] Based on the conflict resolution modal weight information and the current fused semantic features, knowledge reasoning is carried out through a large model and a pre-constructed dynamic hypergraph knowledge network to determine the structured decision-making text, specifically including: determining candidate reasoning paths in the dynamic hypergraph knowledge network according to the current fused semantic features; generating path description information corresponding to each candidate reasoning path, splicing the path description information with the pre-determined corresponding input problem description information, scoring each candidate reasoning path through the large model, and determining the path matching degree score of each candidate reasoning path; correcting the path matching degree score of each candidate reasoning path through the conflict resolution modal weight information to determine the target reasoning path, so as to determine the structured decision-making text.

[0040] In one embodiment of the present specification, based on the currently fused multi-modal semantic features, image defect regions, sensor abnormal data, and text process records, possible reasoning paths are retrieved in the dynamic hypergraph knowledge network. According to predefined industry rules, such as "temperature exceeding the standard needs to be associated with material properties", a basic candidate path, "temperature anomaly → material fatigue → recheck process parameters", is generated. Through the dynamic association relationship of the hypergraph network, such as multi-hop nodes connected by hyperedges, the path range is automatically expanded. For example, if a surface scratch is detected, associated nodes such as "insufficient coating thickness" or "excessive assembly stress" may be retrieved and extended, forming multiple candidate paths. Each candidate path is converted into a natural language description and semantically associated with the user input problem. The node and hyperedge relationships in the path are transformed into readable text. For example, the path "temperature anomaly → material fatigue → recheck process parameters" is described as "temperature exceeding the standard is detected, associating the possibility of material fatigue, and it is recommended to recheck the heat treatment parameters". The user problem (such as "what is the root cause of the current defect?") is spliced with the path description to form a complete input. For example: "Defect root cause analysis: temperature exceeding the standard is detected, associating the possibility of material fatigue, and it is recommended to recheck the heat treatment parameters". The spliced text is input into the pre-trained large model, and the matching degree between the path and the problem is evaluated through the semantic understanding ability inside the model. The model analyzes the logical relevance between the path description and the problem and outputs a confidence score (0-1). For example, the path "temperature anomaly → material fatigue" scores 0.92 in the root cause analysis task, while it may be only 0.45 in the surface detection task. The top 3 paths with the highest scores are retained to ensure both accuracy and diversity of the results. Combining the weights after multi-modal conflict resolution (such as suppressing the image weight when there is a contradiction between the image and sensor data), the path scores are dynamically adjusted. If there is a conflict in the modal data on which a certain path depends (such as the sensor shows normal temperature but the text record is abnormal), the score of this path is reduced. The scores of paths depending on high-weight modalities are increased to ensure that reliable data sources are preferentially adopted in the final decision. When determining the structured decision-making text, the paths before and after conflict resolution can be displayed simultaneously.

[0041] Through the above technical solutions, the hypergraph structure supports multi-hop node expansion, breaks through the limitations of traditional binary relations, covers more complex causal chains, and jointly generates candidate paths with predefined rules and hypergraph dynamic relations to ensure the balance between rule authority and data-driven flexibility; by splicing the path description with the user's question, the large model can specifically evaluate the path relevance, retain high-confidence paths, avoid the one-sidedness of a single path, and ensure decision-making robustness through diversity; when multi-modal data conflicts, reduce the scores of conflicting modal paths; improve the scores of paths that rely on high-confidence modalities to ensure that reliable data sources are preferentially adopted in decision-making; weight adjustment is synchronized with path scoring, and there is no need for offline retraining to meet the real-time requirements of the production line. In scenarios where multi-modal data is noisy or conflicting, the conflict resolution mechanism can improve the decision-making accuracy; integrate the deterministic reasoning of the symbolic rule engine, the semantic understanding ability of the large language model, and the dynamic decision-making mechanism of reinforcement learning to form an interpretable and high-precision hybrid reasoning path generation strategy, which can significantly reduce redundant computing overhead while ensuring detection accuracy and meet the stringent requirements of the high-speed production line for real-time response.

[0042] According to the current fused semantic features, determine candidate inference paths in the dynamic hypergraph knowledge network, specifically including: determining initial candidate paths according to the current fused semantic features and predefined industry rules; performing associated path retrieval in the dynamic hypergraph knowledge network through the path nodes in the initial candidate paths to generate a set of candidate inference paths, where the set of candidate inference paths includes multiple candidate inference paths, and each candidate inference path includes an RDF triple chain connected by a dynamic hyperedge.

[0043] In one embodiment of the present specification, an initial candidate path is generated by combining the currently fused semantic features (such as detecting that the length of the surface scratch exceeds the standard) with predefined industry rules (such as "if the scratch length > 0.5mm, a process review needs to be triggered"). Example rules are as follows: "IF scratch length > 0.5mm THEN associate with material hardness detection". An initial path is generated according to the rules. For example: "Surface scratch exceeds standard → Insufficient material hardness → Recheck heat treatment process". If abnormal temperature rise (sensor data) is currently detected, the path is extended to: "Surface scratch exceeds standard → Abnormal temperature rise → Insufficient material hardness". Based on the nodes in the initial path (such as "Surface scratch exceeds standard", "Insufficient material hardness"), associated paths are retrieved in the dynamic hypergraph knowledge network to form a set of candidate inference paths. Expanding layer by layer along the hyperedge connection relationship, the first hop is "Surface scratch exceeds standard → Insufficient material hardness", the second hop is "Insufficient material hardness → Abnormal heat treatment furnace temperature", and the third hop is "Abnormal heat treatment furnace temperature → Calibrate temperature control parameters". Finally, a path chain "Surface scratch exceeds standard → Insufficient material hardness → Abnormal heat treatment furnace temperature → Calibrate temperature control parameters" is generated. Only the associated paths with weights exceeding the threshold (such as 0.6) are retained to ensure the reliability of the paths. For example, the weak association path "Insufficient material hardness → Excessive assembly stress" with a weight of 0.3 is removed. Each candidate path is composed of a chain of RDF triples connected by dynamic hyperedges, representing a complete causal relationship chain. An example of an RDF triple chain is as follows: triple 1 (Surface scratch exceeds standard, associated defect, Insufficient material hardness), triple 2 (Insufficient material hardness, root cause analysis, Abnormal heat treatment furnace temperature), triple: (Abnormal heat treatment furnace temperature, solution, Calibrate temperature control parameters). The RDF triple chain is converted into a natural language description. For example: "It is detected that the surface scratch exceeds the standard, associated with the problem of insufficient material hardness. The root cause analysis is the abnormal heat treatment furnace temperature, and it is recommended to calibrate the temperature control parameters." Through the above technical solution, by means of pre-defined industry rules, it is possible to make full use of professional knowledge and standards in the industry, ensure that in the face of specific quality problems, an initial judgment can be made based on established industry best practices, generate reasonable initial candidate paths, effectively avoid deviating from industry standards during the reasoning process, improve the accuracy and professionalism of the reasoning results. At the same time, combined with the currently fused semantic features, the initial path is dynamically expanded. This combination of real-time data and industry rules enables the reasoning path to closely fit various situations in the actual production scenario, more comprehensively reflecting the potential associations of product quality problems; using the association relationships between nodes in the dynamic hypergraph knowledge network for path retrieval can fully explore the potential connections between data. By starting from the nodes in the initial path and expanding layer by layer along the hyperedge connection relationships, a long-chain reasoning path is formed, greatly expanding the depth and breadth of reasoning. Compared with traditional simple association reasoning, it can discover causal relationships at more levels and angles, thus providing more comprehensive clues for solving complex quality problems. By setting a weight threshold to screen the associated paths, it is ensured that the retained paths have high reliability, avoiding interference from invalid or low-value paths to the reasoning results, and further improving the quality of the reasoning path set; each candidate path is composed of an RDF triple chain connected by dynamic hyperedges. This structured representation clearly defines the relationships between each node, and the RDF triple chain can accurately describe the complete causal relationship chain.

[0044] Through the conflict resolution modal weight information, the path matching degree score of each candidate reasoning path is corrected to determine the target reasoning path, specifically including: determining at least one reasoning modality in each candidate reasoning path to determine the conflict resolution modal weight corresponding to each reasoning modality; determining the correction factor of the candidate reasoning path through the conflict resolution modal weight corresponding to each reasoning modality; on the basis of the path matching degree score of the candidate reasoning path, using the correction factor to determine the path score corresponding to each candidate reasoning path; among multiple candidate reasoning paths, determining the candidate reasoning path with the highest path score as the target reasoning path.

[0045] In one embodiment of this specification, for each candidate inference path, identify the key modal data it depends on (such as images, sensors, text), and obtain the modal weight information after conflict resolution. For example, path A depends on an image (the conflict resolution weight corresponding to the image modality is 0.38) and sensor data (weight 0.11), and path B depends on a sensor (the conflict resolution weight corresponding to the sensor modality is 0.11) and a text record (the conflict resolution weight corresponding to the text modality is 0.51). Based on the conflict resolution weights of each modality in the path, calculate the comprehensive correction factor of this path. If there is a conflict in the modalities on which the path depends (such as a contradiction between sensor and image data), the corresponding modal weight will be suppressed; if the modal data is reliable (such as the text record is consistent with the sensor), the weight will be maintained or increased. Determine the correction factor corresponding to this candidate path through the sum of the conflict resolution weights of the key modalities on which each candidate path depends, and obtain the path score corresponding to the candidate inference path by multiplying this correction factor by the path matching degree score, and determine the path with the highest path score as the target inference path. The core role of conflict resolution is to suppress unreliable paths. If there are serious conflicts in the modal data on which a path depends (such as the image shows no scratches, but the sensor indicates that the pressure exceeds the limit), its correction factor will be significantly reduced, thus avoiding misselection.

[0046] Through the above technical solution, by determining the inference modalities and their corresponding conflict resolution modal weights in each candidate inference path, the reliability of different modal data can be fully considered. Through the conflict resolution mechanism, the weight of conflicting modal data is suppressed, and the scores of candidate inference paths that depend on conflicting modal data will be affected, while the scores of paths based on reliable and consistent modal data will be relatively increased. Thus, it effectively avoids incorrect inferences caused by modal conflicts, improves the correctness of the inference result, dynamically adjusts the scores of candidate inference paths according to the real-time conflict resolution modal weights, corrects the path matching degree score by updating the modal weights in real time and calculating the correction factor based on this, so that the inference result can better adapt to environmental changes. Among many candidate inference paths, select the one with the highest path score as the target inference path. The path score comprehensively considers the conflict resolution modal weights. Therefore, the target inference path can more reasonably weigh the contributions of different modal data, avoids the limitations of selecting a path solely based on the path matching degree score, makes the finally determined inference path more in line with the actual situation, and enhances the reliability of the inference.

[0047] With the development of industrial production lines, undefined defect types usually occur. Traditional knowledge bases rely on manual rule definition, which is costly to update. The rule engines in traditional solutions cannot cover new types of defects, and static graphs need to be reconstructed offline. After performing association path retrieval in this dynamic hypergraph knowledge network, the method further includes: when the association path is not retrieved in the dynamic hypergraph knowledge network, parsing the multimodal input data through a large model to generate candidate defect description information for constructing candidate nodes; obtaining the existing hypergraph node set and existing hyperedge set in the dynamic hypergraph knowledge network, and calculating the maximum semantic similarity between the candidate nodes and the existing hypergraph node set; according to the maximum semantic similarity, performing a node update operation in the hypergraph node set of the dynamic hypergraph knowledge network to determine the operating hypergraph nodes, where the node update operation includes a similar node merging operation and a new node creation operation; determining at least one associated hyperedge corresponding to the operating hypergraph nodes, and using a temporal graph convolutional network to locally adjust the weights of the dynamic hyperedges to determine the dynamic hyperedge weights corresponding to the associated hyperedges, so as to dynamically update the dynamic hypergraph knowledge network.

[0048] In one embodiment of this specification, when the dynamic hypergraph knowledge network does not retrieve an inference path associated with the current multimodal input data (such as image defect features, sensor anomaly data, text report conclusions), the following processing flow is initiated: jointly parsing the multimodal input data using a pre-trained multimodal model to extract cross-modal semantic features and generate candidate defect information described in natural language (such as "surface microcracks accompanied by abnormal internal temperature rise"). Encoding the candidate defect description as a semantic vector and adding it as a candidate node to the dynamic hypergraph knowledge network, where the node attributes include defect type, associated modal features, and confidence parameters. Calculating the cosine similarity between the candidate node and all nodes in the existing hypergraph node set , and screening the maximum similarity value. If the maximum similarity ≥ 0.6 (configurable), perform similar node merging, merging the candidate node into the most similar existing node and inheriting its associated hyperedges; if the similarity ≥ 0.4 but < 0.6, trigger the manual review process; if < 0.4, perform new node creation and mark it as to be verified, adding the candidate node as a new entity to the hypergraph network. Automatically creating associated hyperedges (such as "microcracks → abnormal temperature rise → material fatigue") for the newly added or merged nodes, with the initial weight set to a default value (such as 0.5); using a temporal graph convolutional network (T-GCN) to locally adjust the weights of the associated hyperedges, and the formula is: , and are the hidden states of the nodes at both ends of the hyperedge, is a learnable parameter matrix. Only the hyperedge weights directly associated with the newly added / merged nodes are adjusted to avoid global network reconstruction and ensure update efficiency. Each time an undefined defect is detected, the knowledge network is dynamically expanded through the above process to gradually cover new patterns that cannot be handled by traditional methods (such as composite material delamination defects, micron-scale deformations). Based on the online learning mechanism, the node similarity threshold and hyperedge weight parameters are automatically optimized using the feedback of actual quality inspection results (such as misjudgment rate, manual re-inspection confirmation).

[0049] That is, when a new defect pattern is detected, such as an undefined candidate defect description generated by the inference large model, calculate its semantic similarity with the existing nodes , if the similarity exceeds the threshold of 0.6, it is automatically added to the network; at the same time, the hyperedge weights are dynamically adjusted using the Temporal Graph Convolutional Network (T-GCN) , where is the hidden state of the nodes at both ends of the hyperedge, is the LeakyReLU activation function. T-GCN compresses the time-consuming of a single update by only making local weight adjustments to the associated hyperedges (instead of global graph reconstruction), combined with the incremental iteration and lightweight design of the temporal hidden state, so as to increase the dynamic update frequency and improve the real-time response ability of quality inspection.

[0050] For example, in the quality inspection of aero-engine blades, when an undefined defect "microcracks in thermal barrier coating accompanied by abnormal high-frequency vibration" is detected for the first time, the large model analyzes multi-modal data and generates a candidate node description "associated defect of coating microcracks - vibration abnormality". Calculate the maximum similarity with the existing nodes. The similarity with the "coating spalling" node is 0.58 (<0.6), triggering the creation of a new node. A new hyperedge "microcracks → vibration abnormality → abnormal coating process parameters" is added with an initial weight of 0.5. After 3 subsequent detections of the same type of defect, T-GCN dynamically increases the hyperedge weight to 0.89, forming a stable inference path.

[0051] In addition, the embodiments of this specification also provide an adaptive feedback optimization method. Taking the industrial quality inspection scenario as an example, the system input end integrates three types of heterogeneous data: high-definition images of the product surface (such as detecting a scratch length of 1.2 mm), point cloud data from a laser sensor (deformation depth of 0.05 mm), and production batch record text (such as "material batch Y, heat treatment temperature deviation ±10°C"). To adapt to the industrial Internet of Things edge computing environment, online knowledge distillation technology is used to transfer the understanding ability of multi-modal knowledge of the large model (teacher model) to the lightweight model (student model), and a distillation loss function including the domain-common KL divergence method and the loss of the defect classification task is designed . By combining knowledge distillation with the collaborative pruning technology of neurons and knowledge nodes, the dual optimization of model complexity and knowledge base redundancy is achieved. On the premise of ensuring the integrity of the core functions of the system, the resource requirements for edge devices can be significantly reduced, providing reliable support for low-power and high-real-time deployment in the industrial Internet of Things environment.

[0052] In terms of dynamic update and conflict resolution of the knowledge base, the dynamic expansion of the knowledge base is realized through semantic similarity calculation. For example, when a new defect pattern (such as microcracks caused by thermal expansion of ceramic coatings) is detected, the large model generates candidate entity descriptions (such as "microcracks → abnormal coating thermal expansion coefficient → adjust sintering process"), and calculates its semantic similarity with existing nodes . After passing the standard, it is automatically added to the knowledge base and merged into the most similar node; if the similarity is ≥0.4 but <0.6, the manual review process is triggered; if <0.4, a new node is created and marked as to be verified. It should be noted that the "conflict resolution" of the present invention refers to resolving the semantic contradictions between multi-modal data (such as image and laser point cloud conflicts), rather than the entity ambiguity processing (knowledge resolution) within the knowledge base. Conflict resolution includes two core links: conflict detection and conflict correction. The hyperedges of the knowledge base are marked with the HYPER_EDGE label and store the associated node ID list. For example, the hyperedge "scratch - temperature - material" is stored as: CREATE (he:HYPER_EDGE {id: "HE001", nodes: ["Node_Image123", "Node_Sensor456", "Node_Text789"]}). Through the built-in distributed query engine, it supports high-performance multi-hop reasoning (such as tracing image defects back to the production process). Through the dynamic pruning mechanism of neurons and knowledge nodes, the dual optimization of model complexity and knowledge base redundancy is achieved, ensuring the efficient operation of the system on edge devices. First, based on the parameter importance evaluation technology, the importance scores of each neuron in the neural network are calculated, and the 30% neurons with the lowest scores are removed, significantly reducing the computational overhead.

[0053] Subsequently, in the optimization of knowledge base nodes, through the node importance score formula, , the value of each node in the knowledge base is dynamically evaluated, refers to the nodes in the knowledge base, such as entities; is the set of all hyperedges associated with node ; is the weight of hyperedge e; represents the degree of node , that is, the number of connected hyperedges. For example, node can represent "scratch length > 0.5mm" or "material batch Y", hyperedge Connect "Scratch Image", "Temperature Abnormality", and "Material Batch Y", hyperedge Weight = 0.9, indicating a high correlation. The degree of the node "Temperature Abnormality" = 5, indicating 5 connected hyperedges. The importance of a node depends not only on the number of connected hyperedges (degree) but also on the semantic weight of each hyperedge. Nodes that frequently appear in industrial quality inspection scenarios and are associated with critical defects (such as "Temperature Abnormality") are more important. It is updated in real time by T-GCN. When a certain material frequently triggers defects, the weight of its associated hyperedges increases, and the importance of the node increases.

[0054] Remove low-score nodes with an importance score < 0.4, such as the outdated inspection standard "Manual Visual Inspection Threshold 0.8mm", and at the same time merge semantically similar nodes with a semantic similarity > 0.7, such as "Ra0.8" and "Ra1.6" merged into the parent node "Excessive Roughness". This optimization reduces a large amount of computational effort in the knowledge base, reduces edge-side latency, overcomes the problem of the rigidity of traditional graph structures, and ensures the efficient operation of the system in a high-load production line environment.

[0055] By constructing a dynamic knowledge modeling framework, a hybrid reasoning optimization mechanism, cross-modal fusion and conflict resolution methods, and an adaptive feedback optimization technology, the unified representation and efficient reasoning of multi-modal data are achieved, and the compatibility of multi-modal data is realized using heterogeneous feature extraction and unified embedding mapping. The unified representation framework is based on dynamic knowledge graph technology, and through the hypergraph structure as an extended form of the knowledge graph, it supports the expression of topological relationships of multi-modal entities, overcoming the limitation of traditional knowledge graphs that only support binary relationships; following the principles of "dynamic embedding, semantic-driven, and real-time optimization", it focuses on core links such as multi-modal dynamic knowledge modeling → hybrid reasoning path generation → cross-modal fusion and conflict resolution → adaptive feedback optimization, forming a closed-loop system from data collection to knowledge update, which can achieve continuous optimization of data → knowledge → decision → feedback, reducing the knowledge base update latency from hours to seconds, supporting real-time process adjustment of the production line. It should be noted that through the collaborative design of hypergraph theory, dynamic evolution algorithms, and multi-modal fusion mechanisms in the embodiments of this specification, a new generation of knowledge base construction and knowledge reasoning method for industrial quality inspection scenarios is proposed.

[0056] Through the embodiments of this specification, by means of a pre-constructed dynamic hypergraph knowledge network, flexible adjustment can be made according to real-time production data and continuously updated industry rules. When new quality problems or data characteristics appear, the relationships between nodes and hyperedges in the hypergraph can be updated to quickly adapt to changes, effectively solving the problem of lack of support for dynamic knowledge evolution. Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, while industrial quality inspection needs to process high-dimensional heterogeneous data such as images, sensor time-series data, and unstructured text. Through the embodiments of this specification, by obtaining the current inference task information, dynamic weight allocation is performed on the multi-modal input data to determine the multi-modal semantic weight information and the current fused semantic features, realizing the real-time fusion of information between different modal data. Traditional fixed weight allocation strategies ignore the differences in task scenarios. This technical solution performs dynamic weight allocation on the multi-modal input data according to the task description text in the current inference task information, enabling the system to pay more attention to the modal data related to the task. Traditional methods use fixed conflict handling rules such as simple threshold filtering, which are prone to over-rely on unreliable features. Through the embodiments of this specification, semantic conflict detection is performed on the multi-modal input data. When there is a semantic conflict, the weights are adjusted based on the multi-modal semantic weight information to determine the conflict resolution modal weight information. By reducing the weights of conflict modal data or increasing the weights of reliable modal data, a reasonable balance can be achieved between conflict data, making the inference process more inclined to rely on reliable data, effectively resolving semantic conflicts, avoiding missed detection of key defects, and improving the accuracy and reliability of the inference results. By comprehensively applying technologies such as dynamic hypergraph knowledge networks, large models, and conflict resolution, knowledge reasoning can be performed quickly and accurately, and structured decision-making texts can be determined. The local update mechanism of the dynamic hypergraph knowledge network only adjusts the weights of associated hyperedges based on the temporal graph convolutional network (T-GCN), without global reconstruction, and the single update takes a short time, ensuring the real-time performance of the system. Through dynamic weight allocation, semantic conflict detection and resolution, and the powerful semantic understanding and reasoning capabilities of the large model, the accuracy of the inference results is improved. In complex industrial quality inspection scenarios, quality problems of products can be discovered in a timely manner, the types and root causes of defects can be accurately analyzed, and reasonable decision-making suggestions and warning signals can be given, effectively meeting the requirements of industrial quality inspection for real-time performance and accuracy.

[0057] The embodiments of this specification also provide a cross-modal knowledge inference device for industrial quality inspection, as Figure 4 shown. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0058] An embodiment of this specification also provides a non-volatile computer storage medium storing computer-executable instructions, which are configured to execute the above method.

[0059] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0060] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. A cross-modal knowledge reasoning method for industrial quality inspection, characterized in that: The method comprises: Acquire current reasoning task information, dynamically assign weights to multimodal input data according to the task description text in the current reasoning task information, determine multimodal semantic weight information, and determine current fused semantic features corresponding to the multimodal input data; Performing semantic conflict detection on the multimodal input data, and when there is a semantic conflict in the multimodal input data, adjusting the weight of the multimodal semantic weight information based on the multimodal semantic weight information to determine conflict resolution modal weight information; According to the conflict resolution modal weight information and the current fusion semantic features, knowledge reasoning is performed through a large model and a pre-built dynamic hypergraph knowledge network to determine a structured decision text, wherein the structured decision text includes any one or more of defect type, root cause analysis, association path, decision suggestion and warning signal.

2. A cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, characterized in that: According to the task description text in the current reasoning task information, dynamically weight the multimodal input data to determine the multimodal semantic weight information, specifically including: Generate a dynamic task query vector corresponding to the current reasoning task through the task description text; Constructing a cross-modal key-value pair corresponding to each modality in the multimodal input data; According to the dynamic task query vector and the cross-modal key-value pair, dynamic weight allocation is performed on multiple modalities to determine the multimodal semantic weight corresponding to each modality.

3. The cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, characterized in that: Based on the multimodal semantic weight information, weight adjustment is performed on the multimodal semantic weight information to determine the conflict resolution modality weight information, specifically including: Determine at least two conflicting modalities among multiple modalities, and obtain a cross-modal distance indicator between the conflicting modalities; Determining a weight penalty factor corresponding to the conflicting modality by using the cross-modal distance index and a preset distance index threshold; According to the weight penalty factor and the multimodal semantic weight information, the weights of the multiple modalities are adjusted respectively to determine the conflict resolution modal weight information.

4. A cross-modal knowledge reasoning method for industrial quality inspection according to claim 3, characterized in that: According to the weight penalty factor and the multimodal semantic weight information, weights of the multiple modalities are adjusted respectively to determine the conflict resolution modal weight information, specifically including: Performing weight penalties on the conflict modes according to the weight penalty factors, and determining current penalty weight parameters corresponding to each conflict mode; Acquire multimodal semantic weight information corresponding to each of the conflicting modes, and determine the conflict penalty weight according to the multimodal semantic weight information corresponding to the conflicting modes and the current penalty weight parameter; Based on the conflict penalty weight, weight-enhancing a reliable modality among the multiple modalities to determine a conflict enhancement weight corresponding to the reliable modality; The conflict resolution modality weight information is determined by using the current penalty weight parameter corresponding to each of the conflict modalities and the conflict enhancement weight corresponding to the reliable modality.

5. The cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, characterized in that: According to the conflict resolution modal weight information and the current fusion semantic features, knowledge reasoning is performed through a large model and a pre-built dynamic hypergraph knowledge network to determine a structured decision text, specifically including: Determining a candidate reasoning path in the dynamic hypergraph knowledge network according to the current fused semantic features; Generate path description information corresponding to each candidate reasoning path, concatenate the path description information with the predetermined corresponding input problem description information, score each candidate reasoning path through the large model, and determine the path matching score of each candidate reasoning path; The path matching score of each candidate reasoning path is corrected through the conflict resolution modal weight information, and the target reasoning path is determined to determine the structured decision text.

6. A cross-modal knowledge reasoning method for industrial quality inspection according to claim 5, characterized in that: Determining a candidate reasoning path in the dynamic hypergraph knowledge network according to the current fused semantic feature specifically includes: Determine an initial candidate path according to the current fused semantic features and predefined industry rules; Through the path nodes in the initial candidate path, an associated path search is performed in the dynamic hypergraph knowledge network to generate a set of candidate reasoning paths, wherein the set of candidate reasoning paths includes multiple candidate reasoning paths, and each candidate reasoning path includes an RDF triple chain connected by dynamic hyperedges.

7. The cross-modal knowledge reasoning method for industrial quality inspection according to claim 5, characterized in that: The path matching score of each candidate reasoning path is modified by using the conflict resolution modal weight information to determine the target reasoning path, specifically including: Determine at least one reasoning modality in each of the candidate reasoning paths to determine a conflict resolution modality weight corresponding to each of the reasoning modalities; Determining the correction factor of the candidate reasoning path by using the conflict resolution modality weight corresponding to each of the reasoning modalities; Based on the path matching scores of the candidate reasoning paths, using the correction factor, determining the path score corresponding to each of the candidate reasoning paths; Among the multiple candidate reasoning paths, determine the candidate reasoning path with the highest path score as the target reasoning path.

8. The cross-modal knowledge reasoning method for industrial quality inspection according to claim 5, characterized in that: After performing the associated path retrieval in the dynamic hypergraph knowledge network, the method further includes: When the associated path is not retrieved in the dynamic hypergraph knowledge network, the multimodal input data is parsed by a large model to generate candidate defect description information to construct a candidate node; Obtaining an existing hypergraph node set and an existing hyperedge set in the dynamic hypergraph knowledge network, and calculating the maximum semantic similarity between the candidate node and the existing hypergraph node set; According to the maximum semantic similarity, a node update operation is performed in a hypergraph node set of the dynamic hypergraph knowledge network to determine an operation hypergraph node, wherein the node update operation includes a similar node merging operation and a new node creation operation; Determine at least one associated hyperedge corresponding to the operation hypergraph node, use the temporal graph convolutional network to perform local adjustment of the dynamic hyperedge weight, and determine the dynamic hyperedge weight corresponding to the associated hyperedge to dynamically update the dynamic hypergraph knowledge network.

9. A cross-modal knowledge reasoning device for industrial quality inspection, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Power grid fault prediction method based on multi-modal knowledge hybrid reasoning

    CN116910633A

  • Short video rumor detection method and system for complex incomplete data scene

    CN119166854A

  • Multimodal entity relationship extraction method based on hypergraph neural network

    CN119830918A

  • Optimization decision-making method of industrial process fusing domain knowledge and multi-source data

    US11409270B1

  • Cross-modal data processing method and device, storage medium, and electronic device

    WO2022068196A1

Cited By

  • Digital RMB related data processing method and device based on large model technology

    CN120256646A

  • Moving bridge type multi-sensor 3D scanning system

    CN120296689A

  • Mobile bridge multi-sensor 3D scanning system

    CN120296689B

  • Network space surveying and mapping threat detection method and system based on causal association privacy protection

    CN120337301A

  • Cyberspace mapping threat detection method and system with causal association privacy protection

    CN120337301B