A Cross-Modal Knowledge Reasoning Method, Device and Medium for Industrial Quality Inspection
Through dynamic weight allocation and semantic conflict detection, combined with dynamic hypergraph knowledge network and large models, the missed inspection problem caused by fixed weights in industrial quality inspection is solved, real-time fusion and accurate reasoning of multimodal data are realized, and real-time fusion and accurate inference of industrial quality inspection are met.
Patent Information
- Application Number
- CN202510542904.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In industrial quality inspection, the traditional cross-modal knowledge reasoning method is ignoring the difference between different inference task scenarios due to the fixed modal allocation method of the static knowledge base, which leads to the possibility of missing key defects when cross-modal data conflicts, which cannot meet the real-time and accuracy requirements in industrial quality inspection scenarios.
By obtaining the current inference task information, dynamic weight allocation and semantic conflict detection, using dynamic hypergraph knowledge networks and large models for knowledge inference, determining structured decision text, dynamically adjusting modal weights to eliminate conflicts, and real-time fusion and accurate inference of multimodal data.
Real-time integration and accuracy improvement of multimodal data is achieved, product quality problems can be discovered in a timely manner, defect types and root causes can be accurately analyzed, reasonable decision-making suggestions are provided, and real-time and accuracy requirements of industrial quality inspection.
Smart Images

Figure CN120069096B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a cross-modal knowledge reasoning method, device, and medium for industrial quality inspection. Background Art
[0002] In recent years, significant progress has been made in the field of artificial intelligence in the direction of cross-modal knowledge reasoning technology, but there are still many technical bottlenecks in industrial quality inspection. Current research focuses on multi-modal pre-training models and cross-modal retrieval technology, whose core goal is to achieve feature space alignment through contrastive learning. However, it relies on the assumption of static data distribution and lacks support for dynamic knowledge evolution.
[0003] In the industrial quality inspection scenario, industrial quality inspection needs to process high-dimensional heterogeneous data such as images (e.g., surface defects), sensor time-series data (e.g., temperature / pressure fluctuations), and unstructured text (e.g., process records) simultaneously, and the data distribution changes dynamically with the adjustment of production line parameters, equipment aging, etc. Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, and the fixed weight allocation strategy ignores the differences in task scenarios. In addition, factors such as sensor noise and light changes in the industrial environment often lead to contradictions in cross-modal data. For example, the image shows no scratches but the sensor indicates that the stress exceeds the limit. Using fixed conflict handling rules such as simple threshold filtering, and the static knowledge base cannot dynamically adjust the modal weights, it is easy to over-rely on unreliable features, resulting in missed detection of key defects.
[0004] Therefore, the traditional cross-modal knowledge reasoning method is affected by the fixed modal allocation weight method in the static knowledge base, ignores the differences in different reasoning task scenarios, and is prone to the problem of missed detection of key defects when cross-modal data conflicts, and cannot meet the real-time and accurate quality inspection requirements in the industrial quality inspection scenario. Summary of the Invention
[0005] One or more embodiments of this specification provide a cross-modal knowledge reasoning method, device, and medium for industrial quality inspection to solve the following technical problems: The traditional cross-modal knowledge reasoning method is affected by the fixed modal allocation weight method in the static knowledge base, ignores the differences in different reasoning task scenarios, and is prone to the problem of missed detection of key defects when cross-modal data conflicts, and cannot meet the real-time and accurate quality inspection requirements in the industrial quality inspection scenario.
[0006] One or more embodiments of this specification adopt the following technical solutions:
[0007] One or more embodiments of this specification provide a cross-modal knowledge reasoning method for industrial quality inspection. The method includes: obtaining current reasoning task information, dynamically allocating weights to multi-modal input data according to the task description text in the current reasoning task information to determine multi-modal semantic weight information, and determining the current fused semantic features corresponding to the multi-modal input data; performing semantic conflict detection on the multi-modal input data, and when there is a semantic conflict in the multi-modal input data, based on the multi-modal semantic weight information, adjusting the weights of the multi-modal semantic weight information to determine conflict resolution modal weight information; according to the conflict resolution modal weight information and the current fused semantic features, performing knowledge reasoning through a large model and a pre-constructed dynamic hypergraph knowledge network to determine structured decision-making text, where the structured decision-making text includes any one or more of defect type, root cause analysis, association path, decision suggestion, and warning signal.
[0008] One or more embodiments of this specification provide a cross-modal knowledge reasoning device for industrial quality inspection, including:
[0009] At least one processor; and,
[0010] A memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method.
[0012] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.
[0013] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Through the embodiments of this specification, by means of a pre-constructed dynamic hypergraph knowledge network, flexible adjustment can be made according to real-time production data and continuously updated industry rules. When new quality problems or data characteristics appear, the relationship between nodes and hyperedges in the hypergraph can be updated to quickly adapt to changes, effectively solving the problem of lack of support for dynamic knowledge evolution. Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, while industrial quality inspection needs to process high-dimensional heterogeneous data such as images, sensor time-series data, and unstructured text. The embodiments of this specification obtain the current inference task information, perform dynamic weight allocation on multi-modal input data, determine the multi-modal semantic weight information, and determine the current fused semantic features, realizing the real-time fusion of information between different modal data. Traditional fixed weight allocation strategies ignore the differences in task scenarios. This technical solution performs dynamic weight allocation on multi-modal input data according to the task description text in the current inference task information, enabling the system to pay more attention to modal data related to the task. Traditional methods adopt fixed conflict handling rules such as simple threshold filtering, which are prone to over-rely on unreliable features. The embodiments of this specification perform semantic conflict detection on multi-modal input data. When there is a semantic conflict, the weights are adjusted based on the multi-modal semantic weight information to determine the conflict resolution modal weight information. By reducing the weight of conflict modal data or increasing the weight of reliable modal data, a reasonable trade-off can be made between conflict data, so that the inference process is more inclined to rely on reliable data, effectively resolving semantic conflicts, avoiding missed detection of key defects, and improving the accuracy and reliability of the inference results. By comprehensively applying technologies such as dynamic hypergraph knowledge networks, large models, and conflict resolution, knowledge reasoning can be carried out quickly and accurately, and structured decision-making texts can be determined. The local update mechanism of the dynamic hypergraph knowledge network only adjusts the weights of associated hyperedges based on the temporal graph convolutional network (T-GCN), without global reconstruction, and the single update takes a short time, ensuring the real-time performance of the system. Through dynamic weight allocation, semantic conflict detection and resolution, and the powerful semantic understanding and reasoning ability of the large model, the accuracy of the inference results is improved. In complex industrial quality inspection scenarios, quality problems of products can be detected in a timely manner, the types and root causes of defects can be accurately analyzed, and reasonable decision-making suggestions and warning signals can be given, effectively meeting the requirements of industrial quality inspection for real-time performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0015] Figure 1 It is a schematic flowchart of a cross-modal knowledge reasoning method for industrial quality inspection provided by an embodiment of this specification;
[0016] Figure 2 It is a schematic flowchart of generating a hybrid reasoning path provided by an embodiment of this specification;
[0017] Figure 3 It is a schematic flowchart of multi-modal dynamic knowledge modeling provided by an embodiment of this specification;
[0018] Figure 4 It is a schematic structural diagram of a cross-modal knowledge reasoning device for industrial quality inspection provided by an embodiment of this specification. Detailed implementation manners
[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0020] An embodiment of this specification provides a cross-modal knowledge reasoning method for industrial quality inspection. It should be noted that the execution subject in the embodiment of this specification can be a server or any device with data processing capabilities. Figure 1 It is a schematic flowchart of a cross-modal knowledge reasoning method for industrial quality inspection provided by an embodiment of this specification, as Figure 1 shown, mainly including the following steps:
[0021] Step S101, obtain the current reasoning task information, perform dynamic weight assignment on the multi-modal input data according to the task description text in the current reasoning task information, determine the multi-modal semantic weight information, and determine the current fused semantic feature corresponding to the multi-modal input data.
[0022] In an embodiment of this specification, obtain the current reasoning task information. Here, the current reasoning task information includes the task description text corresponding to this reasoning task and the multi-modal input data. The task description text is used to distinguish the task type corresponding to the current reasoning task, such as a defect classification task, a root cause analysis task, etc. The multi-modal input data can be various modalities such as image data, sensor data, and text data.
[0023] According to the task description text in the current inference task information, perform dynamic weight allocation on the multi-modal input data to determine the multi-modal semantic weight information, specifically including: generating a dynamic task query vector corresponding to the current inference task through the task description text; constructing cross-modal key-value pairs for each modality in the multi-modal input data; and performing dynamic weight allocation on multiple modalities according to the dynamic task query vector and the cross-modal key-value pairs to determine the multi-modal semantic weight corresponding to each modality.
[0024] In one embodiment of the present specification, a dynamic query vector is generated according to the current inference task type, such as "defect classification task", "root cause analysis task". The semantic features of the task description are extracted through a pre-trained semantic encoder and mapped to a 128-dimensional dynamic task query vector, which represents the core requirements of the current task (for example, "root cause analysis" needs to associate material properties with process parameters). The task description text is mapped to a query vector through a task encoder (Task Encoder): , where is the semantic embedding extracted from the task description by the large model, is a learnable parameter matrix, is a bias term. For example, in the root cause analysis task, to determine the cause of microcracks, the encoder encodes the task text and outputs a semantic vector, which is vectorized into , and then after mapping and adding the bias term, a query vector is generated, pointing to semantics such as "material properties" and "process parameters".
[0025] Extract the scratch area features through a convolutional neural network to generate a key vector and a value vector . The sensor modality processes the temperature and pressure fluctuation data using a temporal encoder (LSTM) to generate a key vector k sensor and a value vector v sensor . Parse the process record text to extract keywords (such as "material batch X has insufficient hardness") to generate a text modality key vector k text and a value vector v text . Taking the process of constructing key-value pairs for image, sensor, and text modality features as an example for illustration. In the image modality, , in the sensor modality, , and in the text modality , where is the feature vector of each modality, and ∈ ℝ^{d×d} is the modality projection matrix. The task query vector q task is respectively matched with the key vectors of each modality, and the initial attention scores are calculated through scaled dot product to obtain the weights of each modality , where \(i\in\{image, sensor, text\}\), and \(d = 64\) is the dimensional scaling factor. The final weighted fusion feature is . For example, in the surface scratch detection task, the system dynamically assigns weights: the image modality = 0.7 (relying on visual features); the sensor modality = 0.2 (monitoring temperature anomalies); the text history record modality = 0.1 (referring to historical quality inspection records).
[0026] Through the above technical solution, based on the generation of dynamic task query vectors, it is possible to automatically adjust the importance weights of multi-modal data according to the current task type, such as "defect classification", "root cause analysis", "process optimization". For example, in the root cause analysis task, the weight of sensor time-series data is increased to 0.65 (default 0.3), and the weight of text process records is increased to 0.25 (default 0.1), effectively capturing deep features strongly related to material properties and process parameters, and increasing the defect root cause localization accuracy to 98.5%; through cross-modal key-value pair construction and dynamic weight allocation, the semantic consistency between different modalities can be quantified.
[0027] Step S102, perform semantic conflict detection on the multi-modal input data. When there is a semantic conflict in the multi-modal input data, based on the multi-modal semantic weight information, adjust the weight of the multi-modal semantic weight information to determine the conflict resolution modal weight information.
[0028] In an embodiment of the present specification, for possible semantic conflicts, such as the image shows "no scratches" while the sensor data indicates "pressure overlimit" or the image shows cracks but the laser point cloud has no deformation features, based on the semantic distance threshold for detection. When the threshold is greater than 1.5, it is determined that a semantic conflict has occurred. When a conflict is detected, it will trigger the backpropagation optimization of the feature alignment module, recalibrate the modal representation, and combine the associated paths in the semantic network to generate a correction suggestion, which can reduce the quality inspection false alarm rate and significantly improve the consistency of multi-modal data.
[0029] In an embodiment of the present specification, the conflict detection trigger conditions are as follows: calculate the cross-modal feature distance and determine the semantic consistency , where is an indicator function that outputs 1 when the condition is met, otherwise 0, is the embedding vector of image, sensor or text features, and \(\tau = 1.5\) is the preset distance threshold determined by expert experience data. When there is any > \(\tau\), it is determined that a semantic conflict has occurred.
[0030] Based on the multi-modal semantic weight information, the multi-modal semantic weight information is adjusted to determine the conflict resolution modal weight information, specifically including: determining at least two conflicting modalities among multiple modalities, and obtaining the cross-modal distance index between the conflicting modalities; determining the weight penalty factor corresponding to the conflicting modalities through the cross-modal distance index and a preset distance index threshold; and adjusting the weights of the multiple modalities respectively according to the weight penalty factor and the multi-modal semantic weight information to determine the conflict resolution modal weight information.
[0031] When there is a semantic conflict in the multi-modal input data, two conflicting modalities with conflicts are determined among the multiple modalities, and the cross-modal distance index between these two conflicting modalities obtained during the conflict judgment process is obtained, that is . Through the cross-modal distance index and a preset distance index threshold τ, the weight penalty factor corresponding to the conflicting modalities is determined, and the penalty factor is , where = 2.0, = 1.5. The weights of the multiple modalities are adjusted respectively using the weight penalty factor and the multi-modal semantic weight information of the corresponding modalities obtained in step S101 to determine the conflict resolution modal weight information.
[0032] According to the weight penalty factor and the multi-modal semantic weight information, the weights of the multiple modalities are adjusted respectively to determine the conflict resolution modal weight information, specifically including: imposing a weight penalty on the conflicting modalities according to the weight penalty factor to determine the current penalty weight parameter corresponding to each of the conflicting modalities; obtaining the multi-modal semantic weight information corresponding to each of the conflicting modalities, and determining the conflict penalty weight through the multi-modal semantic weight information corresponding to the conflicting modality and the current penalty weight parameter; based on the conflict penalty weight, boosting the weights of the reliable modalities among the multiple modalities to determine the conflict boosting weight corresponding to the reliable modality; and determining the conflict resolution modal weight information through the current penalty weight parameter corresponding to each of the conflicting modalities and the conflict boosting weight corresponding to the reliable modality.
[0033] In one embodiment of this specification, according to the weight penalty factor, weight penalty is imposed on the conflicting modality, and the current penalty weight parameter corresponding to the conflicting modality is determined by the product of the weight penalty factor and the original weight in the multi-modal semantic weight information corresponding to this conflicting modality. The conflicting penalty weight corresponding to each conflicting modality is determined by the difference between the multi-modal semantic weight information corresponding to each conflicting modality and the corresponding current penalty weight parameter. Based on the conflicting penalty weight, the weight of the reliable modality among multiple modalities is enhanced, and the conflicting enhancement weight corresponding to the reliable modality is determined. Here, the enhancement amount of the conflicting enhancement weight should be the same as the penalty amount of the conflicting penalty weight. In addition, it should be noted that after summarizing the conflicting penalty weights of each conflicting modality for which weight conflict penalty is imposed, the cumulative weight penalty amount corresponding to the conflicting modality is determined. The conflicting resolution modality weight information is determined by the current penalty weight parameter corresponding to each such conflicting modality and the conflicting enhancement weight corresponding to the reliable modality.
[0034] Assume that the initial weight of the image modality is 0.7, the sensor modality is 0.2, and the text modality is 0.1. It is detected that the L2 distance between the image and the sensor is 1.8, and the distance between the image modality and the sensor modality is greater than 1.5, indicating that there is a semantic conflict. It is necessary to adjust the image weight and the sensor weight. The image weight is adjusted by the formula , that is, adjusted to 0.38, which is the current penalty weight parameter corresponding to the image modality. The sensor weight is also adjusted to 0.11 by the above formula. The original weight of the image modality is 0.7, and after being adjusted to 0.38, it is reduced by 0.32, that is, the conflicting penalty weight under the image modality is 0.32. The original weight of the sensor modality is 0.2, and after being adjusted to 0.11, it is reduced by 0.09, that is, the conflicting penalty weight under the sensor modality is 0.09. The corresponding cumulative conflicting penalty weight of the two is 0.32 + 0.09 = 0.41, that is, the cumulative weight penalty amount is 0.41. Since there is no conflict in the text modality, it is necessary to enhance the decision proportion of the reliable modality (text). The weight enhancement amount should be the same as the cumulative weight penalty amount. Therefore, the text modality weight is 0.1 + 0.41 = 0.51, and the conflicting enhancement weight is obtained. According to the above steps, the current penalty weight parameter corresponding to the conflicting modality and the conflicting enhancement weight corresponding to the reliable modality are obtained, and the conflicting resolution modality weight information corresponding to each modality is determined.
[0035] Through the above technical solution, semantic conflicts are detected based on the semantic distance threshold. When a conflict is detected, the feature alignment module is triggered to perform backpropagation optimization, recalibrate the modal representation, and generate correction suggestions by combining the association paths in the semantic network, which can effectively avoid misjudgments caused by semantic conflicts, thereby reducing the false alarm rate of quality inspection and improving the accuracy of quality inspection. After detecting a semantic conflict, a series of operations are performed to adjust the multi-modal semantic weight information, including determining the conflicting modality, obtaining the cross-modal distance metric, calculating the weight penalty factor, performing weight penalty and promotion, etc., which significantly improves the semantic consistency between multi-modal data and ensures the reliability and accuracy of the data. According to the situation of semantic conflicts, weight penalty is imposed on the conflicting modalities, while weight promotion is performed on the reliable modalities. This dynamic weight adjustment method can reasonably allocate the proportion of each modality in the decision-making, enhance the role of reliable modalities, and reduce the influence of conflicting modalities, thus making the decision-making of the entire multi-modal system more reasonable and reliable.
[0036] Step S103, according to the conflict resolution modal weight information and the current fused semantic features, perform knowledge reasoning through a large model and a pre-constructed dynamic hypergraph knowledge network to determine the structured decision text.
[0037] Among them, the structured decision text includes any one or more of defect type, root cause analysis, association path, decision suggestion, and warning signal.
[0038] In an embodiment of the present specification, a hybrid inference path generation is performed according to the multi-modal input data in the absence of semantic conflicts. Figure 2 It is a schematic flowchart of a hybrid inference path generation provided by an embodiment of the present specification. As Figure 2 shown, it fuses symbolic logic and neural networks to solve the problems of path redundancy and low efficiency of traditional inference methods in complex multi-modal scenarios. Taking the industrial quality inspection scenario as an example, the input data includes a high-definition image of the product surface (such as the detected scratch length is 1.2 mm), laser sensor point cloud data (the deformation depth is 0.05 mm), and production record text (such as "material batch Y, heat treatment temperature deviation ±10°C"). The symbolic reasoning layer constructs a production rule library based on industry detection standards. For example: "scratch length > 0.5 mm → insufficient material hardness → re-inspection required", which is represented symbolically as: "IF (scratch length > 0.5 mm) ∧ (material hardness < 50 HRC) THEN trigger the re-inspection of the heat treatment process", thereby generating candidate inference paths.
[0039] Retrieve associated paths from a multi-dimensional semantic network through querying with the Semantic Web Query Language (SPARQL Protocol and RDF Query Language, SPARQL). For example, when inputting the defect feature "scratches exceeding the standard", retrieve the path "excessive scratch length → abnormal material hardness → trace the heat treatment process → adjust the temperature parameter", and encode it into a natural language description (such as "Detected a scratch length of 1.2 mm, and the material hardness is lower than the standard value. It is recommended to check the temperature of the heat treatment furnace"). For instance, input the RDF triples of the candidate path into a large model and generate a natural language description through template filling. Template example: Detected [defect feature], [root cause analysis], recommended [decision suggestion]. RDF triples are only a temporary expression form of the inference result, generated through dynamic hyper-edge weights and the context understanding of the large model, which is essentially different from the static triple storage of traditional knowledge graphs. The neural inference layer uses the large model to perform semantic scoring on the candidate paths. Concatenate the path description with the input question (such as "Analysis of the cause of scratches?") and input it into the model, and calculate the matching degree through a scoring function , where is the input question embedding vector, such as "Possible causes of a scratch length of 1.2 mm", is the candidate path description, is the semantic matching layer of the large model. Screen the Top-3 high-confidence paths, such as s1 = 0.95, s2 = 0.88, s3 = 0.82, to ensure that the results have both accuracy and interpretability
[0040] To further optimize the inference efficiency, introduce a reinforcement learning mechanism to dynamically adjust the strategy. Design a multi-objective reward function , where the detection accuracy is verified through historical quality inspection results, and the response time is in seconds, which needs to meet the production line beat requirements. Update the policy network parameters through the Proximal Policy Optimization (PPO) algorithm. The policy network is an independently trained reinforcement learning model, and its training goal is to learn how to dynamically select the optimal inference path (such as preferentially performing symbolic inference or neural inference) according to multi-modal input features (such as image scratch length, sensor temperature data, etc.) by interacting with the knowledge base environment. During the training process, the policy network processes state-action pairs (such as the state is a combination of defect features, and the action is to select the path of "re-check process parameters" or "trigger laser repair"), dynamically balancing the detection accuracy and real-time requirements. This hybrid mechanism reduces the path search space by 60%. In the high-speed production line scenario, the response time for complex defect determination is reduced from 5 seconds to 1.5 seconds, and the detection accuracy is increased to 98.5% (85% for traditional methods). To further understand the role of knowledge inference, combined with industrial scenarios, an example of the output of knowledge inference is illustrated as follows. When the system detects a micro-crack (length 0.3 mm) on the surface of an aero-engine blade, the inference engine generates the following results
[0041] (1)Defect classification: Microcracks (confidence 0.95);
[0042] (2)Root cause analysis: Abnormal coating thermal expansion coefficient (probability 89%);
[0043] (3)Decision suggestion: Adjust the sintering process temperature to 1250 °C and perform X-ray flaw detection verification;
[0044] (4)Warning signal: Trigger the automatic shutdown of the production line to avoid batch quality accidents.
[0045] It should be noted that before performing hybrid reasoning, it is necessary to pre-build a multi-modal dynamic knowledge model and construct a dynamically evolving multi-modal knowledge representation framework to solve the static problem of traditional knowledge graphs. Existing knowledge graphs rely on binary relations, head entity - relation - tail entity, and it is difficult to express the joint semantics of multi-modal data. For example, "image scratch + sensor data + text record" needs to be split into multiple triples. In the embodiments of this specification, the knowledge representation framework is implemented using a hypergraph structure. Through the multi-nodes of the hypergraph, it supports complex cross-modal reasoning and improves the path retrieval efficiency.
[0046] Figure 3 FIG. is a schematic flow chart of a multi-modal dynamic knowledge modeling provided by an embodiment of this specification. As Figure 3 shown, in the industrial quality inspection scenario, the input data includes product surface images, time-series data of sensors such as temperature and pressure, and quality inspection report texts. First, through a multi-modal pre-trained large model, here a domestic open-source multi-modal large model can be used, cross-modal feature extraction is realized. For the image modality, the commonly used Vision Transformer (ViT) model in the field can be used to extract local texture features in blocks (such as scratches and depressions), and a 224×224 pixel image is divided into 16×16 blocks. Each block extracts a d = 768-dimensional feature vector through 12 layers of Transformer encoders; the sensor data encodes the time-series features through an LSTM network, with the time step set to t = 50 and the hidden layer dimension h = 256; the text modality uses the BERT model to parse the keywords in the quality inspection report (such as "qualified", "defect type A") and outputs a 768-dimensional semantic vector of the [CLS] token. The multi-modal features interact in the cross-modal attention layer of UNITER to generate a unified semantic representation. At the same time, to achieve cross-modal semantic alignment, a two-tower contrastive learning architecture is designed and trained through the Adam optimizer, using the contrast loss function , to align the image, sensor, and text features in a unified embedding space, where B is the number of samples in the current batch, is the image feature, is the corresponding text description feature, =0.05 is the temperature coefficient, which is determined by grid search. The dual-tower structure includes an image-sensor tower and a text tower, and the semantic alignment between modalities is constrained by contrastive loss. Among them, the positive sample pairs are multi-modal data of the same quality inspection sample (such as image scratch + corresponding sensor temperature anomaly + text report), and the negative sample pairs are randomly sampled different sample combinations.
[0047] Based on feature alignment, a multi-dimensional semantic network is constructed. This network uses a hypergraph structure to dynamically represent cross-modal relationships, serving as the underlying data structure of the multi-dimensional semantic network. A hypergraph allows a hyperedge to connect any number of nodes, breaking through the limitation of traditional knowledge graphs that can only express knowledge through binary relationships (i.e., head entity - relationship - tail entity). For example, in the industrial quality inspection scenario, a hyperedge can connect an image feature node (such as "scratch length 1.2mm"), a sensor node (such as "temperature anomaly 85°C"), and a text node (such as "material batch Y") simultaneously to completely represent the cross-modal association of "scratch - temperature - material". The hypergraph node set includes entities (such as "defect type A", "temperature threshold"), concepts (such as "detection standard"), and modal features (such as "image texture feature"), and the hyperedge set represents cross-modal associations (such as "image scratch → sensor temperature anomaly → quality inspection report conclusion"). In addition, when a new defect pattern is detected, such as an undefined candidate defect description generated by an inference large model, calculate its semantic similarity with existing nodes , calculate the similarity using the cosine function, represents the newly added candidate defect node, u is the set of existing knowledge base nodes, is the candidate node 's feature vector, is the feature vector of the existing node. If the similarity exceeds the threshold of 0.6, it is automatically added to the network; at the same time, a temporal graph convolutional network (T-GCN) is used to dynamically adjust the hyperedge weights , where and are the hidden states of the nodes at both ends of the hyperedge, is the LeakyReLU activation function, is a learnable weight matrix used to map the hidden state of the node to the hyperedge weight space; || is the vector concatenation operation to fuse the semantic information of the two end nodes. Through project the concatenated high-dimensional vector into the weight space; Represents the weight value of the hyperedge e at time step t+1, which is used for subsequent inference path retrieval and conflict resolution, and quantifies the importance of this hyperedge in multimodal semantic association. T-GCN compresses the time consumption of a single update by only making local weight adjustments to the associated hyperedges (instead of global graph reconstruction), combined with the incremental iteration and lightweight design of the temporal hidden state, thereby increasing the dynamic update frequency and enhancing the real-time response ability of quality inspection.
[0048] Through the above technical solutions, the hypergraph structure is adopted to break through the binary relationship limitation of traditional knowledge graphs, support the joint semantic expression of multimodal data such as images, sensors, and texts, making the knowledge representation more suitable for complex industrial scenarios. Cross-modal semantic alignment is achieved through the contrastive learning two-tower structure and the contrastive loss function to ensure the effective interaction of different modal features in a unified space, solving the problem of modal fragmentation in traditional methods; the combination of symbolic logic and neural networks not only retains the interpretability of rule reasoning, such as "scratch length > 0.5mm → material hardness insufficient", but also uses the generalization ability of neural reasoning to handle unknown defect patterns. Reinforcement learning dynamically selects the optimal inference path, reducing the response time of complex defect determination and the path search space, meeting the real-time requirements of high-speed production lines; new defect patterns are automatically added through semantic similarity matching, and the local weight of hyperedges is adjusted in combination with the temporal graph convolutional network, avoiding global graph reconstruction, and compressing the update time consumption to the millisecond level. In the industrial quality inspection scenario, the detection accuracy is improved, the response time is shortened, and the production line beat requirements are met; the reward function combines the detection accuracy and the response time, and balances the contradiction between the two through reinforcement learning; through the multi-node connection characteristics of the hypergraph structure and the evolution mechanism of the temporal graph convolutional network, the limitation of the binary relationship expression of traditional knowledge graphs is solved, and joint semantic modeling of multimodal data such as images, point clouds, and texts is realized, and real-time dynamic update of the knowledge base is supported; the adaptive ability to new defect patterns is significantly improved, and the problems of missed detection caused by lagging knowledge update and cross-modal association breakage of complex defects in industrial scenarios can be effectively solved.
[0049] It should be noted that the conflict resolution mechanism in the embodiments of this specification is essentially different from knowledge resolution. Conflict resolution refers to the semantic contradictions between multi-modal data during the real-time reasoning process, such as the conflict between images and sensor data. Through feature calibration and rule reasoning, correction suggestions are generated. Knowledge resolution belongs to the entity ambiguity processing in the knowledge base construction stage, such as merging synonymous entities. The embodiments of this specification achieve entity ambiguity processing in the knowledge base construction stage through dynamic pruning and merging of hypergraph nodes. Taking the industrial quality inspection scenario as an example, it is necessary to fuse image defect features, sensor abnormal data, and text report conclusions. The present invention designs a multi-head attention mechanism to dynamically allocate modal weights, generate query vectors according to the task type, calculate the attention scores of the modal key vectors, and generate joint features after weighted fusion. For possible semantic conflicts, such as the image showing "no scratch" while the sensor data indicates "pressure overlimit", or the image showing a crack but the laser point cloud having no deformation feature, it is detected based on the semantic distance threshold. When the threshold is greater than 1.5, it is determined that a semantic conflict has occurred. When a conflict is detected, this method will trigger the feature alignment module to perform backpropagation optimization, recalibrate the modal representation, and combine the associated paths in the semantic network to generate correction suggestions. This mechanism can reduce the false alarm rate of quality inspection and significantly improve the consistency of multi-modal data.
[0050] The conflict resolution algorithm adopts a two-stage correction strategy. In the feature calibration stage, the contrast loss function of the multi-modal dynamic knowledge modeling link is used. In the reasoning correction stage, correction suggestions are generated based on the knowledge base rules and the large language model. In the calibration mechanism of the contrast loss function in the feature calibration stage, the feature vectors of the conflicting modalities (such as images and sensors) are aligned to a unified semantic space, reducing the cross-modal feature distance, so that the cross-modal data of the same object (such as a scratch image and the corresponding temperature sensor data) are close in the embedding space, and the features of different objects are far away. The multi-modal data is mapped to a unified embedding space through a two-tower structure (image-sensor tower and text tower). The cosine similarity of the positive sample pairs is calculated and normalized by Softmax. The contrast loss is minimized to push the positive sample pairs closer and the negative sample pairs farther away. In the reasoning correction stage, knowledge-driven semantic correction retrieves the reasoning path associated with the calibrated features from the dynamic hypergraph knowledge base (such as "pressure overlimit → material fatigue → internal damage"). The path is encoded as a natural language instruction (such as "It is recommended to perform X-ray flaw detection"). The correction suggestions in the reasoning correction stage are based on the features calibrated in the feature calibration stage (such as the aligned sensor features). If not calibrated in the feature calibration stage, the sensor features may be associated with the "temperature anomaly" node, generating incorrect suggestions; after calibration, the sensor features are associated with the "internal damage" node, generating correct suggestions. The execution result of the correction suggestion (such as X-ray confirming internal damage) can be used as a supervision signal to adjust the parameters of the contrast loss function (such as the temperature coefficient τ) through reinforcement learning.
[0051] An example of the process of retrieving the inference path associated with the calibrated features from the dynamic hypergraph knowledge base is as follows. First, perform the associated path retrieval to retrieve the top-3 high-weight paths from the multi-dimensional semantic network by extending the SPARQL query. Then, perform the large model semantic generation. Input the path information into the large model and generate natural language suggestions through the structured prompt template. Finally, perform the post-processing of the execution results: utilize the zero-shot classification ability of the large model to perform keyword verification and confidence annotation on the generated suggestions. For example, it can be verified whether the necessary keywords (such as "retest", "ISO standard") are included in the suggestions to ensure compliance with industry standards and operating procedures. The conflict resolution in the inference correction stage will be described in detail below.
[0052] Based on the conflict resolution modality weight information and the current fused semantic features, perform knowledge reasoning through the large model and the pre-constructed dynamic hypergraph knowledge network to determine the structured decision text, specifically including: determining candidate inference paths in the dynamic hypergraph knowledge network according to the current fused semantic features; generating path description information corresponding to each candidate inference path, splicing the path description information with the pre-determined corresponding input problem description information, and scoring each candidate inference path through the large model to determine the path matching degree score of each candidate inference path; correcting the path matching degree score of each candidate inference path through the conflict resolution modality weight information to determine the target inference path so as to determine the structured decision text.
[0053] In one embodiment of the present specification, based on the currently fused multi-modal semantic features, possible inference paths are retrieved in the dynamic hypergraph knowledge network for image defect regions, sensor abnormal data, and text process records. According to predefined industry rules, such as "exceeding the temperature standard needs to be associated with material properties", a basic candidate path is generated: "temperature anomaly → material fatigue → recheck process parameters". Through the dynamic association relationships of the hypergraph network, such as multi-hop nodes connected by hyperedges, the path range is automatically expanded. For example, if a surface scratch is detected, associated nodes such as "insufficient coating thickness" or "excessive assembly stress" may be retrieved, forming multiple candidate paths. Each candidate path is converted into a natural language description and semantically associated with the user's input question. The node and hyperedge relationships in the path are converted into readable text. For example, the path "temperature anomaly → material fatigue → recheck process parameters" is described as "Temperature exceeding the standard is detected, associating the possibility of material fatigue, and it is recommended to recheck the heat treatment parameters". The user's question (such as "What is the root cause of the current defect?") is concatenated with the path description to form a complete input. For example: "Analysis of the root cause of the defect: Temperature exceeding the standard is detected, associating the possibility of material fatigue, and it is recommended to recheck the heat treatment parameters". The concatenated text is input into the pre-trained large model, and the matching degree between the path and the question is evaluated through the semantic understanding ability inside the model. The model analyzes the logical relevance between the path description and the question and outputs a confidence score (0-1). For example, the path "temperature anomaly → material fatigue" scores 0.92 in the root cause analysis task, but may be only 0.45 in the surface detection task. The top 3 paths with the highest scores are retained to ensure the results are both accurate and diverse. Combining the weights after multi-modal conflict resolution (such as suppressing the image weight when there is a contradiction between the image and sensor data), the path scores are dynamically adjusted. If there is a conflict in the modal data on which a certain path depends (such as the sensor shows normal temperature but the text record is abnormal), the score of that path is reduced. The scores of paths depending on high-weight modalities are increased to ensure that reliable data sources are preferentially adopted in the final decision. When determining this structured decision text, the paths before and after conflict resolution can be displayed simultaneously.
[0054] Through the above technical solutions, the hypergraph structure supports multi-hop node expansion, breaks through the limitations of traditional binary relations, covers more complex causal chains, and jointly generates candidate paths with predefined rules and hypergraph dynamic relations to ensure the balance between rule authority and data-driven flexibility; by splicing the path description with the user's question, the large model can specifically evaluate the path relevance, retain high-confidence paths, avoid the one-sidedness of a single path, and ensure the decision-making robustness through diversity; when multi-modal data conflicts, reduce the scores of conflicting modal paths; boost the scores of paths that rely on high-confidence modalities to ensure that reliable data sources are preferentially adopted in decision-making; weight adjustment is synchronized with path scoring, eliminating the need for offline retraining and meeting the real-time requirements of the production line. In scenarios where multi-modal data is noisy or conflicting, the conflict resolution mechanism can improve the decision-making accuracy; by integrating the deterministic reasoning of the symbolic rule engine, the semantic understanding ability of the large language model, and the dynamic decision-making mechanism of reinforcement learning, an interpretable and highly accurate hybrid reasoning path generation strategy is formed. While ensuring detection accuracy, it significantly reduces redundant computational overhead and meets the stringent requirements of high-speed production lines for real-time response.
[0055] According to the current fused semantic features, candidate inference paths are determined in the dynamic hypergraph knowledge network, specifically including: determining initial candidate paths according to the current fused semantic features and predefined industry rules; through the path nodes in the initial candidate paths, performing associated path retrieval in the dynamic hypergraph knowledge network to generate a set of candidate inference paths, where the set of candidate inference paths includes multiple candidate inference paths, and each candidate inference path includes an RDF triple chain connected by a dynamic hyperedge.
[0056] In one embodiment of the present specification, an initial candidate path is generated by combining the currently fused semantic features (such as detecting that the length of the surface scratch exceeds the standard) with predefined industry rules (such as "if the scratch length > 0.5 mm, a process recheck needs to be triggered"). Example rules are as follows: "IF the scratch length > 0.5 mm THEN associate with material hardness detection". An initial path is generated according to the rules. For example: "Surface scratch exceeds the standard → Insufficient material hardness → Recheck heat treatment process". If abnormal temperature rise (sensor data) is currently detected, the path is extended to: "Surface scratch exceeds the standard → Abnormal temperature rise → Insufficient material hardness". Based on the nodes in the initial path (such as "Surface scratch exceeds the standard", "Insufficient material hardness"), associated paths are retrieved in the dynamic hypergraph knowledge network to form a set of candidate inference paths. Expand layer by layer along the hyperedge connection relationship. The first hop is "Surface scratch exceeds the standard → Insufficient material hardness", the second hop is "Insufficient material hardness → Abnormal heat treatment furnace temperature", and the third hop is "Abnormal heat treatment furnace temperature → Calibrate temperature control parameters". Finally, a path chain "Surface scratch exceeds the standard → Insufficient material hardness → Abnormal heat treatment furnace temperature → Calibrate temperature control parameters" is generated. Only the associated paths with weights exceeding the threshold (such as 0.6) are retained to ensure the reliability of the paths. For example, the weak association path "Insufficient material hardness → Excessive assembly stress" with a weight of 0.3 is eliminated. Each candidate path consists of a chain of RDF triples connected by dynamic hyperedges, representing a complete causal relationship chain. An example of an RDF triple chain is as follows: triple 1 (Surface scratch exceeds the standard, associated defect, Insufficient material hardness), triple 2 (Insufficient material hardness, root cause analysis, Abnormal heat treatment furnace temperature), triple: (Abnormal heat treatment furnace temperature, solution, Calibrate temperature control parameters). The RDF triple chain is converted into a natural language description. For example: "It is detected that the surface scratch exceeds the standard, associated with the problem of insufficient material hardness. The root cause analysis is abnormal heat treatment furnace temperature, and it is recommended to calibrate the temperature control parameters."
[0057] Through the above technical solution, by leveraging pre-defined industry rules, it is possible to make full use of professional knowledge and standards within the industry, ensuring that in the face of specific quality problems, an initial judgment can be made based on established industry best practices, generating reasonable initial candidate paths, effectively avoiding deviation from industry standards during the reasoning process, and improving the accuracy and professionalism of the reasoning results. At the same time, by combining the currently integrated semantic features, the initial path is dynamically expanded. This combination of real-time data and industry rules enables the reasoning path to closely fit various situations in the actual production scenario, more comprehensively reflecting the potential associations of product quality problems; using the association relationships between nodes in the dynamic hypergraph knowledge network for path retrieval can fully explore the potential connections between data. By starting from the nodes in the initial path and expanding layer by layer along the hyperedge connection relationships, a long-chain reasoning path is formed, greatly expanding the depth and breadth of reasoning. Compared with traditional simple association reasoning, it can discover causal relationships at more levels and angles, thus providing more comprehensive clues for solving complex quality problems. By setting a weight threshold to screen the associated paths, it is ensured that the retained paths have high reliability, avoiding interference from invalid or low-value paths to the reasoning results and further improving the quality of the reasoning path set; each candidate path consists of a chain of RDF triples connected by dynamic hyperedges. This structured representation clearly defines the relationships between various nodes, and the RDF triple chain can accurately describe the complete causal relationship chain.
[0058] Through the conflict resolution modal weight information, the path matching degree score of each candidate reasoning path is corrected to determine the target reasoning path, specifically including: determining at least one reasoning modality in each candidate reasoning path to determine the conflict resolution modal weight corresponding to each reasoning modality; determining the correction factor of the candidate reasoning path through the conflict resolution modal weight corresponding to each reasoning modality; based on the path matching degree score of the candidate reasoning path, using the correction factor to determine the path score corresponding to each candidate reasoning path; among multiple candidate reasoning paths, determining the candidate reasoning path with the highest path score as the target reasoning path.
[0059] In one embodiment of this specification, for each candidate inference path, identify the key modal data it depends on (such as images, sensors, text), and obtain the modal weight information after conflict resolution. For example, path A depends on an image (the conflict resolution weight corresponding to the image modality is 0.38) and sensor data (weight 0.11), and path B depends on a sensor (the conflict resolution weight corresponding to the sensor modality is 0.11) and a text record (the conflict resolution weight corresponding to the text modality is 0.51). Based on the conflict resolution weights of the modalities in the path, calculate the comprehensive correction factor of this path. If there is a conflict in the modalities on which the path depends (such as a contradiction between sensor and image data), the corresponding modal weight will be suppressed; if the modal data is reliable (such as the text record is consistent with the sensor), the weight will be maintained or increased. Determine the correction factor corresponding to this candidate path through the sum of the conflict resolution weights of the key modalities on which each candidate path depends, and obtain the path score corresponding to the candidate inference path by multiplying this correction factor by the path matching score. The path with the highest path score is determined as the target inference path. The core role of conflict resolution is to suppress unreliable paths. If there are serious conflicts in the modal data on which a certain path depends (such as the image shows no scratches, but the sensor indicates that the pressure exceeds the limit), its correction factor will be significantly reduced, thus avoiding misselection.
[0060] Through the above technical solution, by determining the inference modalities and their corresponding conflict resolution modal weights in each candidate inference path, the reliability of different modal data can be fully considered. Through the conflict resolution mechanism, the weight of conflicting modal data is suppressed, and the scores of candidate inference paths that depend on conflicting modal data will be affected, while the scores of paths based on reliable and consistent modal data will be relatively increased. Thus, it effectively avoids incorrect inferences caused by modal conflicts, improves the correctness of the inference results, dynamically adjusts the scores of candidate inference paths according to the real-time conflict resolution modal weights, corrects the path matching scores by updating the modal weights in real time and calculating the correction factor accordingly, so that the inference results can better adapt to environmental changes. Among many candidate inference paths, the path with the highest path score is selected as the target inference path. The path score comprehensively considers the conflict resolution modal weights. Therefore, the target inference path can more reasonably weigh the contributions of different modal data, avoids the limitations of simply selecting a path based on the path matching score, makes the finally determined inference path more in line with the actual situation, and enhances the reliability of the inference.
[0061] With the development of industrial production lines, undefined defect types usually occur. Traditional knowledge bases rely on manual rule definition, which is costly to update. The rule engines in traditional solutions cannot cover new types of defects, and static graphs need to be reconstructed offline. After performing associated path retrieval in this dynamic hypergraph knowledge network, the method further includes: when the associated path is not retrieved in the dynamic hypergraph knowledge network, parsing the multimodal input data through a large model to generate candidate defect description information for constructing candidate nodes; obtaining the existing hypergraph node set and the existing hyperedge set in the dynamic hypergraph knowledge network, and calculating the maximum semantic similarity between the candidate node and the existing hypergraph node set; according to the maximum semantic similarity, performing a node update operation in the hypergraph node set of the dynamic hypergraph knowledge network to determine the operating hypergraph node, where the node update operation includes a similar node merging operation and a new node creation operation; determining at least one associated hyperedge corresponding to the operating hypergraph node, and using a temporal graph convolutional network to perform local adjustment of the dynamic hyperedge weight to determine the dynamic hyperedge weight corresponding to the associated hyperedge, so as to perform dynamic update on the dynamic hypergraph knowledge network.
[0062] In one embodiment of this specification, when the dynamic hypergraph knowledge network does not retrieve an inference path associated with the current multimodal input data (such as image defect features, sensor anomaly data, text report conclusions), the following processing flow is started: jointly parsing the multimodal input data using a pre-trained multimodal model, extracting cross-modal semantic features, and generating candidate defect information described in natural language (such as "surface microcracks accompanied by abnormal internal temperature rise"). Encoding the candidate defect description as a semantic vector and adding it to the dynamic hypergraph knowledge network as a candidate node, where the node attributes include defect type, associated modal features, and confidence parameters. Calculating the cosine similarity between the candidate node and all nodes in the existing hypergraph node set , and screening the maximum similarity value. If the maximum similarity ≥ 0.6 (configurable), perform similar node merging, merging the candidate node into the most similar existing node and inheriting its associated hyperedges; if the similarity ≥ 0.4 but < 0.6, trigger the manual review process; if < 0.4, perform new node creation and mark it as to be verified, adding the candidate node as a new entity to the hypergraph network. Automatically creating associated hyperedges (such as "microcracks → temperature rise anomaly → material fatigue") for the newly added or merged nodes, with the initial weight set to a default value (such as 0.5); using a temporal graph convolutional network (T-GCN) to perform local weight adjustment on the associated hyperedges, and the formula is: , and are the hidden states of the nodes at both ends of the hyperedge, is a learnable parameter matrix. Only the hyperedge weights directly associated with the newly added / merged nodes are adjusted to avoid global network reconstruction and ensure update efficiency. Every time an undefined defect is detected, the knowledge network is dynamically expanded through the above process to gradually cover new patterns that cannot be handled by traditional methods (such as composite material delamination defects, micron-level deformations). Based on the online learning mechanism, the node similarity threshold and hyperedge weight parameters are automatically optimized using the feedback of actual quality inspection results (such as misjudgment rate, manual re-inspection confirmation).
[0063] That is to say, when a new defect pattern is detected, such as an undefined candidate defect description generated by an inference large model, calculate its semantic similarity with the existing nodes , if the similarity exceeds the threshold of 0.6, it is automatically added to the network; at the same time, a temporal graph convolutional network (T-GCN) is used to dynamically adjust the hyperedge weights , where is the hidden state of the nodes at both ends of the hyperedge, is the LeakyReLU activation function. By only making local weight adjustments to the associated hyperedges (instead of global graph reconstruction), T-GCN combines the incremental iteration of temporal hidden states and lightweight design to compress the time-consuming of a single update, thereby increasing the dynamic update frequency and improving the real-time response ability of quality inspection.
[0064] For example, in the quality inspection of aeroengine blades, when an undefined defect "microcracks in thermal barrier coating accompanied by abnormal high-frequency vibration" is detected for the first time, the large model analyzes multi-modal data and generates a candidate node description "associated defect of coating microcracks - vibration abnormality". Calculate the maximum similarity with the existing nodes. The similarity with the "coating spalling" node is 0.58 (<0.6), triggering the creation of a new node. A new hyperedge "microcracks → vibration abnormality → abnormal coating process parameters" is added with an initial weight of 0.5. After 3 subsequent detections of similar defects, T-GCN dynamically increases the hyperedge weight to 0.89 to form a stable inference path.
[0065] In addition, the embodiments of this specification also provide an adaptive feedback optimization method. Taking the industrial quality inspection scenario as an example, the system input end integrates three types of heterogeneous data: high-definition images of the product surface (such as detecting a scratch length of 1.2 mm), point cloud data from a laser sensor (deformation depth of 0.05 mm), and production batch record text (such as "material batch Y, heat treatment temperature deviation ±10°C"). To adapt to the industrial Internet of Things edge computing environment, an online knowledge distillation technology is used to transfer the understanding ability of multi-modal knowledge of a large model (teacher model) to a lightweight model (student model), and a distillation loss function is designed that includes the commonly used KL divergence method in the field and the loss of the defect classification task . By combining knowledge distillation with the collaborative pruning technology of neurons and knowledge nodes, the dual optimization of model complexity and knowledge base redundancy is achieved. On the premise of ensuring the integrity of the core functions of the system, the resource requirements for edge devices can be significantly reduced, providing reliable support for low-power and high-real-time deployment in the industrial Internet of Things environment.
[0066] In terms of dynamic update and conflict resolution of the knowledge base, the dynamic expansion of the knowledge base is realized through semantic similarity calculation. For example, when a new defect mode (such as microcracks caused by the thermal expansion of a ceramic coating) is detected, the large model generates candidate entity descriptions (such as "microcracks → abnormal coating thermal expansion coefficient → adjust the sintering process"), and calculates its semantic similarity with the existing nodes . After passing the standard, it is automatically added to the knowledge base and merged into the most similar node; if the similarity is ≥ 0.4 but < 0.6, the manual review process is triggered; if < 0.4, a new node is created and marked as to be verified. It should be noted that the "conflict resolution" of the present invention refers to resolving the semantic contradictions between multi-modal data (such as the conflict between images and laser point clouds), rather than the entity ambiguity processing (knowledge resolution) within the knowledge base. Conflict resolution includes two core links: conflict detection and conflict correction. The hyperedges of the knowledge base are marked with the HYPER_EDGE label and store the list of associated node IDs. For example, the hyperedge "scratch - temperature - material" is stored as: CREATE (he:HYPER_EDGE {id: "HE001", nodes: ["Node_Image123", "Node_Sensor456", "Node_Text789"]}). Through the built-in distributed query engine, it supports high-performance multi-hop reasoning (such as tracing image defects back to the production process). Through the dynamic pruning mechanism of neurons and knowledge nodes, the dual optimization of model complexity and knowledge base redundancy is achieved, ensuring the efficient operation of the system on edge devices. First, based on the parameter importance evaluation technology, the importance scores of each neuron in the neural network are calculated, and the 30% neurons with the lowest scores are removed, significantly reducing the computational overhead.
[0067] Subsequently, in the optimization of knowledge base nodes, through the node importance score formula, , the value of each node in the knowledge base is dynamically evaluated, refers to the nodes in the knowledge base, such as entities; is the set of all hyperedges associated with node ; is the weight of hyperedge e; represents the degree of node , that is, the number of connected hyperedges. For example, node can represent "scratch length > 0.5mm" or "material batch Y", and hyperedge Connect "Scratch Image", "Temperature Anomaly", and "Material Batch Y", the hyperedge Weight = 0.9, indicating a high correlation. The degree of the node "Temperature Anomaly" = 5, indicating that it is associated with 5 hyperedges. The importance of a node not only depends on the number of connected hyperedges (degree), but also is related to the semantic weight of each hyperedge. In the industrial quality inspection scenario, nodes that frequently appear and are associated with key defects (such as "Temperature Anomaly") are more important. It is updated in real time by T-GCN. When a certain material frequently triggers defects, the weight of its associated hyperedges increases, and the importance of the node increases.
[0068] Remove low-score nodes with an importance score < 0.4, such as the outdated inspection standard "Manual Visual Inspection Threshold 0.8mm". At the same time, merge semantically similar nodes with a semantic similarity > 0.7, such as "Ra0.8" and "Ra1.6" are merged into the parent node "Excessive Roughness". This optimization reduces a large amount of computational effort in the knowledge base, reduces the edge-side latency, overcomes the problem of the rigidity of the traditional graph structure, and ensures the efficient operation of the system in a high-load production line environment.
[0069] By constructing a dynamic knowledge modeling framework, a hybrid reasoning optimization mechanism, a cross-modal fusion and conflict resolution method, and an adaptive feedback optimization technology, the unified representation and efficient reasoning of multi-modal data are realized. And the compatibility of multi-modal data is achieved using heterogeneous feature extraction and unified embedding mapping. The unified representation framework is based on dynamic knowledge graph technology. As an extended form of the knowledge graph through the hypergraph structure, it supports the expression of the topological relationship of multi-modal entities, overcoming the limitation that the traditional knowledge graph only supports binary relationships; following the principles of "dynamic embedding, semantic-driven, and real-time optimization", it focuses on the core links such as multi-modal dynamic knowledge modeling → hybrid reasoning path generation → cross-modal fusion and conflict resolution → adaptive feedback optimization, forming a closed-loop system from data collection to knowledge update, which can achieve continuous optimization of data → knowledge → decision → feedback, reducing the knowledge base update latency from the hour level to the second level, and supporting real-time process adjustment of the production line. It should be noted that through the collaborative design of hypergraph theory, dynamic evolution algorithms, and multi-modal fusion mechanisms in the embodiments of this specification, a new generation of knowledge base construction and knowledge reasoning method for industrial quality inspection scenarios is proposed.
[0070] Through the embodiments of this specification, with the pre-constructed dynamic hypergraph knowledge network, flexible adjustment can be made according to real-time production data and continuously updated industry rules. When new quality problems or data characteristics emerge, the relationships between nodes and hyperedges in the hypergraph can be updated to quickly adapt to changes, effectively solving the problem of lacking support for dynamic knowledge evolution. Traditional static knowledge graphs cannot fuse multi-modal semantics in real time, while industrial quality inspection needs to process high-dimensional heterogeneous data such as images, sensor time-series data, and unstructured text. Through the embodiments of this specification, by obtaining the current inference task information, dynamic weight allocation is performed on the multi-modal input data to determine the multi-modal semantic weight information and the current fused semantic features, realizing the real-time fusion of information between different modal data. Traditional fixed weight allocation strategies ignore the differences in task scenarios. This technical solution performs dynamic weight allocation on the multi-modal input data according to the task description text in the current inference task information, enabling the system to pay more attention to the modal data related to the task. Traditional methods use fixed conflict handling rules such as simple threshold filtering, which are prone to over-reliance on unreliable features. Through semantic conflict detection of the multi-modal input data, when there is a semantic conflict, the weights are adjusted based on the multi-modal semantic weight information to determine the conflict resolution modal weight information. By reducing the weights of conflicting modal data or increasing the weights of reliable modal data, a reasonable trade-off can be made between conflicting data, making the inference process more inclined to rely on reliable data, effectively resolving semantic conflicts, avoiding missed detection of key defects, and improving the accuracy and reliability of the inference results. By comprehensively applying technologies such as dynamic hypergraph knowledge networks, large models, and conflict resolution, knowledge inference can be quickly and accurately performed to determine structured decision-making text. The local update mechanism of the dynamic hypergraph knowledge network only adjusts the weights of associated hyperedges based on the temporal graph convolutional network (T-GCN) without global reconstruction, and the single update takes a short time, ensuring the real-time performance of the system. Through dynamic weight allocation, semantic conflict detection and resolution, and the powerful semantic understanding and reasoning capabilities of the large model, the accuracy of the inference results is improved. In complex industrial quality inspection scenarios, quality problems of products can be timely detected, the types and root causes of defects can be accurately analyzed, and reasonable decision-making suggestions and warning signals can be given, effectively meeting the requirements of industrial quality inspection for real-time performance and accuracy.
[0071] The embodiments of this specification also provide a cross-modal knowledge inference device for industrial quality inspection, as Figure 4 shown. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0072] An embodiment of this specification also provides a non-volatile computer storage medium storing computer-executable instructions, which are configured to execute the above method.
[0073] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.
[0074] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A cross-modal knowledge reasoning method for industrial quality inspection, characterized in that The method includes: Obtain the current inference task information, and based on the task description text in the current inference task information, perform dynamic weight allocation on the multimodal input data to determine multimodal semantic weight information, and determine the current fused semantic features corresponding to the multimodal input data; Perform semantic conflict detection on the multimodal input data. When there is a semantic conflict in the multimodal input data, based on the multimodal semantic weight information, adjust the weights of the multimodal semantic weight information to determine conflict resolution modal weight information; According to the conflict resolution modal weight information and the current fused semantic features, perform knowledge reasoning through a large model and a pre-constructed dynamic hypergraph knowledge network to determine structured decision-making text, where the structured decision-making text includes any one or more of defect type, root cause analysis, association path, decision suggestion, and warning signal; Based on the multimodal semantic weight information, adjust the weights of the multimodal semantic weight information to determine conflict resolution modal weight information, specifically including: Determine at least two conflicting modalities among multiple modalities, and obtain the cross-modal distance index between the conflicting modalities; Determine the weight penalty factor corresponding to the conflicting modality through the cross-modal distance index and a preset distance index threshold; According to the weight penalty factor and the multimodal semantic weight information, adjust the weights of the multiple modalities respectively to determine conflict resolution modal weight information; According to the weight penalty factor and the multimodal semantic weight information, adjust the weights of the multiple modalities respectively to determine conflict resolution modal weight information, specifically including: Perform weight penalty on the conflicting modality according to the weight penalty factor to determine the current penalty weight parameter corresponding to each conflicting modality; Obtain the multimodal semantic weight information corresponding to each conflicting modality, and determine the conflict penalty weight through the multimodal semantic weight information corresponding to the conflicting modality and the current penalty weight parameter; Based on the conflict penalty weight, enhance the weights of the reliable modalities among the multiple modalities to determine the conflict enhancement weight corresponding to the reliable modality; Determine the conflict resolution modal weight information through the current penalty weight parameter corresponding to each conflicting modality and the conflict enhancement weight corresponding to the reliable modality; According to the conflict resolution modal weight information and the current fused semantic features, perform knowledge reasoning through a large model and a pre-constructed dynamic hypergraph knowledge network to determine structured decision-making text, specifically including: Determine candidate inference paths in the dynamic hypergraph knowledge network according to the current fused semantic features; Generate path description information corresponding to each candidate inference path, splice the path description information with the pre-determined corresponding input problem description information, and score each candidate inference path through the large model to determine the path matching degree score of each candidate inference path; Correct the path matching degree score of each candidate inference path through the conflict resolution modal weight information to determine the target inference path, so as to determine the structured decision-making text; Based on the conflict resolution modality weight information, correct the path matching degree scores of each candidate inference path to determine the target inference path, which specifically includes: Determine at least one inference modality in each candidate inference path to determine the conflict resolution modality weight corresponding to each inference modality; Determine the correction factor of the candidate inference path through the conflict resolution modality weight corresponding to each inference modality; Based on the path matching degree scores of the candidate inference paths, use the correction factor to determine the path score corresponding to each candidate inference path; Among multiple candidate inference paths, determine the candidate inference path with the highest path score as the target inference path.
2. The cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, wherein According to the task description text in the current inference task information, perform dynamic weight allocation on the multi-modal input data to determine the multi-modal semantic weight information, which specifically includes: Generate a dynamic task query vector corresponding to the current inference task through the task description text; Construct cross-modal key-value pairs corresponding to each modality in the multi-modal input data; According to the dynamic task query vector and the cross-modal key-value pairs, perform dynamic weight allocation on multiple modalities to determine the multi-modal semantic weight corresponding to each modality.
3. A cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, characterized in that According to the current fused semantic features, determine candidate inference paths in the dynamic hypergraph knowledge network, which specifically includes: Determine the initial candidate paths according to the current fused semantic features and predefined industry rules; Through the path nodes in the initial candidate paths, perform associated path retrieval in the dynamic hypergraph knowledge network to generate a set of candidate inference paths, where the set of candidate inference paths includes multiple candidate inference paths, and each candidate inference path includes a chain of RDF triples connected by dynamic hyperedges.
4. A cross-modal knowledge reasoning method for industrial quality inspection according to claim 1, characterized in that After performing associated path retrieval in the dynamic hypergraph knowledge network, the method further includes: When the associated path is not retrieved in the dynamic hypergraph knowledge network, parse the multi-modal input data through a large model to generate candidate defect description information to construct candidate nodes; Obtain the existing hypergraph node set and existing hyperedge set in the dynamic hypergraph knowledge network, and calculate the maximum semantic similarity between the candidate nodes and the existing hypergraph node set; According to the maximum semantic similarity, perform a node update operation in the hypergraph node set of the dynamic hypergraph knowledge network to determine the operating hypergraph nodes, where the node update operation includes a similar node merging operation and a new node creation operation; Determine at least one associated hyperedge corresponding to the operating hypergraph nodes, and use a temporal graph convolutional network to perform local adjustment of the dynamic hyperedge weights to determine the dynamic hyperedge weights corresponding to the associated hyperedges, so as to dynamically update the dynamic hypergraph knowledge network.
5. A cross-modal knowledge reasoning device for industrial quality inspection, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.
6. A non-volatile computer storage medium stores computer-executable instructions, characterized in that, The computer-executable instructions are configured to: execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Short video rumor detection method and system for complex incomplete data scene
CN119166854A
Multimodal entity relationship extraction method based on hypergraph neural network
CN119830918A