Anomaly detection fusion dynamic knowledge base retrieval enhancement generation method and system
By constructing a dynamic knowledge base and performing anomaly prototype retrieval and path reasoning, the consistency problem of multi-source evidence under abnormal conditions is solved, improving the accuracy of data processing and the efficiency of resource scheduling, and realizing more granular context control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING HONGTAIJIE TECH CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-21
AI Technical Summary
Under conditions of continuous change in abnormal states, existing technologies struggle to ensure that the organization, retrieval, and generation of multi-source evidence remain consistent with the current anomaly. This can easily introduce irrelevant and conflicting information, affecting the accuracy of computer data processing and the efficiency of resource scheduling.
An enhanced generation method based on anomaly detection and dynamic knowledge base is adopted. Anomalies are detected by receiving multimodal data and real-time data streams. A dynamic knowledge base is constructed, including a regional evidence layer, a semantic vector layer, and a cross-modal knowledge graph layer. The hierarchical anomaly prototype layer is updated according to the current anomaly state. Anomaly prototype retrieval, multimodal semantic retrieval, and path reasoning are performed to generate an evidence set.
It improves the accuracy of processing anomaly-related data, enhances the ability to identify correspondences and conflicts between evidence from different sources, reduces the waste of computing resources caused by irrelevant data participating in retrieval and reasoning, and improves resource scheduling efficiency and context scope control.
Smart Images

Figure CN122087133B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anomaly retrieval enhancement generation technology in industrial intelligent operation and maintenance, and in particular to a retrieval enhancement generation method and system that integrates anomaly detection with a dynamic knowledge base. Background Technology
[0002] As heterogeneous data such as documents, tables, and images are used in both intelligent question answering and anomaly handling, along with real-time business data, existing search enhancement generation solutions are beginning to be applied to scenarios such as fault analysis, operational decision-making, and knowledge-based question answering. These scenarios require a high degree of consistency between search results and the current business status.
[0003] Existing technologies typically employ static knowledge bases and fixed retrieval processes to handle multimodal information. Real-time anomaly data is mostly used only for individual alarms and fails to continuously participate in subsequent knowledge organization and retrieval control. This results in the system still retrieving data from a large amount of historical information in a uniform manner when processing anomaly-related queries, making it difficult to promptly highlight the evidence most relevant to the current anomaly.
[0004] Therefore, the core problem of existing technologies is that, under the condition of continuous changes in abnormal states, it is difficult to keep the organization, retrieval and generation process of multi-source evidence consistent with the current abnormality, which easily introduces irrelevant and conflicting information, thereby affecting the accuracy of computer data processing, conflict identification ability and resource scheduling efficiency in the retrieval and reasoning stages. Summary of the Invention
[0005] To address the aforementioned issues of retrieval mismatch, evidence conflict, and inefficient resource scheduling, this invention proposes a retrieval enhancement generation method and system that integrates anomaly detection with a dynamic knowledge base.
[0006] The present invention achieves the above objectives through the following technical solutions:
[0007] A method for enhancing retrieval generation by fusing anomaly detection with a dynamic knowledge base, the method comprising:
[0008] Receive multimodal data and real-time data stream, perform anomaly detection on the real-time data stream to obtain the current anomaly state, and perform structured extraction on the multimodal data based on the current anomaly state to obtain the object to be parsed;
[0009] Based on the user query and the current abnormal state, the target parsing object is determined from the object to be parsed and parsing is performed to obtain the region-level semantic representation and the document-level semantic representation;
[0010] A dynamic knowledge base is constructed based on the region-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a region evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer. The hierarchical anomaly prototype layer is updated according to the current anomaly state.
[0011] Based on the user query, the current abnormal state, and historical feedback, the retrieval execution order corresponding to the current request is determined, and abnormal prototype retrieval is performed in the hierarchical abnormal prototype layer, multimodal semantic retrieval is performed in the semantic vector layer and the regional evidence layer, and path reasoning is performed in the cross-modal knowledge graph layer to obtain candidate evidence;
[0012] The candidate evidence is fused and sorted to form an evidence set, and the evidence set is used as context input to generate a model to output a response, analysis report, or handling suggestion.
[0013] A further improvement of the present invention is that the multimodal data includes one or more of the following: doc documents, excel spreadsheets, ppt presentations, pdf files, and images;
[0014] Anomaly detection of the real-time data stream includes: extracting temporal change features and object association features within a continuous time window; determining the current anomaly state based on the temporal deviation of the temporal change features relative to the historical baseline and the association deviation of the object association features relative to the association baseline map; wherein the current anomaly state includes at least anomaly category, anomaly degree, associated objects, and anomaly time interval.
[0015] Based on the anomaly category, anomaly severity, associated objects, and anomaly time interval, anomaly association structured extraction is performed on the multimodal data, including: locating document layout areas corresponding to the associated objects and anomaly time intervals from doc documents, pdf files, and ppt presentations, and extracting the corresponding layout hierarchy; locating table cells corresponding to the associated objects and anomaly time intervals from Excel spreadsheets, and extracting the corresponding header dependencies; locating image object areas corresponding to the associated objects and anomaly time intervals from images, and extracting the corresponding spatial relationships; and determining the priority extraction order of document layout areas, table cells, and image object areas according to the anomaly category.
[0016] When the degree of abnormality reaches a preset level, the abnormal time interval is extended forward and backward by a preset time length, and joint extraction is performed on the adjacent document layout area, adjacent table cell or adjacent image object area corresponding to the associated object.
[0017] A cross-modal correspondence is established for the extraction results where the object identifiers after normalization are consistent across different modalities and the normalized time identifiers fall within the abnormal time interval. The extraction results are then associated and stored with the object identifier, time identifier, and modal identifier to obtain the object to be parsed.
[0018] A further improvement of the present invention lies in determining the target parsing object from the object to be parsed and performing parsing, specifically including:
[0019] Calculate the first The target parsing score of the object to be parsed The formula is:
[0020] ;
[0021] In the formula, For user queries and the The semantic relevance of each object to be parsed; For the current abnormal state and the first The degree of abnormal relevance of the object to be parsed; For users to query the current abnormal status of the first The coupling correlation of the objects to be resolved; For the exception category and the first The information type matching degree of the object to be parsed; For the first The completeness of the cross-modal correspondence of the objects to be parsed; For the first The structural integrity of the object to be parsed; For the first Boundary reliability of the object to be resolved; For the first The normalization parsing cost of an object to be parsed; Analyze the scoring weights for the target, and satisfy the following conditions: ;
[0022] According to the target analysis score Sort the objects to be parsed, and select the objects whose target parsing score is greater than a preset threshold or whose ranking is in the top predetermined number as the target parsing objects;
[0023] Fine-grained parsing is performed on the target parsing object to obtain the corresponding region-level semantic representation. ;
[0024] Multiple region-level semantic representations with the same source document identifier, the same object identifier, and time identifier falling within the same anomalous time interval are aggregated to obtain document-level semantic representations. The aggregation formula is as follows: ;and ;
[0025] In the formula, For the first The document-level semantic representation of each target parsing object; This refers to the set of target parsing objects that have the same source document identifier, the same object identifier, and whose time identifiers fall within the same abnormal time interval. For the first The target parsing objects are in the target parsing object set Aggregate weights in; For the first The region-level semantic representation corresponding to each target parsing object; For the first Time consistency of the target parsing objects; For the target parsing object set The target parsing object number in the middle; The aggregate weight parameters satisfy the following conditions: .
[0026] A further improvement of the present invention is that a dynamic knowledge base is constructed based on the region-level semantic representation and the document-level semantic representation, including:
[0027] Each region-level semantic representation, along with its corresponding object identifier, time identifier, modality identifier, and corresponding page hierarchy relationship, header dependency relationship, or spatial location relationship, is constructed into a region evidence entry and stored in the region evidence layer.
[0028] Establish a one-to-one mapping relationship between each region-level semantic representation and its corresponding region evidence entry, establish an aggregate mapping relationship between each document-level semantic representation and its corresponding one or more region evidence entries, and store them in the semantic vector layer;
[0029] The cross-modal knowledge graph layer is constructed based on the cross-modal correspondence, object identifier, time identifier, and abnormal time interval. The nodes in the cross-modal knowledge graph layer include at least text entity nodes, table field nodes, image object nodes, regional evidence nodes, abnormal event nodes, and handling action nodes. The nodes are connected to each other through at least object association edges, time association edges, structural association edges, evidence support edges, and handling association edges.
[0030] Based on the region-level semantic representation, document-level semantic representation, current abnormal state, and the graph sub-path in the cross-modal knowledge graph layer corresponding to the current abnormal state, the hierarchical abnormal prototype layer is constructed, which includes category-level abnormal prototype, semantic-level abnormal prototype, and instance-level abnormal prototype.
[0031] Each instance-level anomaly prototype is associated with at least the anomaly category, anomaly degree, associated object, anomaly time interval, corresponding regional evidence entry, corresponding document-level semantic representation, and corresponding graph sub-path; each semantic-level anomaly prototype is associated with at least multiple instance-level anomaly prototypes with a semantic consistency of not less than a preset threshold; each category-level anomaly prototype is associated with at least one or more semantic-level anomaly prototypes under the same anomaly category.
[0032] A further improvement of the present invention is that updating the hierarchical exception prototype layer according to the current exception state includes:
[0033] Based on the current abnormal state and the corresponding region-level semantic representation, document-level semantic representation, and graph sub-path, generate the current abnormal event representation.
[0034] Calculate the first The prototype matching score between existing instance-level anomaly prototypes and the current anomaly event representation. The formula is:
[0035] ;
[0036] In the formula, Consistency of anomaly categories; The degree of anomaly matching; For the matching degree of the associated objects; This refers to the overlap of abnormal intervals. For the current abnormal event characterization and the first Semantic consistency among existing instance-level exception prototypes; Match weights to the prototype, and satisfy the following conditions: ;
[0037] The existing instance-level exception prototype with the highest prototype matching score is selected as the candidate matching prototype.
[0038] When the prototype matching score of the candidate matching prototype is not less than the first preset threshold, the candidate matching prototype is updated, and the central representation of the semantic-level abnormal prototype and the central representation of the category-level abnormal prototype are corrected according to the updated instance-level abnormal prototype representation.
[0039] When the prototype matching score of the candidate matching prototype is less than the first preset threshold and not less than the second preset threshold, a new instance-level abnormal prototype is created under the semantic-level abnormal prototype to which the candidate matching prototype belongs, and the central representation of the corresponding semantic-level abnormal prototype and the central representation of the corresponding category-level abnormal prototype are corrected.
[0040] When there is no existing instance-level exception prototype with a prototype matching score not less than the second preset threshold, a new instance-level exception prototype and a new semantic-level exception prototype are created based on the current exception event representation, and the new semantic-level exception prototype is attached to the category-level exception prototype of the corresponding exception category; when the category-level exception prototype of the corresponding exception category does not exist, a new category-level exception prototype is created.
[0041] A further improvement of the present invention is that, based on the user query, the current abnormal state, and historical feedback, the retrieval execution order corresponding to the current request is determined, including:
[0042] A strategy experience layer is set in the dynamic knowledge base. Each strategy experience record in the strategy experience layer includes at least historical user query characteristics, historical abnormal state characteristics, historical retrieval execution order, and normalized feedback reward value. ;
[0043] Calculate the current user query and the first Query similarity between historical user query features in the strategy experience record Calculate the current abnormal state and the first abnormal state. State similarity between historical anomalous state features in the policy experience record , and then calculate the first The sequential matching score corresponding to each strategy experience record The formula is: ;
[0044] In the formula, For sequential matching weights, the following conditions must be met: ;
[0045] Selecting the order matching score The historical search execution order that is greater than a first preset threshold or ranks among the top first predetermined number is used as a candidate historical search execution order;
[0046] Based on the current degree of anomaly, the candidate historical retrieval execution order is expanded, pruned, or rearranged to obtain the retrieval execution order corresponding to the current request;
[0047] The expansion includes adding pre-execution of anomaly prototype retrieval or supplementary execution of path reasoning when the current anomaly level is not lower than a first preset level; the pruning includes reducing the number of branches of multimodal semantic retrieval or reducing the expansion depth of path reasoning when the current anomaly level is not higher than a second preset level; and the reordering includes adjusting the execution order of anomaly prototype retrieval, multimodal semantic retrieval, and path reasoning according to the normalized feedback reward value and the priority rules corresponding to the anomaly category in the current anomaly state.
[0048] A further improvement of the present invention is that anomaly prototype retrieval is performed in the hierarchical anomaly prototype layer, and multimodal semantic retrieval is performed in the semantic vector layer and the region evidence layer, including:
[0049] In the hierarchical anomaly prototype layer, based on the matching results between the current anomaly state and each anomaly prototype in terms of anomaly category, anomaly degree, associated object, anomaly time interval and prototype representation, the prototype retrieval score corresponding to each anomaly prototype is calculated, and anomaly prototypes with prototype retrieval scores greater than the second preset threshold or ranked in the top second predetermined number are selected as anomaly prototype retrieval results.
[0050] In the semantic vector layer and the regional evidence layer, based on the semantic similarity between the user query and each semantic vector record, the abnormal correlation between the current abnormal state and each semantic vector record, and the structural completeness of the regional evidence entries corresponding to each semantic vector record, the semantic retrieval score corresponding to each semantic vector record is calculated, and the semantic vector records with a semantic retrieval score greater than the third preset threshold or ranked in the top third predetermined number are selected as multimodal semantic retrieval results.
[0051] Based on the abnormal prototype retrieval results, select the top four predetermined number of abnormal prototypes ranked by prototype similarity. Then, according to the hierarchical type of the abnormal prototype, call the corresponding instance-level abnormal prototype representation, semantic-level abnormal prototype center representation, or category-level abnormal prototype center representation. Finally, weighted aggregation is performed according to the corresponding prototype retrieval score to generate the [number missing]th [item missing]. The prototype guiding vector corresponding to each candidate piece of evidence ;
[0052] The first Each candidate piece of evidence is encoded into a candidate evidence vector. For the candidate evidence vector and the prototype guiding vector Perform denoising fusion to obtain the first... Denoising fusion results corresponding to each candidate piece of evidence The formula is:
[0053] ;
[0054] in, ;
[0055] In the formula, This represents the number of denoising expert functions; and This refers to the index of the denoising expert function; For the first A noise reduction expert function; For the first The denoising expert function for the first... Gating weights for each candidate piece of evidence; This indicates vector concatenation; and This represents the gating parameter vector corresponding to the denoising expert function;
[0056] Based on the denoising fusion results The candidate evidence is reordered, and the candidate evidence whose reordered score meets the preset conditions is selected as the input for path reasoning.
[0057] A further improvement of the present invention is that path reasoning is performed in the cross-modal knowledge graph layer, including:
[0058] The corresponding regional evidence node, text entity node, table field node, or image object node of the path reasoning input is taken as the starting node, and the abnormal event node or handling action node corresponding to the current abnormal state is taken as the target node.
[0059] The starting node is extended under constraints along the object-related edge, time-related edge, structural-related edge, evidence-supporting edge, and disposal-related edge to obtain one or more candidate paths;
[0060] A path reasoning score is calculated for each candidate path, and the path reasoning score is determined at least based on object consistency, temporal consistency, structural consistency and disposal relevance in the path;
[0061] Select candidate paths whose path reasoning scores are greater than the fourth preset threshold or whose ranking is among the top five predetermined number as valid reasoning paths;
[0062] The evidence content corresponding to the regional evidence nodes, abnormal event nodes, and handling action nodes on the effective reasoning path is collected to obtain the path reasoning result;
[0063] The path reasoning result is combined with the abnormal prototype retrieval result and the candidate evidence whose reordering score meets the preset conditions to obtain the candidate evidence.
[0064] A further improvement of the present invention is that the candidate evidence is fused and ranked to form an evidence set, and the evidence set is used as context input to generate a model, including:
[0065] The candidate evidence is clustered according to object identifier, time identifier and source location. Deduplication and merging are performed on candidate evidence with the same object identifier, compatible time identifier and overlapping source location or content semantic similarity not less than the fifth preset threshold to obtain one or more candidate evidence clusters.
[0066] For each candidate evidence cluster, determine the main evidence and supplementary evidence within the cluster. The main evidence is the candidate evidence with the highest fusion and ranking priority in the candidate evidence cluster, and the supplementary evidence is attached to the main evidence. For candidate evidence with the same object identifier and compatible time identifier but inconsistent conclusions, set a conflict flag.
[0067] Based on user query relevance, current anomaly state consistency, anomaly prototype support, path reasoning support, cross-modal complementarity, source completeness, and conflict markers, the main evidence is fused and ranked.
[0068] According to the preset evidence type quota, evidence is selected from each of the main evidences after fusion and sorting to form the evidence set, wherein the evidence types include at least abnormal prototype evidence, multimodal semantic evidence and path reasoning evidence;
[0069] The evidence set is organized into a contextual sequence according to the order of current abnormal state summary, key evidence, conflict explanation and handling basis, and each piece of evidence is attached with a source location identifier and evidence type identifier;
[0070] The task identifier is determined based on the user query type, and the context sequence and the task identifier are input into the generation model to output an answer, analysis report, or handling suggestion.
[0071] A retrieval enhancement and generation system that integrates anomaly detection with a dynamic knowledge base, the system comprising:
[0072] The data access module is used to receive multimodal data and real-time data streams;
[0073] An anomaly detection module is used to perform anomaly detection on the real-time data stream and obtain the current anomaly status;
[0074] The structured extraction module is used to perform structured extraction on the multimodal data based on the current abnormal state to obtain the object to be parsed;
[0075] The parsing module is used to determine the target parsing object from the object to be parsed based on the user query and the current abnormal state, and to perform parsing to obtain the region-level semantic representation and the document-level semantic representation.
[0076] A dynamic knowledge base construction module is used to construct a dynamic knowledge base based on the region-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a region evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer.
[0077] The prototype update module is used to update the hierarchical exception prototype layer according to the current exception state.
[0078] The sequence decision module is used to determine the retrieval execution order corresponding to the current request based on the user query, the current abnormal state, and historical feedback.
[0079] The retrieval execution module is used to perform anomaly prototype retrieval in the hierarchical anomaly prototype layer and multimodal semantic retrieval in the semantic vector layer and the regional evidence layer according to the retrieval execution order.
[0080] The path reasoning module is used to perform path reasoning in the cross-modal knowledge graph layer to obtain candidate evidence;
[0081] The evidence fusion generation module is used to fuse and sort the candidate evidence to form an evidence set, and use the evidence set as context input to generate a model to output a response, analysis report or disposal suggestion.
[0082] The beneficial effects of this invention are as follows: By introducing abnormal states into the entire process of multimodal data processing, knowledge organization, retrieval control, and evidence generation, the evidence extraction and retrieval process can revolve around the current abnormal dynamics. This improves the accuracy of abnormal-related data processing, enhances the ability to identify correspondences and conflicts between evidence from different sources, and reduces the waste of computational resources caused by irrelevant data participating in retrieval and reasoning. Furthermore, through hierarchical knowledge organization, sequential retrieval, and evidence aggregation input, the generative model receives effective evidence that has been filtered, correlated, and sorted, rather than uncontrolled raw context. This improves the orderliness of evidence writing and retrieval, enhances storage and write control, improves resource scheduling efficiency in the retrieval, reasoning, and generation chain, and facilitates finer-grained context scope control and access control. Attached Figure Description
[0083] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.
[0085] like Figure 1 As shown, this is an embodiment of the present invention, which provides a retrieval enhancement generation method for anomaly detection fused with a dynamic knowledge base, including the following steps:
[0086] S1: Receive multimodal data and real-time data stream, perform anomaly detection on the real-time data stream to obtain the current anomaly state, and perform structured extraction on the multimodal data based on the current anomaly state to obtain the object to be parsed.
[0087] Multimodal data includes one or more of the following: doc documents, Excel spreadsheets, PowerPoint presentations, PDF files, and images. Each file is assigned a file number, object identifier, time identifier, modality identifier, and version number during the data acquisition phase. The object identifier represents the device, component, work order, project, indicator, or business entity; the time identifier represents the file creation time, recording time, shooting time, or effective time; and the modality identifier distinguishes between doc documents, Excel spreadsheets, PowerPoint presentations, PDF files, and images. Real-time data streams are continuously fed in by the field acquisition system.
[0088] Taking manufacturing equipment operation and maintenance as an example, multimodal data includes PDF files corresponding to maintenance manuals, DOC documents corresponding to maintenance records, Excel spreadsheets corresponding to spare parts ledgers, PPT presentations corresponding to reporting materials, and inspection photos. Real-time data streams include vibration values, temperature values, current values, pressure values, and status codes continuously collected during equipment operation. Each piece of multimodal data and each piece of real-time data is pre-associated with an object identifier and a time identifier for subsequent cross-modal association.
[0089] Anomaly detection of real-time data streams includes: extracting temporal change features and object association features within a continuous time window; determining the current anomaly state based on the temporal deviation of the temporal change features relative to the historical baseline and the association deviation of the object association features relative to the association baseline map; the current anomaly state includes at least the anomaly category, anomaly degree, associated objects, and anomaly time interval.
[0090] In one implementation, the real-time data stream is first divided into continuous time windows, with a window length of 60 seconds and a sliding step of 10 seconds. For each time window, the system extracts temporal variation features and object association features.
[0091] Temporal variation features reflect the changing state of a single object within the current time window, and include at least the mean, standard deviation, slope of change, kurtosis, energy value, and first-order difference statistics. Object association features reflect the linkage relationship between multiple objects within the current time window, and include at least the degree of synchronous change between objects, the strength of topological adjacency relationships, and the sequence of triggering relationships. To avoid calculation bias caused by different sensor units, missing value imputation, outlier truncation, and standardization are performed on the raw real-time data before extracting the above features.
[0092] Data from the historical normal operation phase is used to construct the historical baseline. The historical baseline includes a baseline vector of time-series variation features and a baseline vector of object association features. The time-series deviation and association deviation values corresponding to the current time window are obtained by comparing the current features with the historical baseline, respectively. The comprehensive deviation score can be defined as: ,in, This represents the overall deviation score corresponding to the current time window; Indicates timing deviation value; Indicates the correlation deviation value; and Represents the weighting coefficients, satisfying In this embodiment, we take , .
[0093] At the implementation level, the anomaly detection module includes a temporal coding branch, an object association coding branch, and a state determination branch. The temporal coding branch receives the real-time sequence within the current time window and outputs a temporal feature vector; the object association coding branch receives the object relationship graph corresponding to the current time window and outputs an association feature vector; the state determination branch concatenates the two feature vectors and inputs them into a fully connected network, outputting the anomaly category, anomaly severity, associated objects, and anomaly time interval. In one implementation, the temporal coding branch uses a two-layer long short-term memory network, with 128 hidden units in both the first and second layers; the object association coding branch uses a two-layer graph attention network, with each layer having an output dimension of 64; the state determination branch uses a two-layer fully connected network, with the first layer having an output dimension of 128, and the second layer connecting to the anomaly category output head and the anomaly severity output head, respectively.
[0094] The model is trained using labeled historical sample windows. Training samples include at least normal sample windows and abnormal sample windows, with the abnormal sample windows accompanied by anomaly category and anomaly severity labels. The training loss can be written as: ,in, Indicates the total loss; This represents the loss for classifying anomaly categories; Regression loss indicates the degree of anomaly; This represents the weighting coefficient, which can be set from 0.1 to 1.0. In this embodiment, it is set to... .
[0095] The anomaly category is determined by the anomaly category output header, and the anomaly severity is determined by the anomaly severity output header and the overall deviation score. Associated objects are determined by the object influence score output from the object association coding branch; the object with the highest influence score is selected as the associated object. Anomaly time intervals are obtained by merging consecutive anomaly windows; if the interval between two adjacent anomaly windows is no greater than one sliding step, they are merged into the same anomaly time interval.
[0096] After obtaining the anomaly category, anomaly severity, associated objects, and anomaly time interval, anomaly association structure extraction is performed on the multimodal data based on these four pieces of information, including:
[0097] For DOC documents, PDF files, and PPT presentations, first perform layout analysis to divide each page or slide into multiple document layout areas. Each document layout area includes at least a title area, body text area, table area, image area, figure caption area, and header / footer area. Each document layout area records the page number, area bounding box coordinates, area type, and hierarchical number. The hierarchical number is used to describe the layout hierarchy relationships between titles and body text, body text and figure captions, and higher-level headings and lower-level headings.
[0098] The target document layout area is located based on associated objects and abnormal time intervals. The location process consists of two steps: object location and time location. Object location is used to find the object name, alias, device number, component number, or business number corresponding to the associated object in the document content; time location is used to find the date, timestamp, time period description, or event time falling within the abnormal time interval in the document content. When both object location and time location are successful, the corresponding document layout area is determined as the valid extraction area. Subsequently, the parent title and adjacent body paragraphs corresponding to the valid extraction area are further traced back to extract the page hierarchy relationship, and the document layout area and page hierarchy relationship are written together into the extraction result. For example, the document layout area location adopts a two-level processing method. The first level performs coarse location within the entire page and outputs candidate areas; the second level performs fine analysis within the candidate areas and outputs the final area content and hierarchy relationship. Through the above processing method, indiscriminate analysis of the entire document can be avoided.
[0099] For Excel spreadsheets, first parse the workbook structure, worksheet structure, merged cell information, formula results, and cell coordinates, and then locate the target table cell based on the associated objects and abnormal time intervals.
[0100] When locating an object, the system searches for related objects in cell text, comment text, named cell names, and formula reference descriptions. When locating a time interval, the system searches for time expressions falling within the abnormal time range in date cells, period cells, header cells, and formula reference areas. When both object and time locators are found, the corresponding cell is determined as a valid extraction unit.
[0101] After locating the valid extraction cell, extract the header dependencies. Header dependencies clarify the business meaning of cell values. For example, header dependencies are extracted according to the following rules: search upwards along the column containing the valid extraction cell to obtain the header; search leftwards along the row containing the valid extraction cell to obtain the row header; if multiple levels of headers exist, continue tracing upwards and leftwards until the topmost and leftmost headers are reached. Finally, combine the valid extraction cell, row header, header, and worksheet identifier into a single header dependency record.
[0102] When merged cells exist in the table, the merged cells are first logically expanded into multiple basic coordinate units, then the search process described above is executed. After the search is complete, the basic coordinate units are mapped back to the original merged area. This method ensures that the table header dependencies are clear and avoids ambiguity caused by merged cells.
[0103] For images, image object detection or instance segmentation is first performed to obtain multiple image object regions. Each image object region records at least the bounding box coordinates, object category, object confidence score, and source image identifier. Object categories can include device body, component area, instrument area, nameplate area, damage area, and alarm marker area. Time stamps are extracted from image metadata, filenames, external index tables, or accompanying descriptive text, and images with time stamps falling within abnormal time intervals are filtered out. Then, valid image object regions are located based on the identifier information of associated objects in the image. For valid image object regions, spatial relationships are further extracted. Spatial relationships include at least left, right, above, below, containment, intersection, and adjacency. For example, spatial relationships are determined by the positional relationships between bounding boxes: when the right boundary of one image object region is to the left of the left boundary of another image object region, the former image object region is determined to be to the left of the latter image object region; when the lower boundary of one image object region is above the upper boundary of another image object region, the former image object region is determined to be above the latter image object region; when one image object region is completely inside another image object region, an inclusion relationship is established between the former and latter image object regions.
[0104] Anomaly categories are used to determine the priority order for extracting document layout areas, table cells, and image object areas. To achieve this functionality, a mapping rule between anomaly categories and information types needs to be established in advance.
[0105] For example, when the anomaly category is mechanical wear, table cells are extracted first, followed by document layout areas, and finally image object areas; when the anomaly category is leakage, image object areas are extracted first, followed by document layout areas, and finally table cells; when the anomaly category is parameter drift, table cells and document layout areas are extracted first. This priority order can be implemented using a preset rule table or a trained priority prediction model. If a prediction model is used, it takes the anomaly category code as input, outputs ranking scores for the three information types, and then determines the extraction priority order based on these ranking scores.
[0106] When the anomaly level reaches a preset threshold, the system extends the anomaly time interval forward and backward by a preset time length, and performs joint extraction. In this embodiment, Level 1 anomalies do not undergo time extension, Level 2 anomalies are extended forward and backward by 15 minutes, and Level 3 anomalies are extended forward and backward by 30 minutes. The extended time interval can be represented as:
[0107] ;
[0108] in, This indicates the expanded abnormal time interval; Indicates the start time of the original anomaly time interval; Indicates the end time of the original abnormal time interval; Indicates the length of time for forward expansion; This indicates the length of time extended backwards. The time unit is consistently minutes.
[0109] Within the expanded abnormal time interval, in addition to extracting directly hit document layout areas, table cells, and image object areas, regions adjacent to the hit objects are also extracted. For document layout areas, adjacency can be defined as adjacent areas sharing boundaries, being consecutive vertically, or being under the same parent heading; for table cells, adjacency can be defined as adjacent in the same row, adjacent in the same column, or located within the same merged area; for image object areas, adjacency can be defined as the distance between the center of the bounding boxes not exceeding a preset threshold or the existence of overlapping bounding boxes. The directly hit areas and adjacent areas are packaged into a joint extraction result to preserve contextual information.
[0110] Multimodal data comes from different sources, and the naming and time representation methods are often inconsistent. Therefore, before establishing cross-modal correspondence, it is necessary to perform object identifier normalization and time identifier normalization.
[0111] Object identifier normalization involves three steps:
[0112] The first step is to use an object dictionary to map standard names, aliases, abbreviations, and historical names to a unified object identifier;
[0113] The second step is to use a number mapping table to convert the equipment number, work order number, and component number into a unified object identifier.
[0114] The third step is to use semantic similarity matching to disambiguate new expressions that do not appear in the dictionary. When the semantic similarity reaches a preset threshold, it is mapped to the target object identifier.
[0115] Time stamp normalization involves four steps:
[0116] The first step is to parse the explicit date and explicit time;
[0117] The second step is to fill in the missing year, month, day, hour, minute, and second;
[0118] The third step is to perform a unified time zone conversion;
[0119] The fourth step is to standardize the processing results into standard timestamps or standard time interval numbers. For times not directly written in the text, file creation time, image capture time, log write time, or external business transaction time can be used to complete the process.
[0120] After normalization, examine the extraction results across different modalities. If the object identifiers after object normalization are consistent, and the time identifiers after time normalization fall within the same anomalous time interval, then establish a cross-modal correspondence between the corresponding extraction results. The cross-modal correspondence record must include at least the source record number, target record number, object identifier, time identifier, and a combination of modal identifiers.
[0121] In one implementation, to ensure the accuracy of the correspondence establishment, in addition to ensuring that the object identifiers are consistent and the time identifiers fall within the same abnormal time interval, a semantic consistency check is added. The semantic consistency check is completed by comparing the content summary vectors of different extraction results. Only when the similarity of the content summary vectors is not lower than a preset threshold is the cross-modal correspondence written.
[0122] The document layout and its hierarchical relationships, table cell and header dependencies, image object regions and their spatial relationships, and cross-modal correspondences are all stored together to form a parsed object. A parsed object must include at least the following fields: object identifier, time identifier, modality identifier, source file identifier, source location, extracted content, structural relationship type, and cross-modal correspondence pointer. The source location is represented by page number and region coordinates in a document, worksheet name and cell coordinates in a table, and image number and bounding box coordinates in an image. The structural relationship type corresponds to the layout hierarchy in a document, the header dependency in a table, and the spatial location in an image.
[0123] Using the above recording method, subsequent steps can directly filter and parse the object to be parsed based on the user query and the current abnormal status, without having to return to the original file to perform repeated location.
[0124] Taking device A as an example, the real-time data stream shows that the vibration value of device A continuously increased between 15:23:10 and 15:24:20, and the correlation with the bearing temperature channel was significantly enhanced. The anomaly detection module outputs the anomaly category as bearing wear, the anomaly level as level three, the associated object as the bearing assembly of device A, and the anomaly time interval as 15:23:10 to 15:24:20.
[0125] Based on the above results, anomaly correlation structure extraction was performed: The bearing disassembly / assembly steps and bearing fault judgment criteria were located in the PDF file corresponding to the maintenance manual, and the page layout hierarchy between the title and the main text was extracted; the bearing inventory unit corresponding to equipment A was located in the Excel spreadsheet corresponding to the spare parts ledger, and the header dependency relationship between the equipment number header, spare parts model header, and inventory value unit was extracted; the bearing area and temperature label area were located in the inspection images, and the spatial position relationship of the temperature label area to the right of the bearing area was extracted. Since the anomaly level was three, the anomaly time interval was extended forward by 30 minutes and backward by 30 minutes, and further, document layout areas, table units, and image object areas adjacent to the above-mentioned content within the extended time interval were jointly extracted. Object names and time representations in different modalities were normalized, cross-modal correspondences were established, and all results were associated and stored as objects to be parsed.
[0126] S2: Based on the user query and the current abnormal state, determine the target parsing object from the objects to be parsed and perform parsing to obtain the region-level semantic representation and the document-level semantic representation.
[0127] In one embodiment, step S2 includes:
[0128] Encode the user query into a query vector, the current anomaly state into an anomaly state vector, and the content summary, structural attributes, and metadata of the object to be parsed into an object vector. Calculate the following formula: The target parsing score of the object to be parsed :
[0129] ;
[0130] In the formula, For user queries and the The semantic relevance of the object to be parsed can optionally be achieved by using the cosine similarity between the query vector and the object vector. The base values are then mapped to the interval [0,1] through linear normalization; For the current abnormal state and the first The anomaly relevance of an object to be parsed can be calculated based on the matching results between the anomaly category, associated objects, and anomaly time interval and the corresponding metadata of the object to be parsed. When the object identifiers match and the time identifiers fall within the anomaly time interval, the anomaly relevance is considered to be true. Take the higher value; For users to query the current abnormal status of the first The coupling relevance of each object to be parsed is used to improve the sorting position of objects to be parsed that simultaneously satisfy query relevance and anomaly relevance. For the exception category and the first The information type matching degree of the object to be parsed is determined by the modality and region type of the object. For example, the fault description paragraph in the document layout area, the parameter record unit in the table cell, and the fault location area in the image are mapped to different information types. A matching rule table between anomaly categories and information types is pre-established, and the results are output according to the rule table. For example, when the anomaly category is mechanical wear, the maintenance procedure information and spare parts parameter information have higher matching values. For the first The completeness of the cross-modal correspondence of a parsing object is determined by whether the parsing object has established a correspondence with at least two objects in the document layout area, table cell, and image object area. Take the higher value; if only a single-modal record exists and there is no cross-modal correspondence, then... Take the lower value; For the first The structural completeness of the object to be parsed is determined by checking whether the object simultaneously contains the source location, structural relationship type, object identifier, and time identifier. The ratio of the number of included fields to the total number of predefined fields is the structural completeness. ; For the first The boundary reliability of the object to be parsed, for a document layout area. It can be determined jointly by the layout detection confidence score and the character recognition confidence score; for table cells, It can be determined jointly by cell location reliability and header traceability stability; for image object regions, It can be directly determined by the object detection confidence score; For the first The normalized parsing cost of an object to be parsed is estimated based on the region area, text length, table complexity, or image resolution, and the estimated result is normalized to the interval [0,1]. The higher the parsing cost, the better. The larger; Analyze the scoring weights for the target, and satisfy the following conditions: In this embodiment, .
[0131] According to the target analysis score The objects to be parsed are sorted, and those with a target parsing score greater than a preset threshold or ranking within the top predetermined number are selected as the target parsing objects. If the number of parsing objects with a score greater than the preset threshold is insufficient, additional parsing objects ranking within the top predetermined number are selected. Preferably, the preset threshold is set to 0.65, and the predetermined number is set to 10. This process ensures that parsing resources are preferentially allocated to parsing objects that are more relevant to the user query and the current anomaly state, have a more complete structure, and have a more moderate parsing cost.
[0132] Fine-grained parsing is performed on the target parsing object to obtain the corresponding region-level semantic representation. In one implementation, fine-grained parsing can be achieved through a text region encoding branch, a table unit encoding branch, an image object region encoding branch, and a unified projection layer. The text region encoding branch consists of a word embedding layer and six Transformer encoding layers concatenated together, used to process text content within the document layout area; the table unit encoding branch consists of a unit content encoding layer, a header relationship encoding layer, and two Transformer encoding layers concatenated together, used to process table units and header dependencies; the image object region encoding branch consists of four concatenated convolutional layers and two Transformer encoding layers concatenated together, used to process image object regions and their spatial relationships. The outputs of the three encoding branches are connected to the same linear projection layer, uniformly mapped to a 768-dimensional vector, serving as a region-level semantic representation of the corresponding target parsing object. The above can be achieved using offline training and online inference. Training samples consist of historical document regions, table units, image object regions, and corresponding semantic labels. During training, the text region encoding branch, table unit encoding branch, and image object region encoding branch each receive samples from their respective modalities. A unified projection layer performs alignment training on the three outputs, ensuring that cross-modal samples of the same object within the same time interval maintain high similarity in the vector space, while samples of different objects or different time intervals maintain low similarity. After training, the model parameters are fixed, and the online phase directly outputs region-level semantic representations. .
[0133] Multiple region-level semantic representations with the same source document identifier, the same object identifier, and time identifier falling within the same anomalous time interval are aggregated to obtain document-level semantic representations. The aggregation formula is as follows: ;and ;
[0134] In the formula, For the first The document-level semantic representation of each target parsing object; This refers to the set of target parsing objects that have the same source document identifier, the same object identifier, and whose time identifiers fall within the same abnormal time interval. For the first The target parsing objects are in the target parsing object set Aggregate weights in; For the first The region-level semantic representation corresponding to each target parsing object; For the first The time consistency of the target parsing objects, in this embodiment, if the first target parsing object... If the timestamp of a target parsing object falls within the original abnormal time range, then Take 1; if the first If the timestamp of a target parsed object falls within the expanded abnormal time interval, then The time difference from the original abnormal time interval boundary is linearly decreased to the interval [0,1); if the first... If the time signature of a target parsing object exceeds the expanded abnormal time interval, it will not be included in the set. ; For the target parsing object set The target parsing object number in the middle; The aggregate weight parameters satisfy the following conditions: In this embodiment .
[0135] Taking the manufacturing equipment operation and maintenance scenario as an example, after the objects to be parsed are formed in step S1, the user inputs "the basis for handling the abnormal bearing of equipment A". First, the target parsing score of each object to be parsed is calculated. The document page area in the maintenance manual that records the bearing disassembly and assembly steps receives a high target parsing score because it is highly relevant to the user's query, matches the current abnormal category, has complete structural fields, and clear boundaries. The table cell in the spare parts ledger that records "6205 bearing inventory" also receives a high target parsing score because it is consistent with the associated object and has a low parsing cost. The image object area corresponding to the bearing part in the inspection image is also included in the target parsing object set because a cross-modal correspondence has been established with the equipment number and alarm time period. The regional semantic representations of the above target parsing objects are output respectively, and they are aggregated within the same source document, the same object identifier, and the same abnormal time interval to form the corresponding document-level semantic representation.
[0136] S3: Construct a dynamic knowledge base based on regional-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a regional evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer, and updates the hierarchical anomaly prototype layer according to the current anomaly state.
[0137] Specifically, regional semantic representations and structured fields are encapsulated into regional evidence entries, and the first in the regional evidence layer... Evidence entries for each region Recorded as: ,in, Indicates the first Regional-level semantic representation; Represents the object identifier; Indicates time stamp; Indicates modal identifier; Indicates structural relationship identifiers; Indicates the source document identifier; Indicates the source location. Structural relationship identifiers are used to record one of the following: page hierarchy, header dependency, or spatial location relationship. In a document, the source location is represented by page number and area coordinates; in a table, by worksheet name and cell coordinates; and in an image, by image number and bounding box coordinates.
[0138] In the semantic vector layer, a one-to-one mapping is established between region-level semantic representations and region evidence entries, while an aggregation mapping is established between document-level semantic representations and one or more corresponding region evidence entries. Each region-level semantic vector record is denoted as: ,in, Indicates a region-specific evidence entry The reference pointer. Each document-level semantic vector record is denoted as: ,in, Indicates the first Document-level semantic representation; Indicates and The corresponding set of evidence items for the region; This represents a pointer to the set of evidence entries for that region. Region-level semantic vector records are used for region-level retrieval, while document-level semantic vector records are used for document-level recall.
[0139] In the cross-modal knowledge graph layer, the graph construction unit builds a graph structure with typed nodes and typed edges based on cross-modal correspondences, object identifiers, time identifiers, and abnormal time intervals. (Cross-modal knowledge graph) Recorded as: ,in, Represents a set of nodes. The node set represents the set of edges. The node set includes at least text entity nodes, table field nodes, image object nodes, area evidence nodes, abnormal event nodes, and action nodes. The edge set includes at least object-related edges, time-related edges, structural-related edges, evidence-supporting edges, and action-related edges. Text entity nodes represent the device names, component names, fault names, parameter names, and operational terms parsed from the document layout area; table field nodes represent header fields and data fields; image object nodes represent device objects, component objects, damaged objects, and marked objects in images; area evidence nodes represent area evidence entries; abnormal event nodes represent event entities corresponding to abnormality categories, abnormality levels, associated objects, and abnormal time intervals; and action nodes represent inspection actions, maintenance actions, shutdown actions, and replacement actions.
[0140] To facilitate the construction of the hierarchical anomaly prototype layer, graph sub-paths corresponding to the current anomaly state are further extracted. In this embodiment, the graph sub-path represents a node connection sequence that starts from the anomaly event node, passes through object association edges, time association edges, or structural association edges, and reaches the region evidence node or the action node. The path of a bar graph is denoted as:
[0141] ,
[0142] In the formula, to Represents nodes on the path. to This indicates the edge type between adjacent nodes.
[0143] In the hierarchical exception prototype layer, instance-level exception prototypes are generated first. An instance-level exception prototype is denoted as:
[0144] ;
[0145] In the formula, Indicates the exception category; Indicates the degree of abnormality; Indicates the associated object; Indicates an abnormal time interval; This represents the set of evidence items for the corresponding region. This represents the corresponding document-level semantic representation; Indicates the corresponding graph sub-path; This represents the prototype representation of an instance-level exception.
[0146] Instance-level exception prototype representations are generated using the following formula: ,in, , This represents the average vector of the semantic representation at the regional level. Represents the current abnormal state vector; Represents the graph subpath encoding vector; For the combined weights, satisfying .
[0147] The semantic-level anomaly prototype consists of multiple instance-level anomaly prototypes with a semantic consistency of no less than a preset threshold. Semantic consistency is calculated using the cosine similarity between the instance-level anomaly prototype representations. ,in, Represents instance-level exception prototypes With instance-level exception prototype The semantic consistency between them. The semantic-level exception prototype is denoted as: ,in, Represents a collection of instance-level exception prototypes; This represents the semantic-level anomaly prototype center representation; This indicates the anomaly category identifier. The semantic-level anomaly prototype center representation is calculated using the following formula: .
[0148] A category-level exception prototype consists of one or more semantic-level exception prototypes within the same exception category. Each category-level anomaly prototype is denoted as: ,in, Indicates the anomaly category identifier; This indicates that it belongs to the abnormal category. A collection of semantic-level exception prototypes; This represents the prototype center representation of a category-level anomaly. The prototype center representation of a category-level anomaly is calculated using the following formula: .
[0149] Through the above construction method, the regional evidence layer is used to store traceable fine-grained evidence, the semantic vector layer is used to store the vector index required for retrieval, the cross-modal knowledge graph layer is used to store the explicit relationships between objects, time, structure and disposal, and the hierarchical anomaly prototype layer is used to store anomaly knowledge organized by anomaly category, semantic pattern and specific event.
[0150] In another specific implementation, updating the hierarchical exception prototype layer according to the current exception state includes:
[0151] Based on the current anomalous state and the corresponding region-level semantic representation, document-level semantic representation, and graph sub-path, a current anomalous event representation is generated. The "current anomalous event representation" is a vector obtained by fusing the region-level semantic representation, document-level semantic representation, anomalous state vector, and graph sub-path encoding corresponding to the current anomalous state. It is used to represent the overall semantics of the current anomalous event and is denoted as: ,in In the formula, This represents the average vector of the region-level semantic representation corresponding to the current abnormal state; This represents the document-level semantic representation corresponding to the current abnormal state; Represents the current abnormal state vector; This represents the graph subpath encoding vector corresponding to the current abnormal state; This represents the set of evidence entries for the region corresponding to the current abnormal state. To integrate weights, satisfy .
[0152] For each existing instance-level exception prototype, calculate the prototype matching score, the first... The prototype matching score between existing instance-level exception prototypes and the current exception event representation. Recorded as: ;
[0153] In the formula, For anomaly category consistency, when the current anomaly category is consistent with the first anomaly category... When multiple existing instance-level exception prototype records have the same exception category Set to 1 when the current exception category is the same as the first exception. When the exception category of an existing instance-level exception prototype record is in a preset set of associated categories Take a correlation value between 0 and 1; when there is no correlation, Set to 0; For the anomaly degree matching degree, if the anomaly degree is represented by discrete levels, then the current anomaly degree matches the level of the [number]th [level]. When the degree of exception is the same for all existing instance-level exception prototype records Set to 1, and decrease by a preset ratio when there is a difference of one level; For the matching degree of associated objects, when the associated object identifiers are the same, When the value is 1, the associated objects are in the same object association cluster. Take a correlation value between 0 and 1; if there is no correlation, Set to 0; The overlap of abnormal intervals can be calculated using the following formula: , Indicates the current abnormal time interval. Indicates the first An exception time interval that is already recorded in an instance-level exception prototype. Indicates the length of time. The value range is [0,1]; For the current abnormal event characterization and the first The semantic consistency between existing instance-level anomaly prototypes can be calculated using cosine similarity. Match weights to the prototype, set them fixed before deployment, and ensure that they meet the following requirements. Preferred .
[0154] The existing instance-level exception prototype with the highest prototype matching score is selected as the candidate matching prototype.
[0155] When the prototype matching score of a candidate matching prototype is not less than the first preset threshold, the candidate matching prototype is updated, with an update coefficient of 0.8. The regional evidence entries, document-level semantic representations, and graph sub-paths corresponding to the current abnormal state are appended to the candidate matching prototype association records, and the abnormal time interval is updated to the union of the original abnormal time interval and the current abnormal time interval. The central representations of the semantic-level abnormal prototype and the category-level abnormal prototype are corrected according to their association relationships. The central representation of the semantic-level abnormal prototype is corrected to the average value of the representations of the instance-level abnormal prototypes, and the central representation of the category-level abnormal prototype is corrected to the average value of the central representations of the semantic-level abnormal prototypes.
[0156] When the prototype matching score of a candidate matching prototype is less than a first preset threshold but not less than a second preset threshold, a new instance-level anomaly prototype is created under the semantic-level anomaly prototype to which the candidate matching prototype belongs. The new instance-level anomaly prototype directly uses the current anomaly event representation as its prototype representation and associates it with the current anomaly category, current anomaly degree, current associated object, current anomaly time interval, current region evidence item set, current document-level semantic representation, and current graph sub-path. After creation, the central representation of the semantic-level anomaly prototype is recalculated, and the central representation of the corresponding category-level anomaly prototype is further corrected. The above processing method is applicable to situations where the anomaly categories are consistent, the semantic patterns are similar, but the conditions for updating the same instance-level anomaly prototype have not yet been met.
[0157] When there is no existing instance-level exception prototype with a prototype matching score not less than the second preset threshold, a new instance-level exception prototype and a new semantic-level exception prototype are created based on the current exception event representation, and the new semantic-level exception prototype is attached to the category-level exception prototype of the corresponding exception category; when the category-level exception prototype of the corresponding exception category does not exist, a new category-level exception prototype is created.
[0158] Preferably, the first preset threshold is set to 0.80, and the second preset threshold is set to 0.55. Through the above two-level threshold settings, the hierarchical exception prototype layer can distinguish between different scenarios such as reusing and updating existing instances, adding new instances under existing semantic patterns, and adding new semantic patterns. This avoids forcibly merging exception events with large semantic differences into the same instance-level exception prototype, and also avoids splitting highly consistent repetitive exception events into multiple independent prototypes.
[0159] Taking the manufacturing equipment operation and maintenance scenario as an example, the current abnormal state corresponds to "bearing wear, level 3, equipment A bearing assembly, 15:23:10 to 15:24:20". First, the current abnormal event representation is generated based on the regional-level semantic representation, document-level semantic representation, and graph sub-path corresponding to the current abnormal state. Then, the prototype matching score is calculated one by one with the existing instance-level abnormal prototypes. If an existing instance-level abnormal prototype also corresponds to equipment A bearing wear and the time interval, handling path, and evidence pattern are highly consistent, the prototype matching score is not less than the first preset threshold, and the corresponding instance-level abnormal prototype is updated. If an existing instance-level abnormal prototype corresponds to the same abnormal category but the associated objects or time intervals are different, resulting in the prototype matching score being between the first and second preset thresholds, then an instance-level abnormal prototype is added under the corresponding semantic-level abnormal prototype. If there is no existing instance-level abnormal prototype in the existing hierarchical abnormal prototype layer that matches the current abnormal event representation to the second preset threshold, then a new semantic-level abnormal prototype is created and attached to the category-level abnormal prototype corresponding to "bearing wear". If the category-level anomaly prototype corresponding to "bearing wear" does not already exist, a new category-level anomaly prototype will be created.
[0160] S4: Based on the user query, the current abnormal status, and historical feedback, determine the retrieval execution order corresponding to the current request, and perform abnormal prototype retrieval in the hierarchical abnormal prototype layer, multimodal semantic retrieval in the semantic vector layer and regional evidence layer, and path reasoning in the cross-modal knowledge graph layer to obtain candidate evidence.
[0161] In one implementation, the retrieval execution order corresponding to the current request is determined based on the user query, the current anomaly status, and historical feedback, and this is implemented by the sequence decision module. The sequence decision module is located within a dynamic knowledge base and includes at least a strategy experience storage unit, a feature encoding unit, a sequence matching unit, and a sequence adjustment unit. The strategy experience storage unit stores historical strategy experience records; the feature encoding unit generates current user query features and current anomaly status features; the sequence matching unit selects candidate historical retrieval execution orders from the historical strategy experience records; and the sequence adjustment unit expands, prunes, or rearranges the candidate historical retrieval execution orders according to the current anomaly level to obtain the retrieval execution order corresponding to the current request.
[0162] In this embodiment, the strategy experience layer represents a data layer used to store historical retrieval decision-making experience. Each strategy experience record includes at least historical user query characteristics, historical abnormal state characteristics, historical retrieval execution order, and normalized feedback reward value. Strategy Experience Record Recorded as: ,in, Indicates the first The historical user query feature vector of each strategy experience record; Indicates the first The feature vector of historical abnormal states recorded in the strategy experience record; Indicates the first The historical retrieval execution order corresponding to each strategy experience record; Indicates the first The normalized feedback reward value corresponding to each strategy experience record. Historical retrieval execution order. This includes at least the execution order of anomaly prototype retrieval, multimodal semantic retrieval, and path reasoning. Normalized feedback reward value. The value range is set to [0,1].
[0163] The feature encoding unit receives the current user query and the current anomaly status. The current user query features are output by the query encoding sub-network, which consists of a word embedding layer and four Transformer encoding layers concatenated, outputting a 256-dimensional query feature vector. The current anomalous state features are output by the state encoding sub-network. This sub-network receives discrete and continuous features corresponding to the anomalous category, anomalous degree, associated objects, and anomalous time interval. After passing through an embedding layer and two fully connected layers, it outputs a 256-dimensional anomalous state feature vector. The parameters of the query encoding subnetwork and the state encoding subnetwork can be determined using offline training. The training samples consist of historical user queries, historical abnormal states, and subsequent reward results.
[0164] The sequential matching unit calculates the matching relationship between the current feature and historical strategy experience records one by one. In one implementation, the current user queries the feature... With the Historical user query characteristics in strategy experience records Query similarity between Cosine similarity is used for calculation; current abnormal state features With the Historical Abnormal State Characteristics in the Policy Experience Record State similarity between Cosine similarity is used for calculation. Query similarity. Similarity to state After linear mapping, all values fall within the interval [0,1]. Then, the th... The sequential matching score corresponding to each strategy experience record :
[0165] ;
[0166] In the formula, For sequential matching weights, the following conditions must be met: Preferably, .
[0167] Sequential matching units are ranked according to their sequential matching scores. All historical strategy experience records are sorted, and historical retrieval execution orders with a sequence matching score greater than a first preset threshold or ranking within a first predetermined number are selected as candidate historical retrieval execution orders. In one embodiment, the first preset threshold is set to 0.70, and the first predetermined number is set to 5. If the number of historical retrieval execution orders that meet the first preset threshold is less than the first predetermined number, the number is supplemented according to the sorting results.
[0168] The order adjustment unit performs expansion, pruning, or rearrangement on the candidate historical retrieval based on the current anomaly level. In this specification, expansion means adding pre-execution of anomaly prototype retrieval or adding supplementary execution of path reasoning; pruning means reducing the number of branches in multimodal semantic retrieval or reducing the expansion depth of path reasoning; rearrangement means adjusting the execution order of anomaly prototype retrieval, multimodal semantic retrieval, and path reasoning.
[0169] The anomaly severity is divided into three ranges: low, medium, and high. When the current anomaly severity is in the high range, anomaly prototype retrieval is performed first, followed by multimodal semantic retrieval, and finally path reasoning. An anomaly prototype supplementary retrieval is then performed after path reasoning. When the current anomaly severity is in the medium range, anomaly prototype retrieval and multimodal semantic retrieval are performed first, followed by path reasoning. When the current anomaly severity is in the low range, multimodal semantic retrieval is performed first, a simplified mode is used for anomaly prototype retrieval, and a shallow expansion mode is used for path reasoning.
[0170] The reordering is determined jointly by the rule table and the reward ranking. The rule table pre-stores the mapping relationship between "anomaly category - priority retrieval order". For example, when the anomaly category is mechanical wear, the priority retrieval order is set as: anomaly prototype retrieval - multimodal semantic retrieval - path reasoning; when the anomaly category is parameter drift, the priority retrieval order is set as: multimodal semantic retrieval - anomaly prototype retrieval - path reasoning. When the execution order of candidate historical retrievals is inconsistent with the output order of the rule table, the one with the higher normalized feedback reward value is given priority; if the difference in normalized feedback reward values does not exceed the preset tolerance, the output order of the rule table is adopted.
[0171] The current user query is "basis for handling abnormalities in bearing A of equipment," and the current abnormality status corresponds to "bearing wear, level three, bearing assembly of equipment A, current alarm period." The sequential decision module recalls multiple historical strategy experience records similar to "fault Q&A" and "mechanical wear" from the strategy experience layer. Since the current abnormality level is in the high-level range, the sequential adjustment unit performs sequential execution expansion on the candidate historical retrieval, so that the abnormality prototype retrieval is performed first, followed by multimodal semantic retrieval, and path reasoning is performed as the third stage. After the path reasoning is completed, a supplementary retrieval based on the graph path results is allowed.
[0172] Anomaly prototype retrieval is performed at the hierarchical anomaly prototype level. During prototype retrieval, the current anomaly state is used as a retrieval constraint. The consistency between the current anomaly state and each instance-level, semantic-level, and category-level anomaly prototype in terms of anomaly category, anomaly severity, associated object, and anomaly time interval is compared. Furthermore, the semantic similarity between the current anomaly event representation and the representations of each anomaly prototype is considered to rank the anomaly prototypes. After ranking, anomaly prototypes with a prototype retrieval score greater than a second preset threshold or ranking within the top second predetermined number are selected as the anomaly prototype retrieval results. Optionally, the second preset threshold is set to 0.65, and the second predetermined number is set to 8. The anomaly prototype retrieval results retain at least the anomaly prototype identifier, hierarchical type, prototype representation, prototype retrieval score, and associated object identifier for subsequent prototype guidance and generation.
[0173] The semantic retrieval unit performs multimodal semantic retrieval at the semantic vector layer and the regional evidence layer. During semantic retrieval, the user query is used as the primary retrieval condition, and the current anomaly state is used as an auxiliary constraint to search for regional-level and document-level semantic vector records. The retrieval process simultaneously considers the semantic relevance between the user query and the semantic vector record, the anomaly relevance between the current anomaly state and the semantic vector record, the structural completeness of the corresponding regional evidence entry, and the information type fit between the modality type and the current anomaly category. After sorting, semantic vector records with a semantic retrieval score greater than a third preset threshold or ranking within the top three predetermined number are selected as multimodal semantic retrieval results. Optionally, the third preset threshold is set to 0.60, and the third predetermined number is set to 20. Each multimodal semantic retrieval result retains its corresponding regional evidence entry, source location, object identifier, time identifier, and modality identifier.
[0174] Based on the anomaly prototype retrieval results, prototype guidance vectors corresponding to candidate evidence are generated. Specifically, firstly, a predetermined number of anomaly prototypes ranked in the top four by prototype retrieval scores are selected from the anomaly prototype retrieval results. Then, the corresponding representations are called according to the hierarchical type of the anomaly prototype: when the anomaly prototype is an instance-level anomaly prototype, the instance-level anomaly prototype representation is called; when the anomaly prototype is a semantic-level anomaly prototype, the semantic-level anomaly prototype center representation is called; when the anomaly prototype is a category-level anomaly prototype, the category-level anomaly prototype center representation is called. Subsequently, the above representations are weighted according to their respective prototype retrieval scores to obtain the prototype guidance vector corresponding to the current candidate evidence. Optionally, the fourth predetermined number is set to 5. To ensure the effectiveness of prototype guidance, only anomaly prototypes with consistent object identifiers, consistent anomaly categories, or time identifiers falling within the same anomaly time interval participate in the prototype guidance generation of the current candidate evidence.
[0175] Candidate evidence vectors and prototype guidance vectors are fused. Candidate evidence vectors are obtained from the regional evidence entries, semantic vector records, and metadata encoding corresponding to the multimodal semantic retrieval results, representing the semantic content, source location, and structural attributes of the candidate evidence itself. Prototype guidance vectors represent the historical anomaly patterns most closely related to the current anomaly state. The purpose of denoising fusion is to suppress candidate evidence components that are irrelevant or have low relevance to the current anomaly, while retaining evidence components consistent with the anomaly prototype, thereby improving the purity and concentration of subsequent path reasoning inputs.
[0176] In one implementation, the denoising fusion unit employs a multi-expert gating structure. For the first... The output vector of each candidate piece of evidence, after denoising and fusion, is calculated using the following formula: ;
[0177] in, ;
[0178] In the formula, Indicates the first The output vector of candidate evidence after denoising and fusion; Indicates the first One candidate evidence vector; Indicates the first The prototype guiding vector corresponding to each candidate piece of evidence; This represents the number of denoising expert functions; and This refers to the index of the denoising expert function; For the first A noise reduction expert function; For the first The denoising expert function for the first... Gating weights for each candidate piece of evidence; This indicates vector concatenation; and This represents the gating parameter vector corresponding to the denoising expert function;
[0179] In one implementation, the number of denoising expert functions The parameter set is 4, where two denoising expert functions are used to suppress noise components inconsistent with the current anomaly category, and the other two are used to enhance evidence components consistent with the prototype guidance vector. The gating parameter vector is obtained through offline training. Training samples consist of a triple of "candidate evidence vector - prototype guidance vector - target evidence label," with the target evidence label generated based on historical retrieval results and human verification results. During training, positive samples are designed to maintain a high similarity to the correct anomaly prototype after denoising and fusion, while negative samples are designed to maintain a low similarity to irrelevant anomaly prototypes after denoising and fusion. After training, the denoising expert function parameters and gating parameter vector are fixed, and inference is directly executed in the online phase.
[0180] Candidate evidence is re-ranked based on the denoising fusion results. The re-ranking considers three factors: first, the consistency between the denoising fusion result and the prototype guiding vector; second, the ranking position of the candidate evidence in the preceding multimodal semantic retrieval; and third, the boundary credibility and structural completeness of the evidence entries corresponding to the candidate evidence in the relevant region. If a candidate evidence performs well in all three aspects, it is given priority as input for path reasoning. Candidate evidence with a re-ranking score greater than a preset condition or ranked within a predetermined number is selected as input for path reasoning; this predetermined number can be set to 10. If multiple candidate pieces of evidence correspond to the same object identifier, the same time identifier, and have overlapping source locations, the one with the higher ranking is retained, and the remaining candidate evidence is used as supplementary evidence.
[0181] For example, the current abnormal state corresponds to "bearing wear of equipment A, level 3 abnormality, current alarm period". The abnormal prototype retrieval results recall multiple instance-level and semantic-level abnormal prototypes related to bearing wear. The multimodal semantic retrieval results recall disassembly and assembly steps in the maintenance manual, inventory units in the spare parts ledger, and bearing parts in inspection images. Based on the prototype retrieval results, prototype guidance vectors corresponding to the above candidate evidence are generated, and then each candidate evidence vector is denoised and fused with its corresponding prototype guidance vector. After denoising and fusion, the weights of irrelevant explanatory paragraphs inconsistent with the current abnormal pattern, historical unrelated work order entries, and non-fault area images are reduced, while the weights of maintenance steps, inventory information, and fault location images directly related to the bearing abnormality are retained. Finally, a set of high-confidence candidate evidence is output for path reasoning.
[0182] The above path reasoning input serves as the initial evidence in the cross-modal knowledge graph layer. The path reasoning module performs initial node determination, constrained expansion, path filtering, and evidence aggregation in the cross-modal knowledge graph layer to obtain the path reasoning result, which is then merged with the previous retrieval result to form candidate evidence.
[0183] The path reasoning module includes at least a starting node determination unit, a path expansion unit, a path filtering unit, and an evidence aggregation unit. The starting node determination unit maps the path reasoning input to a starting node in the cross-modal knowledge graph layer; the path expansion unit performs constrained expansion along the edge types in the graph; the path filtering unit filters valid reasoning paths from multiple candidate paths obtained from the expansion; and the evidence aggregation unit extracts evidence content from the valid reasoning paths and forms the path reasoning result.
[0184] When the starting node is determined, the regional evidence entries corresponding to the path reasoning input are preferentially mapped to regional evidence nodes. If a regional evidence node can directly point to a text entity node, table field node, or image object node, then that text entity node, table field node, or image object node is also used as an auxiliary starting node. If a regional evidence entry fails to form a valid mapping, it degenerates into directly using a text entity node, table field node, or image object node as the starting node. The target node is determined by the abnormal event node or the action node corresponding to the current abnormal state. If the user query is biased towards cause analysis, the abnormal event node is preferentially used as the target node; if the user query is biased towards action suggestions, the action node is preferentially used as the target node.
[0185] The path expansion unit starts from the initial node and expands along object association edges, temporal association edges, structural association edges, evidence support edges, and disposition association edges in the cross-modal knowledge graph layer. The expansion process is not an unconditional traversal but is constrained by both the current anomalous state and the path reasoning input. In one implementation, only next-hop nodes that simultaneously meet the following conditions are allowed to enter the candidate path:
[0186] The object identifier corresponding to the next hop node is consistent with the associated object in the current abnormal state, or is within a preset object association cluster with the associated object;
[0187] The time identifier corresponding to the next hop node intersects with, contains, or falls within the preset extended time window of the current abnormal time interval;
[0188] The structural relationship between the next hop node and the previous node on the current path does not conflict with the existing page hierarchy, header dependency, or spatial position relationship in the initial evidence;
[0189] If the next hop node is a handling action node, then the handling action node must have a preset handling association relationship with the current exception category.
[0190] Path expansion employs a bundle search approach. After each layer of expansion, only a limited number of candidate paths with high scores are retained to control the number of paths and computational complexity. The bundle width can be set based on the current anomaly level: a larger bundle width is used when the anomaly level is high to retain more potential critical paths; a smaller bundle width is used when the anomaly level is low to reduce redundant expansion. The bundle width can be configured between 5 and 20. To avoid introducing noise due to excessively long paths, an upper limit is set for path length. The upper limit for path length is set to 4 hops; paths exceeding this limit are no longer expanded.
[0191] The path filtering unit filters the expanded candidate paths. During filtering, it comprehensively considers the consistency of objects, the consistency of time, the coherence of structural relationships, and the relevance of actions within the path. Object consistency determines whether the path revolves around the same associated object or the same cluster of associated objects; time consistency determines whether the time markers corresponding to each node in the path are compatible with the current anomaly time interval; structural relationship coherence determines whether the page hierarchy, header dependencies, or spatial relationships between nodes in the path can form a reasonable chain of evidence; and action relevance determines whether the action node at the path's endpoint matches the current anomaly category. If a candidate path meets all four preset requirements, it is considered a valid reasoning path. Furthermore, a valid reasoning path must contain at least one regional evidence node, and this regional evidence node must be traceable back to its specific source location in the original document, table, or image to ensure that the subsequently generated results have a citationable evidentiary basis.
[0192] The evidence aggregation unit aggregates the node content along the valid reasoning path. During aggregation, the following node content is prioritized for retention: regional evidence nodes directly corresponding to the associated objects in the current abnormal state; abnormal event nodes directly overlapping with the current abnormal time interval; and action nodes corresponding to the current abnormal category. For multiple nodes with similar content or overlapping source locations on the same path, node content with a clearer source location, more complete structural relationships, and a higher degree of consistency with the current abnormal state is retained. For node content from different modalities but with consistent object and time identifiers, it is saved as a multimodal evidence combination within the same evidence chain. The path reasoning result includes at least the path identifier, the node identifiers involved, the edge types involved, the path source, the regional evidence entries corresponding to the path, and the abnormal event node or action node corresponding to the path endpoint.
[0193] When candidate evidence is formed, the path reasoning results, candidate evidence whose re-ranking scores meet preset conditions, and the previous-level anomaly prototype retrieval results are merged. During merging, if results from different sources correspond to the same object identifier, the same time identifier, and overlapping source locations, the one with higher consistency and clearer source is retained as the primary evidence, and the remaining results are added as supplementary evidence below the primary evidence. If a path reasoning result can simultaneously connect an anomaly event node, a regional evidence node, and a handling action node, it is given priority for retention as high-value candidate evidence.
[0194] For example, the path reasoning input includes three types of evidence: "image area of bearing part of equipment A", "disassembly and assembly steps area in the maintenance manual", and "6205 bearing inventory unit in the spare parts ledger". The starting node determination unit first maps the three types of evidence into image object nodes, area evidence nodes, and table field nodes, and then selects the abnormal event nodes that are consistent with the object identifiers of the three types of nodes and whose time identifiers fall within the current abnormal time interval as the primary target nodes. The path expansion unit first expands along the object association edge and time association edge to the "abnormal event node of bearing wear of equipment A", and then expands along the disposal association edge to disposal action nodes such as "replacing 6205 bearing" and "checking lubrication status". The path filtering unit retains paths that are consistent in object, compatible in time, do not conflict in structural relationships, and match in disposal actions. The evidence collection unit extracts maintenance steps, spare parts information, and disposal actions based on this, forms the path reasoning result, and merges it with the previous retrieval results to form candidate evidence for subsequent fusion ranking and model generation.
[0195] S5: Merge and rank the candidate evidence to form an evidence set, and use the evidence set as context input to generate a model, outputting an answer, analysis report, or handling recommendation.
[0196] In one implementation, after candidate evidence is formed, the evidence fusion generation module performs deduplication and merging, conflict identification, fusion sorting, quota selection, and context assembly on the candidate evidence, and inputs the assembly result into the generation model. The evidence fusion generation module includes at least an evidence clustering unit, a conflict determination unit, a sorting and selection unit, a context assembly unit, and a generation control unit.
[0197] In this embodiment, a candidate evidence cluster represents a set of evidence composed of candidate evidence whose object identifiers, time identifiers, and source locations are close to each other; the main evidence represents the candidate evidence with the highest priority in the candidate evidence cluster, used to represent the candidate evidence cluster in subsequent generation; supplementary evidence represents candidate evidence attached to the main evidence, used to supplement source, modality, or detailed information; conflict marker represents a marker set when candidate evidence has the same object identifier and compatible time identifiers, but the content conclusions are inconsistent; and task identifier represents the output task type of the generation model, used to indicate that the generation model outputs one of the following: a response, an analysis report, or a disposal suggestion.
[0198] The evidence clustering unit clusters all candidate evidence. During clustering, the system uses object identifier as the first constraint, time identifier as the second constraint, and source location and semantic similarity as the third constraint. Only candidate evidence that simultaneously meets the following conditions is merged into the same candidate evidence cluster:
[0199] The object identifiers are the same;
[0200] Time identifier compatibility includes time intervals intersecting, time intervals containing each other, or time intervals not exceeding a preset tolerance.
[0201] The sources overlap, or the semantic similarity of the content is not less than a fifth preset threshold. Semantic similarity can be determined by the cosine similarity between the corresponding vectors of the candidate evidence. Optionally, the fifth preset threshold is set to 0.85.
[0202] For candidate evidence that meets the above conditions, perform deduplication and merging, retaining the one with a clearer source, more complete structure, or higher previous score as the primary candidate, and temporarily storing the remaining candidate evidence as supplementary items.
[0203] The conflict determination unit identifies conflicting evidence within each candidate evidence cluster. If two candidate evidence objects have the same identifier and compatible time identifiers, but one candidate evidence points to "needs replacement" while the other points to "only needs inspection" or "no action required," then the conclusions are deemed inconsistent. For candidate evidence with inconsistent conclusions, a conflict flag is set, and the source location, evidence type, and previous score of both parties are retained. After setting the conflict flag, lower-ranked candidate evidence is not directly deleted; instead, it is retained as a conflict explanation candidate to explicitly indicate the evidence discrepancy during subsequent generation. This method avoids losing contradictory evidence that could affect the judgment result during the candidate evidence formation stage.
[0204] The sorting selection unit performs a fusion sorting of the main evidence in each candidate evidence cluster. During sorting, at least the following factors are considered: user query relevance, current anomaly state consistency, anomaly prototype support, path reasoning support, cross-modal complementarity, source completeness, and conflict flags. User query relevance reflects the semantic proximity between the main evidence and the user query; current anomaly state consistency reflects the consistency between the main evidence and the anomaly category, anomaly degree, associated objects, and anomaly time interval; anomaly prototype support reflects whether the main evidence is supported by the anomaly prototype retrieval results; path reasoning support reflects whether the main evidence appears on a valid reasoning path, or whether it is consistent with the anomaly event node or action node at the end of a valid reasoning path; cross-modal complementarity reflects whether the main evidence can form a complementary evidence chain with other modal evidence; source completeness reflects whether the main evidence carries a clear source location, object identifier, time identifier, and structural relationship; conflict flags are used to reduce the sorting priority of the corresponding main evidence when there is evidence disagreement. In one implementation, a weighted summation method is used to calculate the fusion sorting priority, with each weight preset and fixed before deployment, and the sum being 1. If the main evidence is marked with a conflict, a preset conflict penalty value will be deducted from the original weighted score.
[0205] After determining the fusion and ranking priority of each main piece of evidence, the ranking selection unit forms an evidence set according to a preset quota for each evidence type. Evidence types include at least anomaly prototype evidence, multimodal semantic evidence, and path reasoning evidence. Anomaly prototype evidence is used to demonstrate the correspondence between historical anomaly patterns and the current anomaly state; multimodal semantic evidence is used to demonstrate direct evidence content in documents, tables, and images; path reasoning evidence is used to demonstrate the reasoning link from regional evidence to anomaly events or handling actions. A pre-set upper limit for the total number of evidences is set, and a minimum quota is allocated to different evidence types within this limit. If available evidence exists for a certain evidence type, at least one piece is selected; if no available evidence exists for a certain evidence type, it is supplemented by other evidence types. For high-level anomalies, the quota for path reasoning evidence and anomaly prototype evidence can be increased; for low-level anomalies, the quota for multimodal semantic evidence can be increased. The evidence set formed through this quota system not only retains high-scoring evidence but also ensures that evidence from different sources and with different functions has the opportunity to enter the generation stage, thereby reducing the risk of single-modality or single-path evidence dominating the output.
[0206] The context assembly unit organizes the evidence set into a context sequence that the generative model can directly accept. This context sequence follows the order of "Current Anomaly Summary—Key Evidence—Conflict Explanation—Disposition Basis." The Current Anomaly Summary section includes at least the anomaly category, anomaly severity, related objects, and anomaly time interval. The Key Evidence section lists the main evidence in order of fusion and sorting, with each main evidence appended with a source location identifier and evidence type identifier. The Conflict Explanation section is generated only when conflict markers exist, listing the source, conclusion, and priority order of conflicting evidence. The Disposition Basis section centrally presents anomaly prototype evidence and path reasoning evidence directly related to the disposition action. Each piece of evidence is written into the context sequence as a structured entry, which includes at least the evidence number, object identifier, time identifier, evidence type, source location, and content summary. Through this method, the generative model can clearly identify the source, category, and interrelationships of each piece of evidence without having to manually locate them within a long text.
[0207] The generation control unit determines the task identifier based on the user query type and inputs the task identifier along with the context sequence into the generation model. If the user query aims to provide a factual answer, the task identifier is set to a question-and-answer task; if the user query aims to analyze causes, explain processes, or summarize conclusions, the task identifier is set to an analysis report task; if the user query aims to determine maintenance actions, operating sequences, or handling measures, the task identifier is set to a handling suggestion task. When inputting into the generation model, the generation control unit also writes constraint instructions, requiring the generation model to prioritize outputting content based on the evidence set; when the evidence set does not contain handling evidence that meets the minimum coverage condition, the generation model outputs an insufficient evidence warning instead of providing specific handling suggestions. The minimum coverage condition can be set as follows: the evidence set must contain at least one path reasoning evidence related to the handling action and one sourced multimodal semantic evidence. This method avoids the generation model outputting unfounded handling content when evidence is insufficient.
[0208] In one specific implementation, the generative model employs a decoder-based Transformer structure. The input includes a task identifier and a context sequence, while the output consists of an answer, an analysis report, or a remedial recommendation. For the answer task, the generative model outputs a concise conclusion and retains the evidence number; for the analysis report task, the model outputs the cause of the anomaly, the chain of evidence, and a conclusion explanation; for the remedial recommendation task, the model outputs the inspection actions, remedial actions, and review actions in the order of their execution.
[0209] For example, if the current abnormal state corresponds to "bearing wear of equipment A, level 3 abnormality, current alarm period", and the evidence set simultaneously includes "disassembly and assembly steps in the maintenance manual", "6205 bearing inventory in the spare parts ledger", and "replacement actions obtained through path reasoning", then the generation control unit sets the task identifier as a handling suggestion task, and the generation model outputs handling suggestions for the abnormality of bearing A. If the evidence set only contains an abnormality description but lacks handling basis, the generation model outputs the prompt message "The current evidence supports the abnormality judgment, but lacks sufficient handling basis".
[0210] Another embodiment of the present invention provides a retrieval enhancement generation system that integrates anomaly detection with a dynamic knowledge base, comprising:
[0211] The data access module is used to receive multimodal data and real-time data streams;
[0212] The anomaly detection module is used to detect anomalies in the real-time data stream and obtain the current anomaly status.
[0213] The structured extraction module is used to extract the structured data from the multimodal data based on the current abnormal state to obtain the object to be parsed.
[0214] The parsing module is used to determine the target parsing object from the objects to be parsed based on the user query and the current abnormal state, and to perform parsing to obtain the region-level semantic representation and the document-level semantic representation.
[0215] The dynamic knowledge base construction module is used to build a dynamic knowledge base based on regional-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a regional evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer.
[0216] The prototype update module is used to update the hierarchical exception prototype layer according to the current exception status.
[0217] The sequence decision module is used to determine the retrieval execution order corresponding to the current request based on the user query, the current abnormal status, and historical feedback.
[0218] The retrieval execution module is used to perform anomaly prototype retrieval in the hierarchical anomaly prototype layer and multimodal semantic retrieval in the semantic vector layer and the regional evidence layer, according to the retrieval execution order.
[0219] The path reasoning module is used to perform path reasoning in the cross-modal knowledge graph layer to obtain candidate evidence;
[0220] The evidence fusion and generation module is used to fuse and sort candidate evidence to form an evidence set, and use the evidence set as context input to generate a model, outputting answers, analysis reports or disposal suggestions.
[0221] In summary, by enabling real-time abnormal states to continuously participate in the structured extraction, knowledge organization, retrieval order control, candidate evidence formation, and output generation of multimodal data, this invention can improve the accuracy of computer data processing in abnormal scenarios, enhance the correspondence and conflict detection capabilities of cross-source evidence, reduce resource consumption caused by irrelevant data participating in retrieval, reasoning, and generation, and improve the orderliness of evidence writing, retrieval, and context organization, which is conducive to achieving more stable storage / write control and access control security.
[0222] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for enhancing retrieval and generation by fusing anomaly detection with a dynamic knowledge base, characterized in that, The method includes: Receive multimodal data and real-time data stream, perform anomaly detection on the real-time data stream to obtain the current anomaly state, and perform structured extraction on the multimodal data based on the current anomaly state to obtain the object to be parsed; Based on the user query and the current abnormal state, the target parsing object is determined from the object to be parsed and parsing is performed to obtain the region-level semantic representation and the document-level semantic representation; A dynamic knowledge base is constructed based on the region-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a region evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer. The hierarchical anomaly prototype layer is updated according to the current anomaly state. Based on the user query, the current abnormal state, and historical feedback, the retrieval execution order corresponding to the current request is determined, and abnormal prototype retrieval is performed in the hierarchical abnormal prototype layer, multimodal semantic retrieval is performed in the semantic vector layer and the regional evidence layer, and path reasoning is performed in the cross-modal knowledge graph layer to obtain candidate evidence; The candidate evidence is fused and sorted to form an evidence set, and the evidence set is used as context input to generate a model to output an answer, analysis report or handling suggestion. A dynamic knowledge base is constructed based on the aforementioned region-level semantic representation and document-level semantic representation, including: Each region-level semantic representation, along with its corresponding object identifier, time identifier, modality identifier, and corresponding page hierarchy relationship, header dependency relationship, or spatial location relationship, is constructed into a region evidence entry and stored in the region evidence layer. Establish a one-to-one mapping relationship between each region-level semantic representation and its corresponding region evidence entry, establish an aggregate mapping relationship between each document-level semantic representation and its corresponding one or more region evidence entries, and store them in the semantic vector layer; The cross-modal knowledge graph layer is constructed based on cross-modal correspondence, object identifier, time identifier, and abnormal time interval. The nodes in the cross-modal knowledge graph layer include at least text entity nodes, table field nodes, image object nodes, regional evidence nodes, abnormal event nodes, and handling action nodes. The nodes are connected to each other at least through object association edges, time association edges, structural association edges, evidence support edges, and handling association edges. Based on the region-level semantic representation, document-level semantic representation, current abnormal state, and the graph sub-path in the cross-modal knowledge graph layer corresponding to the current abnormal state, the hierarchical abnormal prototype layer is constructed, which includes category-level abnormal prototype, semantic-level abnormal prototype, and instance-level abnormal prototype. Each instance-level anomaly prototype is associated with at least the anomaly category, anomaly degree, associated object, anomaly time interval, corresponding regional evidence entry, corresponding document-level semantic representation, and corresponding graph sub-path; each semantic-level anomaly prototype is associated with at least multiple instance-level anomaly prototypes with a semantic consistency of not less than a preset threshold; each category-level anomaly prototype is associated with at least one or more semantic-level anomaly prototypes under the same anomaly category.
2. The method for enhancing retrieval and generating data by fusing anomaly detection with a dynamic knowledge base according to claim 1, characterized in that, The multimodal data includes one or more of the following: doc documents, excel spreadsheets, ppt presentations, pdf files, and images; Anomaly detection of the real-time data stream includes: extracting temporal change features and object association features within a continuous time window; determining the current anomaly state based on the temporal deviation of the temporal change features relative to the historical baseline and the association deviation of the object association features relative to the association baseline map; wherein the current anomaly state includes at least anomaly category, anomaly degree, associated objects, and anomaly time interval. Based on the anomaly category, anomaly severity, associated objects, and anomaly time interval, anomaly association structured extraction is performed on the multimodal data, including: locating document layout areas corresponding to the associated objects and anomaly time intervals from doc documents, pdf files, and ppt presentations, and extracting the corresponding layout hierarchy; locating table cells corresponding to the associated objects and anomaly time intervals from Excel spreadsheets, and extracting the corresponding header dependencies; locating image object areas corresponding to the associated objects and anomaly time intervals from images, and extracting the corresponding spatial relationships; and determining the priority extraction order of document layout areas, table cells, and image object areas according to the anomaly category. When the degree of abnormality reaches a preset level, the abnormal time interval is extended forward and backward by a preset time length, and joint extraction is performed on the adjacent document layout area, adjacent table cell or adjacent image object area corresponding to the associated object. A cross-modal correspondence is established for the extraction results where the object identifiers after normalization are consistent across different modalities and the normalized time identifiers fall within the abnormal time interval. The extraction results are then associated and stored with the object identifier, time identifier, and modal identifier to obtain the object to be parsed.
3. The method for enhancing retrieval and generation by fusing anomaly detection with a dynamic knowledge base according to claim 2, characterized in that, Determining the target object to be parsed from the object to be parsed and performing parsing specifically includes: Calculate the first The target parsing score of the object to be parsed The formula is: ; In the formula, For user queries and the The semantic relevance of each object to be parsed; For the current abnormal state and the first The degree of abnormal relevance of the object to be parsed; For users to query the current abnormal status of the first The coupling correlation of the objects to be resolved; For the exception category and the first The information type matching degree of the object to be parsed; For the first The completeness of the cross-modal correspondence of the objects to be parsed; For the first The structural integrity of the object to be parsed; For the first Boundary reliability of the object to be resolved; For the first The normalization parsing cost of an object to be parsed; Analyze the scoring weights for the target, and satisfy the following conditions: ; According to the target analysis score Sort the objects to be parsed, and select the objects whose target parsing score is greater than a preset threshold or whose ranking is in the top predetermined number as the target parsing objects; Fine-grained parsing is performed on the target parsing object to obtain the corresponding region-level semantic representation. ; Multiple region-level semantic representations with the same source document identifier, the same object identifier, and time identifier falling within the same anomalous time interval are aggregated to obtain document-level semantic representations. The aggregation formula is as follows: ;and ; In the formula, For the first The document-level semantic representation of each target parsing object; This refers to the set of target parsing objects that have the same source document identifier, the same object identifier, and whose time identifiers fall within the same abnormal time interval. For the first The target parsing objects are in the target parsing object set Aggregate weights in; For the first The region-level semantic representation corresponding to each target parsing object; For the first Time consistency of the target parsing objects; For the target parsing object set The target parsing object number in the middle; The aggregate weight parameters satisfy the following conditions: .
4. The retrieval enhancement generation method based on anomaly detection and dynamic knowledge base according to claim 3, characterized in that, Updating the hierarchical exception prototype layer based on the current exception state includes: Based on the current abnormal state and the corresponding region-level semantic representation, document-level semantic representation, and graph sub-path, generate the current abnormal event representation. Calculate the first The prototype matching score between existing instance-level anomaly prototypes and the current anomaly event representation. The formula is: ; In the formula, Consistency of anomaly categories; The degree of anomaly matching; For the matching degree of the associated objects; This refers to the overlap of abnormal intervals. For the current abnormal event characterization and the first Semantic consistency among existing instance-level exception prototypes; Match weights to the prototype, and satisfy the following conditions: ; The existing instance-level exception prototype with the highest prototype matching score is selected as the candidate matching prototype. When the prototype matching score of the candidate matching prototype is not less than the first preset threshold, the candidate matching prototype is updated, and the central representation of the semantic-level abnormal prototype and the central representation of the category-level abnormal prototype are corrected according to the updated instance-level abnormal prototype representation. When the prototype matching score of the candidate matching prototype is less than the first preset threshold and not less than the second preset threshold, a new instance-level abnormal prototype is created under the semantic-level abnormal prototype to which the candidate matching prototype belongs, and the central representation of the corresponding semantic-level abnormal prototype and the central representation of the corresponding category-level abnormal prototype are corrected. When there is no existing instance-level exception prototype with a prototype matching score not less than the second preset threshold, a new instance-level exception prototype and a new semantic-level exception prototype are created based on the current exception event representation, and the new semantic-level exception prototype is attached to the category-level exception prototype of the corresponding exception category; when the category-level exception prototype of the corresponding exception category does not exist, a new category-level exception prototype is created.
5. The method for enhancing retrieval and generating anomaly detection based on a dynamic knowledge base according to claim 4, characterized in that, Based on the user query, the current abnormal status, and historical feedback, determine the retrieval execution order corresponding to the current request, including: A strategy experience layer is set in the dynamic knowledge base. Each strategy experience record in the strategy experience layer includes at least historical user query characteristics, historical abnormal state characteristics, historical retrieval execution order, and normalized feedback reward value. ; Calculate the current user query and the first Query similarity between historical user query features in the strategy experience record Calculate the current abnormal state and the first abnormal state. State similarity between historical anomalous state features in the policy experience record , and then calculate the first The sequential matching score corresponding to each strategy experience record The formula is: ; In the formula, For sequential matching weights, the following conditions must be met: ; Selecting the order matching score The historical search execution order that is greater than a first preset threshold or ranks among the top first predetermined number is used as a candidate historical search execution order; Based on the current degree of anomaly, the candidate historical retrieval execution order is expanded, pruned, or rearranged to obtain the retrieval execution order corresponding to the current request; The expansion includes adding pre-execution of anomaly prototype retrieval or supplementary execution of path reasoning when the current anomaly level is not lower than a first preset level; the pruning includes reducing the number of branches of multimodal semantic retrieval or reducing the expansion depth of path reasoning when the current anomaly level is not higher than a second preset level; and the reordering includes adjusting the execution order of anomaly prototype retrieval, multimodal semantic retrieval, and path reasoning according to the normalized feedback reward value and the priority rules corresponding to the anomaly category in the current anomaly state.
6. The method for enhancing retrieval and generation by fusing anomaly detection with a dynamic knowledge base according to claim 5, characterized in that, Performing anomaly prototype retrieval in the hierarchical anomaly prototype layer and multimodal semantic retrieval in the semantic vector layer and region evidence layer includes: In the hierarchical anomaly prototype layer, based on the matching results between the current anomaly state and each anomaly prototype in terms of anomaly category, anomaly degree, associated object, anomaly time interval and prototype representation, the prototype retrieval score corresponding to each anomaly prototype is calculated, and anomaly prototypes with prototype retrieval scores greater than the second preset threshold or ranked in the top second predetermined number are selected as anomaly prototype retrieval results. In the semantic vector layer and the regional evidence layer, based on the semantic similarity between the user query and each semantic vector record, the abnormal correlation between the current abnormal state and each semantic vector record, and the structural completeness of the regional evidence entries corresponding to each semantic vector record, the semantic retrieval score corresponding to each semantic vector record is calculated, and the semantic vector records with a semantic retrieval score greater than the third preset threshold or ranked in the top third predetermined number are selected as multimodal semantic retrieval results. Based on the abnormal prototype retrieval results, select the top four predetermined number of abnormal prototypes ranked by prototype similarity. Then, according to the hierarchical type of the abnormal prototype, call the corresponding instance-level abnormal prototype representation, semantic-level abnormal prototype center representation, or category-level abnormal prototype center representation. Finally, weighted aggregation is performed according to the corresponding prototype retrieval score to generate the [number missing]th [item missing]. The prototype guiding vector corresponding to each candidate piece of evidence ; The first Each candidate piece of evidence is encoded into a candidate evidence vector. For the candidate evidence vector and the prototype guiding vector Perform denoising fusion to obtain the first... Denoising fusion results corresponding to each candidate piece of evidence The formula is: ; in, ; In the formula, This represents the number of denoising expert functions; and This refers to the index of the denoising expert function; For the first A noise reduction expert function; For the first The denoising expert function for the first... Gating weights for each candidate piece of evidence; This indicates vector concatenation; and This represents the gating parameter vector corresponding to the denoising expert function; Based on the denoising fusion results The candidate evidence is reordered, and the candidate evidence whose reordered score meets the preset conditions is selected as the input for path reasoning.
7. The method for enhancing retrieval and generating anomaly detection based on a dynamic knowledge base according to claim 6, characterized in that, Performing path reasoning in the cross-modal knowledge graph layer includes: The corresponding regional evidence node, text entity node, table field node, or image object node of the path reasoning input is taken as the starting node, and the abnormal event node or handling action node corresponding to the current abnormal state is taken as the target node. The starting node is extended under constraints along the object-related edge, time-related edge, structural-related edge, evidence-supporting edge, and disposal-related edge to obtain one or more candidate paths; A path reasoning score is calculated for each candidate path, and the path reasoning score is determined at least based on object consistency, temporal consistency, structural consistency and disposal relevance in the path; Select candidate paths whose path reasoning scores are greater than the fourth preset threshold or whose ranking is among the top five predetermined number as valid reasoning paths; The evidence content corresponding to the regional evidence nodes, abnormal event nodes, and handling action nodes on the effective reasoning path is collected to obtain the path reasoning result; The path reasoning result is combined with the abnormal prototype retrieval result and the candidate evidence whose reordering score meets the preset conditions to obtain the candidate evidence.
8. The method for enhancing retrieval and generation by fusing anomaly detection with a dynamic knowledge base according to claim 1, characterized in that, The candidate evidence is fused and ranked to form an evidence set, and the evidence set is used as context input to generate a model, including: The candidate evidence is clustered according to object identifier, time identifier and source location. Deduplication and merging are performed on candidate evidence with the same object identifier, compatible time identifier and overlapping source location or content semantic similarity not less than the fifth preset threshold to obtain one or more candidate evidence clusters. For each candidate evidence cluster, determine the main evidence and supplementary evidence within the cluster. The main evidence is the candidate evidence with the highest fusion and ranking priority in the candidate evidence cluster, and the supplementary evidence is attached to the main evidence. For candidate evidence with the same object identifier and compatible time identifier but inconsistent conclusions, set a conflict flag. Based on user query relevance, current anomaly state consistency, anomaly prototype support, path reasoning support, cross-modal complementarity, source completeness, and conflict markers, the main evidence is fused and ranked. According to the preset evidence type quota, evidence is selected from each of the main evidences after fusion and sorting to form the evidence set, wherein the evidence types include at least abnormal prototype evidence, multimodal semantic evidence and path reasoning evidence; The evidence set is organized into a contextual sequence according to the order of current abnormal state summary, key evidence, conflict explanation and handling basis, and each piece of evidence is attached with a source location identifier and evidence type identifier; The task identifier is determined based on the user query type, and the context sequence and the task identifier are input into the generation model to output an answer, analysis report, or handling suggestion.
9. A retrieval enhancement generation system based on anomaly detection and dynamic knowledge base integration, wherein the retrieval enhancement generation method based on anomaly detection and dynamic knowledge base integration as described in any one of claims 1-8 is characterized in that, The system includes: The data access module is used to receive multimodal data and real-time data streams; An anomaly detection module is used to perform anomaly detection on the real-time data stream and obtain the current anomaly status; The structured extraction module is used to perform structured extraction on the multimodal data based on the current abnormal state to obtain the object to be parsed; The parsing module is used to determine the target parsing object from the object to be parsed based on the user query and the current abnormal state, and to perform parsing to obtain the region-level semantic representation and the document-level semantic representation. A dynamic knowledge base construction module is used to construct a dynamic knowledge base based on the region-level semantic representation and document-level semantic representation. The dynamic knowledge base includes at least a region evidence layer, a semantic vector layer, a cross-modal knowledge graph layer, and a hierarchical anomaly prototype layer. The prototype update module is used to update the hierarchical exception prototype layer according to the current exception state. The sequence decision module is used to determine the retrieval execution order corresponding to the current request based on the user query, the current abnormal state, and historical feedback. The retrieval execution module is used to perform anomaly prototype retrieval in the hierarchical anomaly prototype layer and multimodal semantic retrieval in the semantic vector layer and the regional evidence layer according to the retrieval execution order. The path reasoning module is used to perform path reasoning in the cross-modal knowledge graph layer to obtain candidate evidence; The evidence fusion generation module is used to fuse and sort the candidate evidence to form an evidence set, and use the evidence set as context input to generate a model to output an answer, analysis report or disposal suggestion.