Intelligent Data Analysis Method and System Based on Industry Large Model

Through intelligent data analysis methods based on industry big models, unified expression and correlation mining of multimodal data are realized, the problem of insufficient multi-source data processing framework in the existing technology is solved, the accuracy and interpretability of analysis results are improved, and the fusion strategy is dynamically adjusted to support the efficient use of industry knowledge.

CN120086266BActive Publication Date: 2025-07-11GUOTOU INTELLIGENT (NANJING) INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510581675.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-11
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When the existing technology faces the coexistence of structured and unstructured data, it lacks a unified and efficient processing framework, making it difficult to explore the implicit correlations and contextual dependencies between various modalities, and the accuracy and interpretability of the analysis results are relatively weak. Industry knowledge utilization only supports search based on keywords or shallow semantics, and cannot dynamically retrieve the most appropriate expert knowledge fragments. Attribution analysis methods cannot objectively measure the similarity between events and standard processes, and it is difficult to adapt to rapidly changing business scenarios.

Method used

The intelligent data analysis method based on industry big models is adopted, and the scaling dot product attention score matrix is calculated through the self-attention mechanism to perform multi-modal fusion, an industry knowledge database is constructed, and the gated input vector is used for fusion calculation, and the attribution path is determined in combination with the LCS algorithm to generate an abnormal attribution path list to realize the unified expression and correlation mining of multi-modal data.

Benefits of technology

It realizes unified expression and correlation mining of heterogeneous information such as structured data, text information and images, improves the accuracy and context fit of knowledge retrieval, dynamically adjusts the fusion strategy, provides objective attribution path alignment standards, suppresses pseudo-match, and improves the accuracy and interpretability of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086266B_ABST
    Figure CN120086266B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent data analysis method and system based on an industry large model, which relates to the technical field of data analysis. It includes collecting business event data, extracting feature vectors according to data types, calculating a scaled dot-product attention score matrix using a self-attention mechanism, performing weighted output of multi-modal fusion vectors, and constructing an industry knowledge database to determine embedding vectors; calculating correlation values to screen combinations of multi-modal fusion vectors and knowledge fragments, and forming gated input vectors for fusion vector calculation. The method of the present invention can achieve unified expression and correlation mining of heterogeneous information such as structured data, text information, and images through a multi-modal feature fusion and knowledge fragment retrieval scheme. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, realize information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to an intelligent data analysis method and system based on an industry big model. Background Art

[0002] With the continuous improvement of information infrastructure and the deepening of enterprise digital transformation, a large amount of structured, unstructured and multimodal business data has been accumulated in various fields of the industry, covering production sensor signals, equipment operation logs, business operation data, process text records, monitoring images and other types. As artificial intelligence technology, especially deep learning models, natural language processing algorithms and computer vision tools are gradually introduced into industrial data analysis scenarios, the ability to process complex business data has been improved;

[0003] However, existing methods are still mainly based on single data mode processing or weak correlation analysis, and have not yet achieved deep integration of multimodal information. The construction and application of industry knowledge bases are mostly at the level of static documentation or limited case reasoning, lacking dynamic knowledge mining and intelligent attribution support combined with real-time business data. Faced with the reality of coexistence of multiple sources of structured and unstructured data, traditional technologies lack a unified and efficient processing framework at the feature extraction, representation and fusion levels; it is difficult to mine the implicit associations and contextual dependencies between modes using simple splicing, static weighting and other means, resulting in weak accuracy and interpretability of analysis results. Secondly, regarding the use of industry knowledge, most related technologies only support retrieval based on keywords or shallow semantics, and cannot dynamically and accurately retrieve the most appropriate expert knowledge fragments based on the actual combination of multimodal business events. In addition, attribution analysis methods mostly use template matching or shallow sequence comparison, which can neither objectively measure the similarity between each event and the standard process, nor adapt to rapidly changing business scenarios. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an intelligent data analysis method based on an industry big model to solve the reality of the coexistence of multiple sources of structured and unstructured data. Traditional technologies lack a unified and efficient processing framework in terms of feature extraction, representation and fusion. It is difficult to mine the implicit associations and contextual dependencies between modalities using simple splicing, static weighting and other means, resulting in weak accuracy and interpretability of the analysis results. Secondly, regarding the use of industry knowledge, most related technologies only support retrieval based on keywords or shallow semantics, and cannot dynamically and accurately retrieve the most appropriate expert knowledge fragments based on the actual combination of multimodal business events. In addition, attribution analysis methods mostly use template matching or shallow sequence comparison, which can neither objectively measure the similarity between each event and the standard process, nor can they adapt to the problem of rapidly changing business scenarios.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an intelligent data analysis method based on an industry large model, which includes:

[0008] Collect business event data, extract feature vectors according to data types, calculate a scaled dot-product attention score matrix using a self-attention mechanism, perform weighted output of a multi-modal fusion vector, construct an industry knowledge database to determine embedding vectors;

[0009] Calculate relevance values to screen combinations of multi-modal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process;

[0010] Determine the category label to which the business event belongs to generate an event directed graph, extract a set of category abnormal state features, sort out the template path and the attribution path of the actual business, use the LCS algorithm to determine the path overlap length, and screen the highly similar aligned attribution paths;

[0011] Generate a direct index table of the attribution path and the template according to the abnormal state features, count the attribution paths that cannot be aligned in the statistical period to calculate the alignment deviation rate, and generate a list of abnormal attribution paths.

[0012] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the extracting feature vectors according to data types and constructing an industry knowledge database to determine embedding vectors includes,

[0013] Collecting business event data includes structured data, text data, and graphic data, preprocessing the collected data, and respectively performing data dimensionality reduction on the processed structured data, text data, and graphic data;

[0014] Among them, for structured data, a multi-layer perceptron MLP is used, and feature vectors are extracted through a ReLU non-linear activation function. For text data, a pre-trained BERT model is used for encoding, and the output text feature vectors are selected. For graphic data, a pre-trained ResNet50 model is used to extract features, and the output of its global average pooling layer is taken to obtain image feature vectors;

[0015] Perform normalization processing on the feature vectors of the three collected data using linear transformation, stack them to form a modal feature matrix, calculate Query, Key, and Value using a self-attention mechanism, calculate a scaled dot-product attention score matrix, apply softmax normalization, obtain an attention matrix, and weight the Value;

[0016] Subsequently, average pooling is adopted to fuse the information of the three modalities into a single vector, output the multi-modal fusion vector, and serialize it according to the industry business process to form a fused multi-modal vector sequence;

[0017] Construct an industry knowledge database based on enterprise regulations and expert documents. Uniformly embed each knowledge fragment in the knowledge base using the text encoder Enc to obtain the corresponding embedding vector, and the entire knowledge base corresponds to a set of embedding matrices;

[0018] For the multi-modal fusion vector of the business event, retrieve the K most relevant knowledge fragments in the knowledge base embedding matrix E to form a knowledge vector set, and calculate the relevance value using cosine similarity to screen K groups of combinations of multi-modal fusion vectors and knowledge fragments.

[0019] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: calculating the relevance value to screen the combination of the multi-modal fusion vector and the knowledge fragment, forming a gated input vector for fusion vector calculation, and obtaining an enhanced fusion vector sequence of the complete business process, including,

[0020] Concatenate according to the combination of the multi-modal fusion vector and the knowledge fragment as the gated input vector, construct a knowledge gated fusion model based on the feed-forward fully connected neural network, and obtain the gating factor through a single-layer neural network for the gated input vector;

[0021] Select the cross-entropy loss function to calculate the difference between the gating factor and the actual label of the training set, use the Adam optimizer for gradient descent optimization, update the parameters including weights and bias terms, and stop the iteration when the loss difference calculated in the continuous iteration process no longer decreases significantly;

[0022] Use the gating factor to weigh the information from two sources for fusion vector calculation. At the same time, for each knowledge fragment, obtain the relationship strength value with the initial fusion vector in the form of a dot product, and perform Softmax processing to obtain the normalized weight. Determine the final enhanced fusion vector using all K candidate enhanced vectors and the corresponding normalized weights;

[0023] Repeat the above operations for each event t in the business process, and finally obtain an enhanced fusion vector sequence of the complete business process.

[0024] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: determining the category label to which the business event belongs to generate an event directed graph, extracting the set of category abnormal state features, using the LCS algorithm to determine the path coincidence length, and screening the highly similar aligned attribution paths, including,

[0025] Determine the time point of the event according to different events in the enhanced fusion vector sequence, determine the category label of the business event and the event scene metadata, construct the time node, and determine the node set of all business events;

[0026] Obtain the main line of the industry business process according to the process standard table, construct directed edges according to the standard flow relationship from event i to event j in the business process, generate candidate directed edge sets for each adjacent business event based on the actual business occurrence sequence and the theoretical node connection sequence, and merge the edges generated by the standard process and the edges collected by the actual business process to generate an event directed graph;

[0027] According to the detailed domain labels given by the industry knowledge base, all event nodes are traversed, and nodes with the same domain labels and connected in the directed graph are determined as business sub-processes, and all business sub-processes are constructed into sub-process sets;

[0028] For each directed edge in the sub-process, a set of abnormal state features is extracted, including time series features, state jump metrics, mutation amplitude, process standard violation labels, and real-time monitoring alarm signs. Each abnormal state feature of each directed edge is counted and normalized and summed as the attribution score of each directed edge, and the total attribution score of each sub-process is counted.

[0029] Based on the score threshold set by the empirical method, the sub-process with a score higher than the score threshold is taken as the main attribution path. Based on the industry knowledge database, the archived paths of the enterprise over the years are collected and sorted as template paths. The abnormal state characteristics of the node step edges in the attribution path are simultaneously collected to generate a standard path template library.

[0030] For each attribution path obtained in actual business and each template path in the industry knowledge database, the node order and the abnormal state characteristics of each directed edge are encoded into a set of ordered feature sequence vectors, which are the structured feature sequence of the attribution path and the structured feature sequence of the expert knowledge template path respectively;

[0031] For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the longest common subsequence LCS algorithm to calculate the length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T as the two-dimensional array L;

[0032] If all fields of the attribution path step i and the template path step j are completely consistent, the consistent overlap length is determined. Otherwise, the node edges that do not match the attribution path and the template path are excluded, and the maximum overlap length is selected;

[0033] The ratio of the determined LCS coincidence length to the total length of the path is taken as the structural similarity score;

[0034] The sum of the mean and standard deviation of historical data is used as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, it is determined that the attribution path is highly similar and aligned with the template path.

[0035] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the generation of the direct index table of the attribution path and the template according to the abnormal state characteristics includes,

[0036] For the highly similar and aligned attribution paths screened, map the abnormal state characteristics of the attribution path to the abnormal state characteristic items of the template path and store them in a structured format to form a direct index table of the attribution path and the template path.

[0037] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: calculating the alignment deviation rate for the attribution paths that cannot be aligned in the statistical period to generate a list of abnormal attribution paths, including,

[0038] Count the number of attribution paths in the statistical period that cannot be aligned with any template path, and the total number of attribution paths in this period, calculate the alignment deviation rate of the current period. If the deviation rate is higher than the deviation threshold, output a list of abnormal attribution paths to form a feedback report.

[0039] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the acquisition of business event data includes,

[0040] Collect data of business events according to the industry business process. The structured data includes production sensors, device logs, and business system data. The text data includes work records, fault reports, and original text data. The graphic data includes on-site device monitoring and product image data.

[0041] In the second aspect, the present invention provides a system for the intelligent data analysis method based on the industry large model, including,

[0042] A multi-source data acquisition module responsible for collecting raw data;

[0043] A multi-modal fusion module that extracts features from the collected data and outputs a single multi-modal fusion vector;

[0044] An embedding management module that converts expert documents and historical knowledge fragments into a vectorized embedding library by a unified text encoder method;

[0045] A gated weighting module that obtains a dynamic gating factor according to the gated input, realizes the adaptive weighted fusion between the fusion vector and the knowledge embedding, and outputs an enhanced fusion vector sequence for the full-process event;

[0046] Anomaly feature encoding module, which generates a set of event nodes and directed edges to form a composite event directed graph;

[0047] Template matching module, which traverses the sorted attribution paths and template paths, performs the LCS algorithm on the path sequences to calculate the structural overlap length, and screens the key attribution paths with highly similar alignments;

[0048] Deviation feedback module, which establishes a traceable direct index table and outputs a list and report of anomaly attribution paths.

[0049] Thirdly, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the intelligent data analysis method based on the industry large model as described in the first aspect of the present invention is implemented.

[0050] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the intelligent data analysis method based on the industry large model as described in the first aspect of the present invention is implemented.

[0051] The beneficial effects of the present invention are as follows: through the multi-modal feature fusion and knowledge fragment retrieval scheme, the unified expression and correlation mining of heterogeneous information such as structured data, text information, and images can be realized. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, realize information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task. Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragments accumulated in history. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, so as to amplify or suppress the knowledge input as needed. The structural similarity score obtained by the ratio of the overlapping sequence length to the total length provides an objective and quantitative standard for the dynamic alignment of the attribution path and the template path, and suppresses the false matching caused by surface noise or local fluctuations. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic flowchart of the intelligent data analysis method based on the industry large model in Embodiment 1.

[0054] Figure 2This is a structural diagram of the intelligent data analysis system based on the industry big model in Example 1. DETAILED DESCRIPTION

[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0056] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0057] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0058] Example 1, reference Figures 1 to 2 , which is the first embodiment of the present invention, and provides an intelligent data analysis method based on an industry big model, comprising the following steps:

[0059] S1, collect business event data, extract feature vectors according to data types, and use the self-attention mechanism to calculate the scaled dot product attention score matrix, perform weighted output of multimodal fusion vectors, and build an industry knowledge database to determine the embedding vector;

[0060] Preferably, collecting business event data includes:

[0061] Data collection of business events is performed based on industry business processes. The collected business event data includes structured data, text data, and graphic data. Structured data includes production sensors, equipment logs, and business system data. Text data includes work records, fault reports, and original text data. Graphic data includes on-site equipment monitoring and product image data.

[0062] By collecting structured data, the production sensors, device logs, and business system data can comprehensively restore the production operation status and the dynamics of each link, facilitating subsequent fine-grained process analysis, anomaly location, and performance tracking. The collection of text data provides rich semantic support for the subjective descriptions of the business site, problem reports, and event details, helping to restore on-site operations, the occurrence and handling of faults, enhancing the capabilities of automatic analysis and natural language parsing, and providing a solid foundation for knowledge extraction and root cause reasoning. The collection of graphic data enables the operation of on-site equipment, product status, and visual anomalies to be intuitively reflected in the data flow, paving the data path for automatic visual inspection, anomaly warning, and process supervision, and achieving multi-angle and multi-modal perception of complex on-site situations. The synchronous acquisition of the three types of data greatly enhances subsequent multi-modal correlation analysis, knowledge fusion, and high-dimensional data-driven attribution effects, enabling the transparency and intelligence of business processes to reach a higher level.

[0063] Furthermore, feature vector extraction is performed according to the data type, and an industry knowledge database is constructed to determine the embedding vector, including

[0064] Preprocess the collected data, and perform data dimensionality reduction on the processed structured data, text data, and graphic data respectively;

[0065] Among them, for structured data, a multi-layer perceptron MLP is used, and feature vector extraction is performed through the ReLU non-linear activation function. For text data, a pre-trained BERT model is used for encoding, and the output text feature vector is selected. For graphic data, a pre-trained ResNet50 model is used to extract features, and the output of its global average pooling layer is taken as the image feature vector;

[0066] The feature vectors of the three collected data are standardized through linear transformation. Based on historical experience, a unified dimension value D is determined, stacked to form a modal feature matrix, and the self-attention mechanism is used to calculate Query, Key, and Value, and the scaled dot-product attention score matrix is calculated, expressed as:

[0067] ;

[0068] Among them represents the attention score between the i-th and j-th modalities, represents the Query value of the i-th modality, represents the transpose of the Key value of the j-th modality, and D represents the unified dimension value;

[0069] Apply softmax normalization to ensure that the dimension of softmax is set for each row, so that the attention distribution of each modality is reasonable, obtain the attention matrix and weight the Value, expressed as:

[0070] ;

[0071] ;

[0072] where represents the attention weight of the i-th and j-th modalities, represents the attention score between the i-th and k-th modalities, and R represents the total number of modalities, represents the i-th modality fusion vector, represents the Value value of the j-th modality;

[0073] Subsequently, average pooling is adopted to fuse the information of the three modalities into a single vector, output the multi-modal fusion vector and serialize it according to the industry business process to form a fused multi-modal vector sequence;

[0074] Construct an industry knowledge database based on enterprise regulations and expert documents, and uniformly embed each knowledge fragment in the knowledge base using the text encoder Enc to obtain the corresponding embedding vectors. The entire knowledge base corresponds to a set of embedding matrices;

[0075] For the multi-modal fusion vector of the business event, retrieve the K most relevant knowledge fragments in the knowledge base embedding matrix E to form a knowledge vector set, and use cosine similarity to calculate the relevance value to screen K groups of combinations of multi-modal fusion vectors and knowledge fragments, where the value of K can be set according to historical experience.

[0076] Through the multi-modal feature fusion and knowledge fragment retrieval scheme, unified expression and correlation mining of heterogeneous information such as structured data, text information, and images can be achieved. Through preprocessing and dimensionality reduction, effective compression and redundancy removal of various types of raw data are ensured, improving the utilization rate of downstream features. Different types of data are respectively used with neural network models suitable for their semantic structures for feature extraction. For example, structured data uses MLP to obtain deep abstract expressions, text data uses BERT to capture context and semantics, and images use ResNet50 to refine complex spatial structures, which can maximize the retention of their respective information advantages. Linear mapping normalization solves the problems of multi-modal feature different scales and different distributions, laying a foundation for subsequent unified fusion;

[0077] Through the self-attention mechanism, the model can adaptively adjust the attention of each modality feature according to the business scenario, achieve information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task; softmax row-wise normalization and weighted processing ensure that the calculated attention distribution has physical and business rationality. This fusion method eliminates the uncertainty and inefficiency brought by subjective weighting or manual feature engineering, and realizes deep coupling of information from different sources;

[0078] By adopting the average pooling method, a single vector with high information integration degree is finally obtained, which is automatically aligned with the business process to form a standard fusion feature sequence expression, which is convenient for subsequent business sequence modeling and traceability. After the knowledge base content is uniformly embedded in the text and represented by the vector, the relationship between structured business data and domain knowledge can be established, which greatly improves the accuracy of knowledge retrieval and the fit of context business.

[0079] Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragment accumulated in history, effectively shortening the response time and manual search costs, ensuring the model's ability to discover and utilize knowledge in series, and the dynamic K value setting combines the actual business scale and historical case density to make the model neither overfitting nor flexible enough. The overall process greatly improves the depth and accuracy of business event understanding, and also provides a strong information foundation and significant global optimization effect for process compliance, automatic attribution and intelligent decision-making.

[0080] S2, calculate the correlation value to screen the combination of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process;

[0081] Preferably, the correlation value is calculated to screen the combination of multimodal fusion vectors and knowledge fragments, and the gated input vector is formed to perform fusion vector calculation to obtain an enhanced fusion vector sequence of the complete business process, including:

[0082] According to the combination of multimodal fusion vector and knowledge fragments, they are spliced ​​as the gated input vector, and a knowledge gated fusion model is constructed based on a feedforward fully connected neural network. The gated factor of the single-layer neural network of the gated input vector is obtained, which is expressed as:

[0083] ;

[0084] in represents the gating factor for the t-th business event and its j-th retrieved knowledge fragment, represents the Sigmoid activation function, represents the transpose calculation of the weight vector, represents the concatenated vector of the t-th business event multimodal fusion vector and the j-th retrieved knowledge fragment embedding, represents the bias term;

[0085] Select the cross entropy loss function to calculate the difference between the gating factor and the actual label of the training set, use the Adam optimizer for gradient descent optimization, update the parameters including weights and bias terms, and stop the iteration if the calculated loss difference no longer decreases significantly during the continuous iteration process;

[0086] The gating factor is used to weigh two sources of information for fusion vector calculation, expressed as:

[0087] ;

[0088] where represents the candidate enhancement vector after the fusion of the t-th business event and the j-th retrieved knowledge fragment, represents the multimodal fusion vector of the t-th business event, represents the embedding vector of the j-th knowledge fragment retrieved by the t-th business event;

[0089] At the same time, for each knowledge fragment, the relationship strength value with the initial fusion vector is obtained in the form of dot product, expressed as:

[0090] ;

[0091] where represents the relationship strength value between the t-th business event and the j-th retrieved knowledge fragment, represents the transpose calculation of the multimodal fusion vector;

[0092] And Softmax processing is performed to obtain the normalized weight, and the final enhanced fusion vector is determined by all K candidate enhancement vectors and the corresponding normalized weights;

[0093] The above operations are repeated for each event t in the business process, and finally an enhanced fusion vector sequence of the complete business process is obtained.

[0094] When the multimodal features of business events are combined with the knowledge fragment embeddings, it is not simply a direct superposition, but rather the contribution of each type of information is selectively determined through a gating mechanism. This method utilizes the non-linear modeling ability of neural networks to automatically mine the most valuable correlation relationships between business events and knowledge fragments in different scenarios, avoiding the subjective biases of static or artificial prior weighting. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, thereby amplifying or suppressing knowledge input as needed, making the fusion vector more in line with the actual business scenario and empirical rules;

[0095] Through the training set labels and cross-entropy loss, the model continuously learns from past data what kind of fusion ratio and method can bring optimal decisions, so that the finally obtained enhanced fusion vector not only contains the most core multimodal information of the current business, but also fully integrates the essence of historical knowledge, facilitating downstream analysis and intelligent reasoning; the dot product relationship strength and Softmax normalization realize the automatic ranking of the importance of different candidate fusion vectors, and the model can effectively filter redundant or low-correlation knowledge, achieve focusing and utilize the most valuable information fragments, and improve the pertinence and effect of knowledge-assisted reasoning;

[0096] For each event in the business process, repeat the above fusion steps. The enhanced fusion vector sequence of the entire process can more precisely capture the dynamic relationship between each link in the event chain and the knowledge base, as well as the latest business situation. The final sequence can not only reflect the global business evolution but also have profound guiding significance of industry knowledge, providing a solid and logically interpretable information foundation for subsequent tasks such as complex business attribution, process analysis, and risk identification.

[0097] S3. Determine the category label to which the business event belongs, generate an event directed graph, extract the set of category abnormal state features, organize the template path and the attribution path of the actual business, use the LCS algorithm to determine the path overlap length, and screen the highly similar and aligned attribution paths.

[0098] Preferably, determining the category label to which the business event belongs, generating an event directed graph, extracting the set of category abnormal state features, using the LCS algorithm to determine the path overlap length, and screening the highly similar and aligned attribution paths includes:

[0099] Determine the time point of the event according to different events in the enhanced fusion vector sequence, and determine the category label to which the business event belongs (such as process stage number, equipment type, alarm type, etc., taken from actual business data) and the event site metadata (such as timestamp, operator, production line ID, physical location, relevant process parameters, all from the available structured data in the business system), construct time nodes, and determine the node set of all business events.

[0100] According to the process standard, clearly obtain the main line of the industry business process. Construct directed edges according to the standard transfer relationship that event i directly flows to event j in the business process. For each adjacent business event, generate a candidate directed edge set according to the actual occurrence order of the real business and the theoretical connection order between nodes. For the standard process, the directed edges are generated by directly connecting processes, links, or business instructions according to industry regulations. For the actual event flow, the direct transfer relationship between events is automatically generated according to the occurrence order of adjacent events in the real business, and each generated directed edge is attached with detailed attributes including category labels and site metadata.

[0101] Merge the edges generated by the standard process and the edges collected from the actual business process to generate an event directed graph, that is, it contains both the theoretical transfer relationship and the actual transfer that actually occurred on site. In this way, whether it is process supervision, anomaly detection, root cause analysis, or dynamic process optimization, the "ideal" and "real" all-round situations can be considered simultaneously.

[0102] According to the detailed domain tags given by the industry knowledge base, traverse all event nodes, and determine the nodes with the same domain tags and connected in the directed graph as business sub-processes, and construct a sub-process set for all business sub-processes;

[0103] For each directed edge within the sub-process, extract a set of category anomaly status features, including timing features (the time interval between the occurrences of consecutive events), state transition metrics (the mutation amplitude of device core parameters (temperature, pressure, load, etc.)), process flow standard violation labels (the program automatically compares whether the transfer nodes of the process skip the necessary processes, whether they flow backward, and whether there are illegal or non-compliant links), real-time monitoring alarm flags (faults, overlimits, or alarm events generated by the ERP / MES / sensor system). For each anomaly status feature of each directed edge, perform statistics and normalized summation as the attribution score of each directed edge, and count the total attribution score of each sub-process;

[0104] Based on the score threshold set by the empirical method, and take the sub-process with a score higher than the score threshold as the main attribution path. Based on the industry knowledge database, collect and organize the enterprise's archived paths over the years as template paths, including classic defect root cause paths, typical problem flows, expert-reviewed anomaly propagation chains, and business standard propagation templates. Synchronously collect the anomaly status features of the nodes, steps, and edges in the attribution path to generate a standard path template library;

[0105] For each attribution path obtained in the actual business and each template path in the industry knowledge database, encode the node order and the anomaly status features of each directed edge into a set of ordered feature sequence vectors, namely the structured feature sequence of the attribution path and the structured feature sequence of the expert knowledge template path, which can be encoded and stored using a JSON array;

[0106] For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the longest common subsequence LCS algorithm. Take the length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T as the two-dimensional array L;

[0107] If all fields of the i-th step of the attribution path and the j-th step of the template path are exactly the same, determine the consistent overlap length. Otherwise, select to exclude the nodes and edges where the attribution path and the template path do not match, and select the maximum overlap length, which is expressed as:

[0108] ;

[0109] ;

[0110] where represents the consistent overlap length between the first i steps of the attribution path and the first j steps of the template, Indicates the length of the longest common subsequence that the previous pair has completed alignment. represents the maximum overlap length of the first i steps of the attribution path and the first j steps of the template, and They represent the longest common subsequence lengths of only the first i-1 steps of the attribution path and the first j steps of the template path, and only the first i steps of the attribution path and the first j-1 steps of the template path;

[0111] The ratio of the determined LCS coincidence length to the total length of the path is taken as the structural similarity score;

[0112] The sum of the mean and standard deviation of historical data is used as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, the attribution path is judged to be highly similar and aligned with the template path.

[0113] Business events are divided into strict node sets, each with traceable business metadata and event categories, so that the description of the event chain has clear logic and data authenticity. Directed edges are extracted from the theoretical process and actual execution respectively and merged into a unified graph structure, which allows subsequent process optimization, deviation identification and root cause tracing to be carried out based on the overall business logic, avoiding the blind spots of a single perspective and providing the possibility of dynamic comparison between ideal and reality.

[0114] Sub-process division is completed through the linkage of domain labels and node connectivity, which can accurately aggregate similar business links and automatically adapt to on-site process changes, reducing the deviation that may be caused by manual division. For each directed edge, multi-dimensional abnormal features such as time, state, process, and alarm are counted, and the sum is normalized into an objective and measurable attribution score. The degree of abnormality of each sub-process is further quantified, realizing rapid focus on the core abnormal chain that is most likely to affect business results.

[0115] The score threshold set by the empirical method can avoid high-noise and low-correlation processes from occupying analysis resources, so that the analysis can focus on the process sections that are truly abnormal and may generate business risks; the standardized template path library and synchronous feature vector storage are used to form a highly consistent structured comparison of the actual attribution results with the typical problems and experiences accumulated in history. This comparison does not rely on manual experience, nor will it miss new abnormal patterns in the business process, which fundamentally improves the applicability and real-time performance of automated intelligent analysis.

[0116] After comparing the LCS algorithms, the structural similarity score obtained from the ratio of the overlapping sequence length to the total length provides an objective and quantitative standard for the automatic "classification" and dynamic alignment of the attribution path and the template path, and suppresses the false matching caused by surface noise or local fluctuations. Using the threshold synthesized by the historical mean and standard deviation as the final criterion enables the identification of the main attribution path and highly reliable automatic archiving to have a data-driven adaptive mechanism, truly empowering the process anomaly localization and knowledge closed-loop improvement in large-scale industrial scenarios.

[0117] S4. Generate a direct index table of the attribution path and the template according to the abnormal state characteristics. Calculate the alignment deviation rate for the attribution paths that could not be aligned during the statistical period, and generate a list of abnormal attribution paths.

[0118] Preferably, generating a direct index table of the attribution path and the template according to the abnormal state characteristics includes

[0119] For the highly similar aligned attribution paths that are screened, map the abnormal state characteristics of the attribution path to the abnormal state characteristic items of the template path and store them in a structured format to form a direct index table of the attribution path and the template path.

[0120] By mapping the abnormal state characteristics of the highly similar aligned attribution paths and the template paths one by one and storing them in a structured manner, the precise information association between the attribution result and the standard template is realized, ensuring that all abnormal details of the attribution result can be accurately retrieved and searched by subsequent automated systems, operation and maintenance personnel, or decision-making systems, improving the data reuse and query efficiency. The direct index relationship between the attribution path and the template path brings seamless connection between business attribution and historical knowledge, facilitating the statistics of attribution times, monitoring the occurrence trend of specific abnormal types, and discovering the high-incidence links of process optimization, providing a solid structural foundation for the subsequent abnormal tracking, attribution explanation, compliance audit, and continuous expansion of the knowledge base of the enterprise, greatly enhancing the automation and controllability of management and analysis.

[0121] Furthermore, calculating the alignment deviation rate for the attribution paths that could not be aligned during the statistical period and generating a list of abnormal attribution paths includes

[0122] Save all node sequences, edge abnormal state characteristics of the attribution path, and the data of its matching structural similarity score to generate a log file.

[0123] Count the number of attribution paths in the statistical period that could not be aligned with any template path, and the total number of attribution paths in this period. Calculate the alignment deviation rate of the current period. If the deviation rate is higher than the deviation threshold (set according to historical experience), output the list of abnormal attribution paths to form a feedback report.

[0124] By detailedly recording the node sequences of each attribution path, the abnormal state characteristics of all edges, and the specific structural similarity scores and generating logs, it enables subsequent traceability, playback, and problem review to have the support of the original data source, ensuring that each attribution analysis can be verified and reused conveniently. Periodically count the number of unaligned attribution paths and the total number of attribution paths, and calculate the alignment deviation rate accordingly. Through the quantitative index, it can dynamically reflect the coverage degree of the current knowledge template for the actual business complexity, and timely expose the "blind spots" of the knowledge base or the trend of new abnormal patterns in the business.

[0125] When the deviation rate exceeds the predetermined threshold, the system can automatically screen out all abnormal attribution paths, form a feedback report, and accurately prompt the management and analysis team to pay attention to possible production line changes, equipment aging, or new problems, supporting rapid response and the expansion and optimization of knowledge templates. This process realizes the automatic archiving, statistics, and reporting of abnormal information, ensuring that the enterprise can continuously optimize the business process and knowledge assets in a closed loop, effectively improving the problem warning and adaptive adjustment capabilities.

[0126] This embodiment also provides a system for an intelligent data analysis method based on an industry large model, including,

[0127] A multi-source data acquisition module, responsible for collecting raw data from structured data sources, text data sources, and graphic data sources;

[0128] A multimodal fusion module, which extracts features from the collected data, performs standardization mapping and feature alignment, and dynamically models the information relationship between multimodals through the self-attention mechanism, and outputs a single multimodal fusion vector;

[0129] An embedding management module, which uses a unified text encoder method to convert expert documents and historical knowledge fragments into a vectorized embedding library;

[0130] A gating and weighting module, which obtains a dynamic gating factor according to the gating input, realizes the adaptive weighted fusion between the fusion vector and the knowledge embedding, and outputs an enhanced fusion vector sequence for the full-process event;

[0131] An abnormal feature encoding module, which generates an event node set and directed edges to form a composite event directed graph;

[0132] A template matching module, which traverses the sorted attribution paths and template paths, calculates the structural coincidence length of the path sequence using the LCS algorithm, and combines the attribution scores and structural similarity scores to screen out the key attribution paths with highly similar alignments;

[0133] A deviation feedback module, which establishes a traceable direct index table and outputs a list and report of abnormal attribution paths.

[0134] This embodiment also provides a computer device applicable to the situation of the intelligent data analysis method based on an industry large model, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent data analysis method based on the industry large model proposed in the above embodiment.

[0135] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0136] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the intelligent data analysis method based on the industry large model proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Red-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0137] In summary, through the multi-modal feature fusion and knowledge fragment retrieval scheme, the present invention can achieve the unified expression and correlation mining of heterogeneous information such as structured data, text information, and images. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, realize information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task. Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragments accumulated in history. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, thereby amplifying or suppressing the knowledge input as needed. The structural similarity score obtained by the ratio of the overlapping sequence length to the total length provides an objective and quantitative standard for the automatic "classification" and dynamic alignment of the attribution path and the template path, and suppresses the false matching caused by surface noise or local fluctuations.

[0138] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An intelligent data analysis method based on an industry large model, characterized in that, Including: Collect business event data including structured data, text data, and graphical data, extract feature vectors according to the data type, and use the self-attention mechanism to calculate the scaled dot-product attention score matrix, perform weighted output of the multimodal fusion vector, and construct an industry knowledge database to determine the embedding vector; Calculate the relevance value to screen the combinations of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process; Determine the category label to which the business event belongs to generate an event directed graph, extract the set of category abnormal state features, sort out the template path and the attribution path of the actual business, use the LCS algorithm to determine the path coincidence length, and screen the highly similar aligned attribution paths; Generate a direct index table of the attribution path and the template according to the abnormal state features, calculate the alignment deviation rate for the attribution paths that cannot be aligned in the statistical period, and generate a list of abnormal attribution paths.

2. The intelligent data analysis method based on the industry large model according to claim 1, wherein: The feature vector extraction according to the data type and the construction of the industry knowledge database to determine the embedding vector include Collect business event data including structured data, text data, and graphical data, preprocess the collected data, and perform data dimensionality reduction on the processed structured data, text data, and graphical data respectively; Among them, for structured data, use a multi-layer perceptron MLP to extract feature vectors through the ReLU non-linear activation function, use a pre-trained BERT model to encode text data, and select the output text feature vectors. For graphical data, use a pre-trained ResNet50 model to extract features, and take the output of its global average pooling layer to obtain the image feature vector; Perform standardization processing on the feature vectors of the three collected data using linear transformation, stack them to form a modal feature matrix, use the self-attention mechanism to calculate Query, Key, and Value, calculate the scaled dot-product attention score matrix, apply softmax normalization to obtain the attention matrix and weight the Value; Subsequently, use average pooling to fuse the information of the three modalities into a single vector, output the multimodal fusion vector and serialize it according to the industry business process to form a fused multimodal vector sequence; Build an industry knowledge database based on enterprise regulations and expert documents, uniformly use a text encoder Enc to embed each knowledge fragment in the knowledge base to obtain the corresponding embedding vector, and the entire knowledge base corresponds to a set of embedding matrices; For the multimodal fusion vector of the business event, retrieve the most relevant K knowledge fragments in the knowledge base embedding matrix E to form a knowledge vector set, and use cosine similarity to calculate the relevance value to screen K combinations of multimodal fusion vectors and knowledge fragments.

3. The intelligent data analysis method based on an industry large model according to claim 2, wherein: The calculation of the relevance value to screen the combinations of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and obtain an enhanced fusion vector sequence of the complete business process, including Concatenate according to the combination of the multimodal fusion vector and the knowledge fragment as the gated input vector, construct a knowledge gating fusion model based on a feed-forward fully connected neural network, and obtain a gating factor for the gated input vector through a single-layer neural network; The cross-entropy loss function is selected to calculate the difference between the gating factor and the actual labels of the training set. The Adam optimizer is used for gradient descent optimization, and the updated parameters include weights and bias terms. The iteration stops when the loss difference calculated in consecutive iterations no longer decreases significantly. The gating factor is used to weigh the two source information for the calculation of the fusion vector. At the same time, for each knowledge fragment, the relationship strength value with the initial fusion vector is obtained in the form of a dot product and is processed by Softmax to obtain the normalized weights. The final enhanced fusion vector is determined by all K candidate enhanced vectors and the corresponding normalized weights. The above operations are repeated for each event t in the business process, and finally an enhanced fusion vector sequence of the complete business process is obtained.

4. The intelligent data analysis method based on an industry large model according to claim 3, wherein: The category label to which the business event belongs is determined to generate an event directed graph, the set of category abnormal state features is extracted, the LCS algorithm is used to determine the path coincidence length, and the highly similar aligned attribution paths are screened. including, The time point of the event is determined according to different events in the enhanced fusion vector sequence, and the category label to which the business event belongs and the event site metadata are determined, and a time node is constructed, and the node set of all business events is determined. The main line of the industry business process is clearly obtained according to the process standard. The directed edges are constructed according to the standard transfer relationship that event i directly flows to event j in the business process. For each adjacent business event, candidate directed edge sets are generated respectively according to the actual business occurrence order and the theoretical specified connection order between nodes. The edges generated by the standard process and the edges collected from the actual business process are combined to generate an event directed graph. According to the detailed domain labels given by the industry knowledge base, all event nodes are traversed, and the nodes with the same domain label and connected in the directed graph are determined as business sub-processes, and all business sub-processes are constructed into a sub-process set. For each directed edge in the sub-process, the set of category abnormal state features is extracted, including time series features, state jump metrics, mutation amplitudes, process flow standard violation labels, and real-time monitoring alarm flags. Each abnormal state feature of each directed edge is statistically analyzed and normalized and summed as the attribution score of each directed edge, and the total attribution score of each sub-process is statistically analyzed. Based on the score threshold set by the empirical method, the sub-processes with scores higher than the score threshold are used as the main attribution paths. Based on the industry knowledge database, the archived paths of the enterprise over the years are collected and sorted as template paths, and the abnormal state features of the nodes, steps, and edges in the attribution paths are synchronously collected to generate a standard path template library. For each attribution path obtained in the actual business and each template path in the industry knowledge database, the node order and the abnormal state features of each directed edge are encoded into a group of ordered feature sequence vectors, which are respectively the structured feature sequence of the attribution path and the structured feature sequence of the expert knowledge template path. For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the longest common subsequence LCS algorithm. According to the length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T, it is used as the two-dimensional array L. If all the fields in the i-th step of the attribution path and the j-th step of the template path are exactly the same, then determine the consistent overlap length; otherwise, select and exclude the node edges where the attribution path and the template path do not match, and select the maximum overlap length; Use the ratio of the determined LCS overlap length to the total length of the path as the structural similarity score; Use the sum of the mean and standard deviation of the historical data as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, it is determined that the attribution path and the template path are highly similar and aligned.

5. The intelligent data analysis method based on the industry large model according to claim 4, characterized in that: Generate a direct index table of the attribution path and the template according to the abnormal state characteristics, including, For the highly similar and aligned attribution paths screened, map the abnormal state characteristics of the attribution path to the abnormal state characteristic items of the template path and store them in a structured format to form a direct index table of the attribution path and the template path.

6. The intelligent data analysis method based on the industry large model according to claim 5, characterized in that: Calculate the alignment deviation rate for the attribution paths that cannot be aligned in the statistical period, and generate a list of abnormal attribution paths, including, Count the number of attribution paths in the statistical period that cannot be aligned with any template path, and the total number of attribution paths in this period, calculate the alignment deviation rate of the current period. If the deviation rate is higher than the deviation threshold, output the list of abnormal attribution paths to form a feedback report.

7. The intelligent data analysis method based on the industry large model according to claim 6, characterized in that: Collect business event data, including, Collect data of business events according to the industry business process. The structured data includes production sensors, device logs, and business system data. The text data includes work records, fault reports, and raw text data. The graphic data includes on-site device monitoring and product image data.

8. A system for an intelligent data analysis method based on an industry large model, based on the intelligent data analysis method based on an industry large model according to any one of claims 1 to 7, characterized in that: including, A multi-source data collection module that collects business event data including structured data, text data, and graphic data; A multi-modal fusion module that extracts features from the collected data and outputs a single multi-modal fusion vector; An embedding management module that converts expert documents and historical knowledge fragments into a vectorized embedding library by a unified text encoder method; A gating and weighting module that obtains a dynamic gating factor according to the gating input, realizes the adaptive weighted fusion between the fusion vector and the knowledge embedding, and outputs an enhanced fusion vector sequence for the full-process event; An abnormal feature encoding module that generates a set of event nodes and directed edges to form a composite event directed graph; A template matching module that traverses the sorted attribution paths and template paths, performs the LCS algorithm on the path sequence to calculate the structural overlap length, and screens the key attribution paths that are highly similar and aligned; A deviation feedback module that establishes a traceable direct index table and outputs a list of abnormal attribution paths and reports.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent data analysis method based on the industry large model according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent data analysis method based on the industry large model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Inspection method, device and equipment based on digital airspace system and medium

    CN119989283A