Intelligent data analysis method and system based on industry large model
Through the intelligent data analysis method based on industry big models, the lack of correlation mining and knowledge utilization in multimodal data processing is solved, efficient data analysis and knowledge retrieval are achieved, and the accuracy and interpretability of the analysis results are improved.
Patent Information
- Application Number
- CN202510581675.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art lacks a unified and efficient processing framework when processing multimodal data, making it difficult to explore the implicit correlations and contextual dependencies between each modal, resulting in weak accuracy and interpretability of the analysis results. At the same time, the utilization of industry knowledge mainly relies on search based on keywords or shallow semantics, and the most appropriate expert knowledge fragments cannot be retrieved dynamically and accurately, and the attribution analysis method is difficult to adapt to rapidly changing business scenarios.
Using an intelligent data analysis method based on industry big models, we collect business event data, extract multimodal feature vectors, and use a self-attention mechanism to weight output multimodal fusion vectors. At the same time, an industry knowledge database is constructed, the correlation value is calculated to filter the combination of multimodal fusion vectors and knowledge fragments, and an enhanced fusion vector sequence is generated. The LCS algorithm is used to determine the similarity between the attribution path and the template path, generate a direct index table and calculate the alignment deviation rate.
It realizes unified expression and correlation mining of heterogeneous information such as structured data, text information and images, and improves the accuracy and interpretability of the analysis results. Through the fusion of dynamic knowledge retrieval and adaptive weighting, the accuracy and adaptability of knowledge utilization are improved. At the same time, through objective alignment standards and direct index tables, the traceability and analysis efficiency of attribution paths are enhanced.
Smart Images

Figure CN120086266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to an intelligent data analysis method and system based on an industry big model. Background Art
[0002] With the continuous improvement of information infrastructure and the deepening of enterprise digital transformation, a large amount of structured, unstructured and multimodal business data has been accumulated in various fields of the industry, covering production sensor signals, equipment operation logs, business operation data, process text records, monitoring images and other types. As artificial intelligence technology, especially deep learning models, natural language processing algorithms and computer vision tools are gradually introduced into industrial data analysis scenarios, the ability to process complex business data has been improved; However, existing methods are still mainly based on single data mode processing or weak correlation analysis, and have not yet achieved deep integration of multimodal information. The construction and application of industry knowledge bases are mostly at the level of static documentation or limited case reasoning, lacking dynamic knowledge mining and intelligent attribution support combined with real-time business data. Faced with the reality of coexistence of multiple sources of structured and unstructured data, traditional technologies lack a unified and efficient processing framework at the feature extraction, representation and fusion levels; it is difficult to mine the implicit associations and contextual dependencies between modes using simple splicing, static weighting and other means, resulting in weak accuracy and interpretability of analysis results. Secondly, regarding the use of industry knowledge, most related technologies only support retrieval based on keywords or shallow semantics, and cannot dynamically and accurately retrieve the most appropriate expert knowledge fragments based on the actual combination of multimodal business events. In addition, attribution analysis methods mostly use template matching or shallow sequence comparison, which can neither objectively measure the similarity between each event and the standard process, nor adapt to rapidly changing business scenarios. Summary of the invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides an intelligent data analysis method based on an industry big model to solve the reality of the coexistence of multiple sources of structured and unstructured data. Traditional technologies lack a unified and efficient processing framework in terms of feature extraction, representation and fusion. It is difficult to mine the implicit associations and contextual dependencies between modalities using simple splicing, static weighting and other means, resulting in weak accuracy and interpretability of the analysis results. Secondly, regarding the use of industry knowledge, most related technologies only support retrieval based on keywords or shallow semantics, and cannot dynamically and accurately retrieve the most appropriate expert knowledge fragments based on the actual combination of multimodal business events. In addition, attribution analysis methods mostly use template matching or shallow sequence comparison, which can neither objectively measure the similarity between each event and the standard process, nor can they adapt to the problem of rapidly changing business scenarios.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides an intelligent data analysis method based on an industry large model, which includes: Collect business event data, extract feature vectors according to data types, and use the self-attention mechanism to calculate the scaled dot-product attention score matrix, perform weighted output of multimodal fusion vectors, and construct an industry knowledge database to determine embedding vectors; Calculate the relevance value to screen the combinations of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process; Determine the category label to which the business event belongs to generate an event directed graph, extract the set of category abnormal state features, sort out the template path and the attribution path of the actual business, use the LCS algorithm to determine the path coincidence length, and screen the highly similar aligned attribution paths; Generate a direct index table of the attribution path and the template according to the abnormal state features, calculate the alignment deviation rate for the attribution paths that cannot be aligned in the statistical period, and generate a list of abnormal attribution paths.
[0006] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the extraction of feature vectors according to data types and the construction of an industry knowledge database to determine embedding vectors include, Collecting business event data includes structured data, text data, and graphic data, preprocessing the collected data, and respectively performing data dimensionality reduction on the processed structured data, text data, and graphic data; Among them, for structured data, a multi-layer perceptron MLP is used, and feature vectors are extracted through the ReLU non-linear activation function. For text data, a pre-trained BERT model is used for encoding, and the output text feature vectors are selected. For graphic data, a pre-trained ResNet50 model is used to extract features, and the output of its global average pooling layer is taken to obtain image feature vectors; Perform standardization processing on the feature vectors of the three collected data through linear transformation, stack them to form a modal feature matrix, use the self-attention mechanism to calculate Query, Key, and Value, calculate the scaled dot-product attention score matrix, apply softmax normalization, obtain the attention matrix and weight the Value; Subsequently, average pooling is used to fuse the information of the three modalities into a single vector, output the multimodal fusion vector and serialize it according to the industry business process to form a fused multimodal vector sequence; Construct an industry knowledge database based on enterprise regulations and expert documents. For each knowledge fragment in the knowledge base, use the text encoder Enc to embed it to obtain the corresponding embedding vector, and the entire knowledge base corresponds to a set of embedding matrices. For the multi-modal fusion vector of a business event, retrieve the K most relevant knowledge fragments in the knowledge base embedding matrix E to form a knowledge vector set, and use cosine similarity to calculate the relevance value to screen K groups of combinations of multi-modal fusion vectors and knowledge fragments.
[0007] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: calculating the relevance value to screen the combination of the multi-modal fusion vector and the knowledge fragment, forming a gated input vector for fusion vector calculation, and obtaining an enhanced fusion vector sequence of the complete business process, including, Concatenate according to the combination of the multi-modal fusion vector and the knowledge fragment as the gated input vector, construct a knowledge gating fusion model based on a feed-forward fully connected neural network, and obtain a gating factor through a single-layer neural network for the gated input vector; Select a cross-entropy loss function to calculate the difference between the gating factor and the actual label of the training set, use the Adam optimizer for gradient descent optimization, update the parameters including weights and bias terms, and stop the iteration when the loss difference calculated during continuous iteration no longer decreases significantly; Use the gating factor to weigh the two source information for fusion vector calculation. At the same time, for each knowledge fragment, obtain the relationship strength value with the initial fusion vector in the form of a dot product, and perform Softmax processing to obtain the normalized weight. Determine the final enhanced fusion vector using all K candidate enhanced vectors and the corresponding normalized weights; Repeat the above operations for each event t in the business process, and finally obtain an enhanced fusion vector sequence of the complete business process.
[0008] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: determining the category label to which the business event belongs to generate an event directed graph, extracting a set of category abnormal state features, using the LCS algorithm to determine the path coincidence length, and screening highly similar aligned attribution paths, including, Determine the time point of the event according to different events in the enhanced fusion vector sequence, and determine the category label to which the business event belongs and the event site metadata, construct time nodes, and determine the node set of all business events; Specify the main line of the industry business process according to the process standard, construct a directed edge according to the standard transfer relationship from event i directly flowing to event j in the business process. For each adjacent business event, generate a candidate directed edge set according to the actual business occurrence order and the theoretically specified connection order between nodes, and combine the edges generated by the standard process and the edges collected from the actual business process to generate an event directed graph; According to the detailed domain tags given by the industry knowledge base, traverse all event nodes, and determine the nodes with the same domain tags and connected in the directed graph as business sub-processes, and construct a sub-process set for all business sub-processes; For each directed edge within the sub-process, extract the set of category anomaly status features, including timing features, state transition metrics, mutation amplitude, process flow standard violation labels, and real-time monitoring alarm flags. Statistically analyze each anomaly status feature of each directed edge and perform normalized summation as the attribution score for each directed edge, and statistically analyze the total attribution score of each sub-process; Based on the score threshold set by the empirical method, and take the sub-process with a score higher than the score threshold as the main attribution path. Based on the industry knowledge database, collect and organize the enterprise's archived paths over the years as template paths, and synchronously collect the anomaly status features of the nodes, steps, and edges in the attribution path to generate a standard path template library; For each attribution path obtained in the actual business and each template path in the industry knowledge database, encode the node order and the anomaly status features of each directed edge into a group of ordered feature sequence vectors, which are the structured feature sequences of the attribution path and the structured feature sequences of the expert knowledge template path respectively; For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the longest common subsequence LCS algorithm. According to the length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T, it is used as the two-dimensional array L; If all fields of the i-th step of the attribution path and the j-th step of the template path are exactly the same, determine the consistent coincidence length. Otherwise, select to exclude the nodes and edges that do not match the attribution path and the template path, and select the maximum coincidence length; Take the ratio of the determined LCS coincidence length to the total length of the path as the structural similarity score; Based on the sum of the mean and standard deviation of historical data as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, it is determined that the attribution path and the template path are highly similar and aligned.
[0009] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the generation of the direct index table of the attribution path and the template according to the anomaly status features includes, For the screened highly similar and aligned attribution paths, map the anomaly status features of the attribution path and the anomaly status feature items of the template path and store them in a structured format to form a direct index table of the attribution path and the template path.
[0010] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: for the attribution path where the statistical periods are not aligned, calculate the alignment deviation rate to generate a list of abnormal attribution paths, including Calculate the number of attribution paths in the attribution path within the statistical period that cannot be aligned with any template path, and the total number of attribution paths in the current period, calculate the alignment deviation rate of the current period. If the deviation rate is higher than the deviation threshold, output the list of abnormal attribution paths to form a feedback report.
[0011] As a preferred solution of the intelligent data analysis method based on the industry large model of the present invention, wherein: the collection of business event data includes Collect data of business events according to the industry business process, wherein the structured data includes production sensors, device logs, and business system data, the text data includes work records, fault reports, and original text data, and the graphic data includes on-site device monitoring and product image data.
[0012] In a second aspect, the present invention provides a system for the intelligent data analysis method based on the industry large model, including A multi-source data collection module responsible for collecting raw data; A multi-modal fusion module extracts features from the collected data and outputs a single multi-modal fusion vector; An embedding management module converts expert documents and historical knowledge fragments into a vectorized embedding library by a unified text encoder method; A gating empowerment module obtains a dynamic gating factor according to the gating input, realizes the adaptive weighted fusion between the fusion vector and the knowledge embedding, and outputs an enhanced fusion vector sequence for the full-process event; An abnormal feature encoding module generates a set of event nodes and directed edges to form a composite event directed graph; A template matching module traverses the sorted attribution paths and template paths, performs the LCS algorithm on the path sequence to calculate the structure coincidence length, and screens the key attribution paths with highly similar alignment; A deviation feedback module establishes a traceable direct index table and outputs a list of abnormal attribution paths and a report.
[0013] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the intelligent data analysis method based on the industry large model as described in the first aspect of the present invention is implemented.
[0014] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, any step of the intelligent data analysis method based on an industry large model as described in the first aspect of the present invention is implemented.
[0015] The beneficial effects of the present invention are as follows: Through the multi-modal feature fusion and knowledge fragment retrieval scheme, the unified expression and correlation mining of heterogeneous information such as structured data, text information, and images can be realized. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, achieve information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task. Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragments accumulated in history. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, thereby amplifying or suppressing the knowledge input as needed. The structural similarity score obtained by the ratio of the overlapping sequence length to the total length provides an objective and quantitative standard for the dynamic alignment of the attribution path and the template path, and suppresses the false matching caused by surface noise or local fluctuations. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of the intelligent data analysis method based on an industry large model in Embodiment 1.
[0018] Figure 2 It is a schematic structural diagram of the intelligent data analysis system based on an industry large model in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.
[0020] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that are mutually exclusive with other embodiments.
[0022] Embodiment 1, referring to From Figure 1 to Figure 2 , is the first embodiment of the present invention. This embodiment provides an intelligent data analysis method based on an industry large model, including the following steps: S1, Collect business event data, extract feature vectors according to the data type, and use the self-attention mechanism to calculate the scaled dot-product attention score matrix, perform weighted output of the multi-modal fusion vector, and construct an industry knowledge database to determine the embedding vector; Preferably, collecting business event data includes Collecting business event data according to the industry business process. The collected business event data includes structured data, text data, and graphic data. Among them, the structured data includes production sensors, device logs, and business system data, the text data includes work records, fault reports, and original text data, and the graphic data includes on-site device monitoring and product image data.
[0023] Through the collection of structured data, the production sensors, device logs, and business system data can comprehensively restore the production operation status and the dynamics of each link, facilitating subsequent fine-grained process analysis, anomaly location, and performance tracking. The collection of text data provides rich semantic support for the subjective descriptions, problem reports, and event details of the business site, helps to restore the on-site operations, fault occurrence, and handling process, enhances the automatic analysis and natural language parsing capabilities, and provides a solid foundation for knowledge extraction and root cause reasoning. The collection of graphic data enables the on-site device operation, product status, and visual anomalies to be intuitively reflected in the data flow, paving the data path for automatic visual inspection, anomaly warning, and process supervision, and realizing multi-angle and multi-modal perception of complex on-site situations. The synchronous acquisition of the three types of data greatly enhances the subsequent multi-modal correlation analysis, knowledge fusion, and high-dimensional data-driven attribution effect, making the transparency and intelligence of the business process reach a higher level.
[0024] Furthermore, extracting feature vectors according to the data type and constructing an industry knowledge database to determine the embedding vector includes Preprocessing the collected data, and respectively performing data dimensionality reduction on the processed structured data, text data, and graphic data; Among them, for structured data, a multi-layer perceptron (MLP) is used, and feature vectors are extracted through the ReLU non-linear activation function. For text data, a pre-trained BERT model is used for encoding, and the output text feature vectors are selected. For graphic data, a pre-trained ResNet50 model is used to extract features, and the output of its global average pooling layer is taken to obtain image feature vectors; The feature vectors of the three collected data are standardized by linear transformation. Based on historical experience, a unified dimensionality value D is determined, stacked to form a modal feature matrix, and the self-attention mechanism is used to calculate Query, Key, and Value, and the scaled dot-product attention score matrix is calculated, expressed as: ; where represents the attention score between the i-th and j-th modalities, represents the Query value of the i-th modality, represents the transpose of the Key value of the j-th modality, and D represents the unified dimensionality value; Apply softmax normalization to ensure that the dimension of softmax is set for each row, so that the attention distribution of each modality is reasonable, and the attention matrix is obtained and the Value is weighted, expressed as: ; ; where represents the attention weight between the i-th and j-th modalities, represents the attention score between the i-th and k-th modalities, R represents the total number of modalities, represents the fused vector of the i-th modality, represents the Value value of the j-th modality; Subsequently, average pooling is used to fuse the information of the three modalities into a single vector, and the multi-modal fused vector is output and serialized according to the industry business process to form a fused multi-modal vector sequence; Based on enterprise regulations and expert documents, an industry knowledge database is constructed. Each knowledge fragment in the knowledge base is uniformly embedded using the text encoder Enc to obtain the corresponding embedded vector, and the entire knowledge base corresponds to a set of embedded matrices; For the multi-modal fused vector of the business event, the K most relevant knowledge fragments are retrieved from the knowledge base embedding matrix E to form a knowledge vector set, and the cosine similarity is used to calculate the relevance value to screen the combinations of K groups of multi-modal fused vectors and knowledge fragments, where the value of K can be set according to historical experience.
[0025] Through multimodal feature fusion and knowledge fragment retrieval solutions, the unified expression and correlation mining of heterogeneous information such as structured data, text information and images can be achieved. Through preprocessing and dimensionality reduction, the effective compression and redundancy removal of various types of original data are ensured, and the downstream feature utilization rate is improved. Different types of data use neural network models suitable for their semantic structures for feature extraction. For example, MLP is used to obtain deep abstract expressions for structured data, BERT is used to capture context and semantics for text data, and ResNet50 is used to extract complex spatial structures for images. Each information advantage can be retained to the maximum extent. Linear mapping standardization solves the problems of heterogeneous scales and distributions of multimodal features, laying the foundation for subsequent unified fusion. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, achieve information complementarity, and automatically highlight the multimodal dimensions most relevant to the current task; softmax row-by-row normalization and weighted processing ensure that the calculated attention distribution is physically and business reasonable. This fusion method eliminates the uncertainty and inefficiency caused by subjective weighting or manual feature engineering, and achieves deep coupling for information from different sources; By adopting the average pooling method, a single vector with high information integration degree is finally obtained, which is automatically aligned with the business process to form a standard fusion feature sequence expression, which is convenient for subsequent business sequence modeling and traceability. After the knowledge base content is uniformly embedded in the text and represented by the vector, the relationship between structured business data and domain knowledge can be established, which greatly improves the accuracy of knowledge retrieval and the fit of context business. Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragment accumulated in history, effectively shortening the response time and manual search costs, ensuring the model's ability to discover and utilize knowledge in series, and the dynamic K value setting combines the actual business scale and historical case density to make the model neither overfitting nor flexible enough. The overall process greatly improves the depth and accuracy of business event understanding, and also provides a strong information foundation and significant global optimization effect for process compliance, automatic attribution and intelligent decision-making.
[0026] S2, calculate the correlation value to screen the combination of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process; Preferably, the correlation value is calculated to screen the combination of multimodal fusion vectors and knowledge fragments, and the gated input vector is formed to perform fusion vector calculation to obtain an enhanced fusion vector sequence of the complete business process, including: The concatenation is performed according to the combination of the multi-modal fusion vector and the knowledge fragment as the gated input vector. A knowledge gated fusion model is constructed based on a feed-forward fully connected neural network. The gated input vector is input into a single-layer neural network to obtain a gating factor, which is expressed as: ; where represents the gating factor for the t-th business event and its j-th retrieved knowledge fragment, represents the Sigmoid activation function, represents the transpose calculation of the weight vector, represents the concatenated vector of the multi-modal fusion vector of the t-th business event and the embedding of the j-th retrieved knowledge fragment, represents the bias term; The cross-entropy loss function is selected to calculate the difference between the gating factor and the actual label of the training set. The Adam optimizer is used for gradient descent optimization. The updated parameters include the weight and the bias term. The iteration stops when the loss difference calculated in the continuous iteration process no longer decreases significantly; The gating factor is used to weigh the two sources of information for the fusion vector calculation, which is expressed as: ; where represents the candidate enhanced vector after the fusion of the t-th business event and the j-th retrieved knowledge fragment, represents the multi-modal fusion vector of the t-th business event, represents the embedding vector of the j-th knowledge fragment retrieved by the t-th business event; At the same time, for each knowledge fragment, the relationship strength value with the initial fusion vector is obtained in the form of a dot product, which is expressed as: ; where represents the relationship strength value between the t-th business event and the j-th retrieved knowledge fragment, represents the transpose calculation of the multi-modal fusion vector; And Softmax processing is performed to obtain the normalized weight. The final enhanced fusion vector is determined by all K candidate enhanced vectors and the corresponding normalized weights; The above operations are repeated for each event t in the business process, and finally an enhanced fusion vector sequence of the complete business process is obtained.
[0027] When the multi-modal features of business events are combined with knowledge fragment embeddings, it is not simply a direct superposition. Instead, through a gating mechanism, it selectively determines the contribution of each type of information. This approach utilizes the non-linear modeling ability of neural networks to automatically mine the most valuable correlation relationships between business events and knowledge fragments in different scenarios, avoiding the subjective biases of static or artificial prior weighting. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, thereby amplifying or suppressing the knowledge input as needed, making the fusion vector more in line with the actual business scenario and empirical rules; Through the training set labels and cross-entropy loss, the model continuously learns from past data what fusion ratios and methods can bring optimal decisions, so that the finally obtained enhanced fusion vector contains both the most core multi-modal information of the current business and fully integrates the essence of historical knowledge, facilitating downstream analysis and intelligent reasoning; The dot product relationship strength and Softmax normalization automatically rank the importance of different candidate fusion vectors, and the model can effectively filter redundant or low-correlated knowledge, focus on and utilize the most valuable information fragments, and improve the pertinence and effectiveness of knowledge-assisted reasoning; As each event in the business process repeats the above fusion steps, the sequence of enhanced fusion vectors for the entire process can more precisely capture the dynamic relationships between each link in the event chain and the knowledge base and the latest business situation; The final sequence can not only reflect the overall business evolution but also have the profound guiding significance of industry knowledge, providing a solid and logically interpretable information foundation for subsequent tasks such as complex business attribution, process analysis, and risk identification.
[0028] S3. Determine the category label to which the business event belongs, generate an event directed graph, extract the set of category abnormal state features, sort out the template path and the attribution path of the actual business, use the LCS algorithm to determine the path coincidence length, and screen the highly similar and aligned attribution paths; Preferably, determining the category label to which the business event belongs, generating an event directed graph, extracting the set of category abnormal state features, using the LCS algorithm to determine the path coincidence length, and screening the highly similar and aligned attribution paths includes Determine the time point of the event according to different events in the enhanced fusion vector sequence, and determine the category label to which the business event belongs (such as process stage number, equipment type, alarm type, etc., taken from actual business data) and the event site metadata (such as timestamp, operator, production line ID, physical location, relevant process parameters, all from the available structured data in the business system), construct time nodes, and determine the node set of all business events; Identify the main line of the industry business process clearly according to the process standard. Construct directed edges based on the standard transfer relationship that event i directly flows to event j in the business process. For each adjacent business event, generate a candidate directed edge set respectively according to the actual business occurrence order and the theoretical specified connection order between nodes. For the standard process, the directed edges are generated based on the directly connected processes, links or business instructions stipulated by the industry regulations. For the actual event flow, the direct transfer relationship between events is automatically generated according to the occurrence order of adjacent events in the real business, and detailed attributes including category labels and on-site metadata are attached to each generated directed edge; Merge the edges generated by the standard process and the edges collected from the actual business process to generate an event directed graph, which contains both the theoretically expected transfer relationship and the actual transfer that really occurred on-site. In this way, whether it is process supervision, anomaly detection, root cause analysis or dynamic process optimization, the overall situations of "ideal" and "real" can be considered simultaneously; Traverse all event nodes according to the detailed domain labels given by the industry knowledge base. Determine the nodes with the same domain label and connected in the directed graph as business sub-processes, and construct a sub-process set for all business sub-processes; Extract the set of category anomaly status features for each directed edge within the sub-process, including timing features (the time interval between the occurrence of consecutive events), state jump metrics (the mutation amplitude of the core parameters of the device (temperature, pressure, load, etc.)), process flow standard violation labels (the program automatically compares whether the transfer node pair skips the necessary process, whether it flows backward, and whether there are illegal or non-compliant links), real-time monitoring alarm flags (faults, over-limits or alarm events generated by the ERP / MES / sensor system). Statistically analyze each anomaly status feature of each directed edge and perform normalized summation as the attribution score of each directed edge, and statistically analyze the total attribution score of each sub-process; Set a score threshold based on the empirical method, and regard the sub-processes with scores higher than the score threshold as the main attribution paths. Based on the industry knowledge database, collect and organize the enterprise's archived paths over the years as template paths, including classic defect root cause paths, typical problem flows, expert-reviewed anomaly propagation chains, and business standard propagation templates. Synchronously collect the anomaly status features of the nodes, steps and edges in the attribution paths to generate a standard path template library; For each attribution path obtained in the actual business and each template path in the industry knowledge database, encode the node order and the anomaly status features of each directed edge into a group of ordered feature sequence vectors, which are the structured feature sequences of the attribution path and the structured feature sequences of the expert knowledge template path respectively, and can be encoded and stored using a JSON array; For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the Longest Common Subsequence (LCS) algorithm. The length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T is used as the two-dimensional array L; If all fields of the i-th step of the attribution path and the j-th step of the template path are exactly the same, determine the consistent overlap length. Otherwise, select and exclude the node edges where the attribution path and the template path do not match, and select the maximum overlap length, which is expressed as: ; ; where represents the consistent overlap length between the first i steps of the attribution path and the first j steps of the template, represents the length of the longest common subsequence where the previous pair has been completely compared, represents the maximum overlap length between the first i steps of the attribution path and the first j steps of the template, and represent the lengths of the longest common subsequences when only looking at the first i - 1 steps of the attribution path and the first j steps of the template path, and when only looking at the first i steps of the attribution path and the first j - 1 steps of the template path, respectively; Use the ratio of the determined LCS overlap length to the total length of the path as the structural similarity score; Based on the sum of the mean and standard deviation of historical data as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, it is determined that the attribution path and the template path are highly similar and aligned.
[0029] Business events are divided into strict node sets, and each node carries traceable business metadata and event categories, ensuring clear logic and true data in the description of the event chain; directed edges are extracted from the theoretical process and the actual execution respectively and merged into a unified graph structure, enabling subsequent process optimization, deviation identification, and root cause tracing to be carried out based on the overall picture of business logic, avoiding the blind spots of a single perspective and providing the possibility of dynamic comparison between the ideal and the reality.
[0030] The sub - process division is completed through the linkage of domain labels and node connectivity, which can accurately aggregate similar business links and automatically adapt to on - site process changes, reducing the deviation that may be caused by manual division. For each directed edge, multi - dimensional abnormal features such as time, status, process, and alarm are statistically analyzed, and the normalized sum is an objectively measurable attribution score. Further, the abnormal degree of each sub - process is quantified, achieving a quick focus on the core abnormal chain that is most likely to affect the business result.
[0031] The score threshold set by the empirical method can avoid high-noise and low-correlation processes from occupying analysis resources, enabling the analysis to focus on the truly abnormally intense process segments that may pose business risks; adopting a standardized template path library and synchronous feature vectorized storage allows the actual attribution results to form a highly consistent structured comparison with the typical problems and experiences precipitated in history. This comparison neither relies on manual experience nor misses new abnormal patterns in business processes, fundamentally enhancing the applicability and real-time performance of automated intelligent analysis.
[0032] After the LCS algorithm comparison, the structural similarity score obtained from the ratio of the length of the overlapping sequence to the total length provides an objective and quantitative standard for the automatic "classification" and dynamic alignment of the attribution path and the template path, and suppresses the pseudo-matching caused by surface noise or local fluctuations. Using the threshold synthesized from the historical mean and standard deviation as the final criterion enables the identification of the main attribution path and the highly reliable automatic archiving to have a data-driven adaptive mechanism, truly empowering process anomaly localization and knowledge closed-loop improvement in large-scale industrial scenarios.
[0033] S4. Generate a direct index table of the attribution path and the template based on the abnormal state characteristics. Calculate the alignment deviation rate for the attribution paths that could not be aligned during the statistical period, and generate a list of abnormal attribution paths. Preferably, generating a direct index table of the attribution path and the template based on the abnormal state characteristics includes For the highly similar aligned attribution paths selected, map the abnormal state characteristics of the attribution path to the abnormal state characteristic items of the template path and store them in a structured format to form a direct index table of the attribution path and the template path.
[0034] By mapping the abnormal state characteristics of the highly similar aligned attribution paths and the template paths one by one and storing them in a structured manner, the precise information association between the attribution result and the standard template is realized, ensuring that all abnormal details of the attribution result can be accurately retrieved and searched by subsequent automated systems, operation and maintenance personnel, or decision-making systems, improving the data reuse and query efficiency. The direct index relationship between the attribution path and the template path brings about the seamless connection between business attribution and historical knowledge, facilitating the statistical attribution times, monitoring the occurrence trend of specific abnormal types, and discovering the high-incidence links of process optimization, providing a solid structural foundation for the subsequent anomaly tracking, attribution explanation, compliance audit, and continuous expansion of the knowledge base of the enterprise, and greatly enhancing the automation and controllability of management and analysis.
[0035] Furthermore, calculating the alignment deviation rate for the attribution paths that could not be aligned during the statistical period and generating a list of abnormal attribution paths includes Save all the node sequences, edge abnormal state characteristics of the attribution path, and the data of its matching structural similarity score to generate a log file. Count the number of attribution paths in the attribution path during the statistical period that fail to align with any template path, as well as the total number of attribution paths in this period, calculate the alignment deviation rate of the current period. If the deviation rate is higher than the deviation threshold (set according to historical experience), output a list of abnormal attribution paths to form a feedback report.
[0036] By recording in detail the node sequence of each attribution path, the abnormal state characteristics of all edges, and the specific structure similarity score and generating logs, there is a source of original data to support subsequent traceability, playback, and problem review. Ensure that each attribution analysis can be verified and reused conveniently. Periodically count the number of unaligned attribution paths and the total number of attribution paths, and calculate the alignment deviation rate accordingly. Through quantitative indicators, it can dynamically reflect the coverage of the current knowledge template for the actual business complexity, and timely expose the "blind spots" in the knowledge base or the trend of new abnormal patterns in the business.
[0037] When the deviation rate exceeds the predetermined threshold, the system can automatically screen out all abnormal attribution paths, form a feedback report, and accurately prompt the management and analysis team to pay attention to possible production line changes, equipment aging, or new problems, supporting rapid response and the expansion and optimization of knowledge templates. This process realizes the automatic archiving, statistics, and reporting of abnormal information, ensuring that enterprises can continuously optimize business processes and knowledge assets in a closed-loop manner, effectively improving problem warning and adaptive adjustment capabilities.
[0038] This embodiment also provides a system for an intelligent data analysis method based on an industry large model, including A multi-source data acquisition module responsible for collecting raw data from structured data sources, text data sources, and graphic data sources; A multi-modal fusion module that extracts features from the collected data, performs standardized mapping and feature alignment, and dynamically models the information relationship between multi-modalities through a self-attention mechanism, outputting a single multi-modal fusion vector; An embedding management module that converts expert documents and historical knowledge fragments into a vectorized embedding library using a unified text encoder method; A gated weighting module that obtains a dynamic gating factor according to the gated input, realizes the adaptive weighted fusion between the fusion vector and the knowledge embedding, and outputs an enhanced fusion vector sequence for the full-process event; An abnormal feature encoding module that generates a set of event nodes and directed edges to form a composite event directed graph; A template matching module that traverses the sorted attribution paths and template paths, calculates the structure overlap length of the path sequence using the LCS algorithm, and combines the attribution score and the structure similarity score to screen out the key attribution paths with highly similar alignment; A deviation feedback module that establishes a traceable direct index table and outputs a list and report of abnormal attribution paths.
[0039] This embodiment also provides a computer device applicable to the scenario of an intelligent data analysis method based on an industry large model, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent data analysis method based on the industry large model proposed in the above embodiment.
[0040] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0041] This embodiment also provides a storage medium on which a computer program is stored, and when the program is executed by a processor, it implements the intelligent data analysis method based on the industry large model proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0042] In summary, through the multi-modal feature fusion and knowledge fragment retrieval scheme, the present invention can achieve the unified expression and correlation mining of heterogeneous information such as structured data, text information, and images. Through the self-attention mechanism, the model can adaptively adjust the attention of each modal feature according to the business scenario, realize information complementarity, and automatically highlight the multi-modal dimensions most relevant to the current task. Through cosine similarity screening, each business event fusion vector is automatically associated with the most similar and most valuable knowledge fragments accumulated in history. The introduction of the gating factor enables the model to dynamically adjust the fusion strategy according to the matching situation between the current business event and the retrieved knowledge fragments, thereby amplifying or suppressing the knowledge input as needed. The structural similarity score obtained by the ratio of the overlapping sequence length to the total length provides an objective and quantitative standard for the automatic "classification" and dynamic alignment of the attribution path and the template path, and suppresses the false matching caused by surface noise or local fluctuations.
[0043] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. Intelligent data analysis method based on industry big model, characterized by: include: Collect business event data, extract feature vectors based on data types, and use the self-attention mechanism to calculate the scaled dot product attention score matrix, perform weighted output of multimodal fusion vectors, and build an industry knowledge database to determine the embedding vector; Calculate the correlation value to filter the combination of multimodal fusion vectors and knowledge fragments, form a gated input vector for fusion vector calculation, and calculate the relationship strength value between the knowledge fragment and the fusion vector to obtain an enhanced fusion vector sequence of the complete business process; Determine the category label of the business event to generate an event directed graph, extract the category abnormal state feature set, organize the template path and the attribution path of the actual business, use the LCS algorithm to determine the path overlap length, and screen the attribution paths with high similarity and alignment; Generate a direct index table of attribution paths and templates based on abnormal status characteristics, calculate the alignment deviation rate of attribution paths that fail to align during the statistical period, and generate a list of abnormal attribution paths.
2. The intelligent data analysis method based on the industry big model as claimed in claim 1, characterized in that: The feature vector is extracted according to the data type, and the industry knowledge database is constructed to determine the embedding vector, including: Collecting business event data including structured data, text data and graphic data, preprocessing the collected data, and performing data dimension reduction on the processed structured data, text data and graphic data respectively; For structured data, a multi-layer perceptron (MLP) is used to extract feature vectors through the ReLU nonlinear activation function. The pre-trained BERT model is used to encode text data and the output text feature vector is selected. For graphic data, the pre-trained ResNet50 model is used to extract features and the output of the global average pooling layer is used to obtain the image feature vector. The feature vectors of the three collected data are standardized by linear change, stacked to form a modal feature matrix, and the self-attention mechanism is used to calculate the query, key, and value, and the scaled dot product attention score matrix is calculated. Softmax normalization is applied to obtain the attention matrix and weight the value; Then, average pooling is used to fuse the information of the three modalities into a single vector, and the multimodal fusion vector is output and serialized according to the industry business process to form a fused multimodal vector sequence; Build an industry knowledge database based on enterprise regulations and expert documents, and use the text encoder Enc to embed each knowledge fragment in the knowledge base to obtain the corresponding embedding vector. The entire knowledge base corresponds to a set of embedding matrices. For the multimodal fusion vector of business events, the most relevant K knowledge fragments are retrieved from the knowledge base embedding matrix E to form a knowledge vector set, and the cosine similarity is used to calculate the correlation value to screen the combination of K groups of multimodal fusion vectors and knowledge fragments.
3. The intelligent data analysis method based on the industry big model as claimed in claim 2 is characterized by: The calculation of the correlation value screens the combination of the multimodal fusion vector and the knowledge fragment to form a gated input vector for fusion vector calculation, and obtains an enhanced fusion vector sequence of a complete business process, including: According to the combination of multimodal fusion vector and knowledge fragments, they are spliced as gated input vectors, a knowledge gated fusion model is constructed based on a feedforward fully connected neural network, and a single-layer neural network of the gated input vector is used to obtain a gating factor; Select the cross entropy loss function to calculate the difference between the gating factor and the actual label of the training set, use the Adam optimizer for gradient descent optimization, update the parameters including weights and bias terms, and stop the iteration if the calculated loss difference no longer decreases significantly during the continuous iteration process; The gating factor is used to weigh the two source information to calculate the fusion vector. At the same time, the dot product form is used to obtain the relationship strength value with the initial fusion vector for each knowledge fragment, and the Softmax processing is performed to obtain the normalized weight. The final enhanced fusion vector is determined using all K candidate enhanced vectors and the corresponding normalized weights. The above operation is repeated for each event t in the business process, and finally the enhanced fusion vector sequence of the complete business process is obtained.
4. The intelligent data analysis method based on the industry big model as claimed in claim 3 is characterized by: The process of determining the category label to which the business event belongs generates an event directed graph, extracts a category abnormal state feature set, uses the LCS algorithm to determine the path overlap length, and screens highly similar and aligned attribution paths. include, Determine the time point of the event according to different events in the enhanced fusion vector sequence, determine the category label of the business event and the event scene metadata, construct the time node, and determine the node set of all business events; Obtain the main line of the industry business process according to the process standard table, construct directed edges according to the standard flow relationship from event i to event j in the business process, generate candidate directed edge sets for each adjacent business event based on the actual business occurrence sequence and the theoretical node connection sequence, and merge the edges generated by the standard process and the edges collected by the actual business process to generate an event directed graph; According to the detailed domain labels given by the industry knowledge base, all event nodes are traversed, and nodes with the same domain labels and connected in the directed graph are determined as business sub-processes, and all business sub-processes are constructed into sub-process sets; For each directed edge in the sub-process, a set of abnormal state features is extracted, including time series features, state jump metrics, mutation amplitude, process standard violation labels, and real-time monitoring alarm signs. Each abnormal state feature of each directed edge is counted and normalized and summed as the attribution score of each directed edge, and the total attribution score of each sub-process is counted. Based on the score threshold set by the empirical method, the sub-process with a score higher than the score threshold is taken as the main attribution path. Based on the industry knowledge database, the archived paths of the enterprise over the years are collected and sorted as template paths. The abnormal state characteristics of the node step edges in the attribution path are simultaneously collected to generate a standard path template library. For each attribution path obtained in actual business and each template path in the industry knowledge database, the node order and the abnormal state characteristics of each directed edge are encoded into a set of ordered feature sequence vectors, which are the structured feature sequence of the attribution path and the structured feature sequence of the expert knowledge template path respectively; For each path to be attributed, traverse the template library, compare the feature sequences one by one, and use the longest common subsequence LCS algorithm to calculate the length of the longest common subsequence between the first i steps of the attribution path P and the first j steps of the template path T as the two-dimensional array L; If all fields of the attribution path step i and the template path step j are completely consistent, the consistent overlap length is determined. Otherwise, the node edges that do not match the attribution path and the template path are excluded, and the maximum overlap length is selected; The ratio of the determined LCS coincidence length to the total length of the path is taken as the structural similarity score; The sum of the mean and standard deviation of historical data is used as the structural similarity threshold. If the structural similarity score is greater than or equal to the structural similarity threshold, the attribution path is judged to be highly similar and aligned with the template path.
5. The intelligent data analysis method based on the industry big model as claimed in claim 4 is characterized by: The direct index table of attribution paths and templates is generated according to the abnormal state characteristics, include, For the screened highly similar aligned attribution paths, the abnormal state features of the attribution paths are mapped with the abnormal state feature items of the template paths and stored in a structured format to form a direct index table of the attribution paths and template paths.
6. The intelligent data analysis method based on the industry big model as claimed in claim 5, characterized in that: The attribution paths that failed to align during the statistical period calculate the alignment deviation rate and generate a list of abnormal attribution paths, including: The number of attribution paths that fail to align with any template path in the statistical cycle, as well as the total number of attribution paths in this cycle, are counted. The alignment deviation rate of the current cycle is calculated. If the deviation rate is higher than the deviation threshold, a list of abnormal attribution paths is output to form a feedback report.
7. The intelligent data analysis method based on the industry big model as claimed in claim 6 is characterized by: The collected business event data includes: Data collection of business events is carried out according to the industry business process, where structured data includes production sensors, equipment logs and business system data, text data includes work records, fault reports and original text data, and graphic data includes on-site equipment monitoring and product image data.
8. A system for intelligent data analysis method based on industry big model, based on the intelligent data analysis method based on industry big model according to any one of claims 1 to 7, characterized in that: include, Multi-source data acquisition module, responsible for collecting raw data; The multimodal fusion module extracts features from the collected data and outputs a single multimodal fusion vector; The embedding management module converts expert documents and historical knowledge fragments into vectorized embedding libraries with a unified text encoder approach; The gated weighting module obtains dynamic gated factors according to gated inputs, realizes adaptive weighted fusion between fusion vectors and knowledge embedding, and outputs enhanced fusion vector sequences for full-process events; The abnormal feature encoding module generates a set of event nodes and directed edges to form a composite event directed graph; The template matching module traverses the sorted attribution paths and template paths, applies the LCS algorithm to the path sequences to calculate the structural overlap length, and selects the key attribution paths with high similarity and alignment; The deviation feedback module establishes a traceable direct index table and outputs an abnormal attribution path list and report.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent data analysis method based on the industry big model as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent data analysis method based on the industry big model described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video emotion classification method based on gating fusion and multi-task learning
CN115203409A
Text and image-based multi-modal named entity recognition system
CN118052235A
Video analysis method based on deep learning platform and related equipment
CN118918522A
Inspection method, device and equipment based on digital airspace system and medium
CN119989283A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Cited By
Lean operation method based on production scheduling optimization and equipment OEE monitoring
CN120525140A
Adaptive scene intelligent interaction system based on AI
CN120704532A
AI-based adaptive scene intelligent interaction system
CN120704532B
Multi-data fusion method and system based on edge calculation
CN120974414A
A Multi-Data Fusion Method and System Based on Edge Computing
CN120974414B