Collection and analysis method and system for industrial multi-source heterogeneous data

By extracting operating condition behavior features, constructing an operating condition state transition model, and generating a unified time benchmark, the problems of timestamp deviation and data alignment of multi-source heterogeneous data were solved, achieving accurate alignment and efficient analysis of multi-source data, and improving the level of data intelligence and automation in industrial production.

CN122087261APending Publication Date: 2026-05-26WUXI KUAI INFORMATION SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI KUAI INFORMATION SCI & TECH
Filing Date
2025-12-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing industrial data analysis methods struggle to effectively handle multi-source heterogeneous data, particularly due to difficulties in data alignment caused by timestamp discrepancies, differences in data types and structures, complex cross-data source feature extraction, and the inability to dynamically adjust processing methods, thus failing to achieve efficient and real-time data acquisition and analysis.

Method used

By extracting operating condition behavior features, performing feature similarity matching across data sources, constructing an operating condition state transition model, generating a unified time benchmark, and adjusting data mapping relationships when mapping conflicts occur, accurate alignment of multi-source data is achieved.

Benefits of technology

It improves the accuracy and flexibility of multi-source data analysis, can adaptively handle new changes in the production process, optimizes data analysis quality, and enhances the level of data intelligence and automation in industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087261A_ABST
    Figure CN122087261A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of industrial data analysis, and particularly discloses an industrial multi-source heterogeneous data acquisition and analysis method and system, and the method comprises the steps: obtaining industrial data from a plurality of data sources, forming a data sequence, extracting working condition behavior characteristics, and generating a working condition change event sequence. Then, on the basis of the feature sequences, cross-data-source similarity matching is executed, a matching relation set is obtained, and a working condition state transition model is constructed; based on this, a unified time reference is generated and the data is mapped to a logical time interval on the reference. According to the method, the problem of inconsistent timestamps among data sources can be effectively solved, accurate alignment of multi-source data is ensured, and adjustment is performed according to the matching confidence coefficient during mapping conflict. The method is suitable for industrial data analysis and intelligent manufacturing, and the accuracy and efficiency of data fusion and processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial data analysis technology, specifically to a method and system for acquiring and analyzing multi-source heterogeneous industrial data. Background Technology

[0002] With the rapid development of the Industrial Internet and intelligent manufacturing, industrial site monitoring and data acquisition have gradually shifted from single data sources to the integration of multi-source heterogeneous data. In industrial production, various equipment, sensors, and management systems generate a wide variety of industrial data, with diverse sources and formats. To effectively support industrial big data analytics and enable real-time monitoring and optimization of production processes, efficiently integrating, processing, and analyzing heterogeneous data from different data sources has become a pressing technical challenge.

[0003] Currently, many industrial data analysis methods rely on a single data source and lack the integration and unified processing of multi-source heterogeneous data. Traditional industrial data acquisition and analysis methods often face the following problems: First, discrepancies in timestamps between different data sources make data alignment difficult; second, differences in the type and structure of heterogeneous data sources complicate feature extraction and matching across data sources; third, existing methods are mostly limited to a single data source or rule-based static processing, unable to dynamically adjust and update data processing methods, making it difficult to achieve efficient and real-time data acquisition and analysis in complex production environments.

[0004] Furthermore, existing industrial data processing methods often lack unified standards for extracting operational behavior features across data sources and aligning time-series data when dealing with multi-source heterogeneous data. With the increasing variability of production processes, the adaptability and accuracy of existing methods are limited, making it difficult to effectively handle nonlinear relationships and complex interactions between multi-source data. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for acquiring and analyzing multi-source heterogeneous industrial data, so as to solve the problems mentioned in the background art.

[0006] According to one aspect of the present invention, a method for acquiring and analyzing multi-source heterogeneous industrial data is provided, applied to a chip interconnect system including a master control device and at least one slave device, characterized by comprising the following method steps: Acquire industrial data from multiple data sources, wherein the industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and form a data sequence corresponding to each data source based on the data source identifier; For each of the data sequences, extract the operating condition behavior features that characterize the changes in operating conditions, and generate the operating condition behavior feature sequence corresponding to each data source. The operating condition behavior features include operating condition change events, and the operating condition change events include at least feature types and corresponding original timestamps. Based on the characteristic sequences of each working condition, feature similarity matching across data sources is performed to obtain a set of matching relationships for working condition change events across data sources, and a matching confidence level is determined for the matching relationships in the set of matching relationships. A working condition state transition model is constructed based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. A unified time reference is generated based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events, and does not use the original timestamp of any single data source as the reference for cross-data source alignment. Data sequences from various data sources are mapped to logical time intervals on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when mapping conflicts exist, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

[0007] Preferably, the step of extracting operating condition behavior features that characterize the change in operating conditions for each of the data sequences includes calculating the rate of change or difference component of adjacent sampling points for continuous numerical data, and determining the corresponding operating condition change event when the rate of change or difference component meets a preset threshold condition. The system detects state value transitions for status-type or flag-type data and determines the corresponding operating condition change event when the state value transition is detected.

[0008] Preferably, the operating condition change event further includes a characteristic intensity or amplitude, wherein the characteristic intensity or amplitude is determined by the change amplitude, rate of change, or step amplitude of the corresponding continuous numerical data.

[0009] Preferably, the step of performing cross-data source feature similarity matching based on each of the operating condition behavior feature sequences includes: using dynamic time warping to non-linearly align the operating condition behavior feature sequences of different data sources, calculating the similarity based on the alignment path, and generating the matching relationship in the matching relationship set when the similarity meets the preset matching threshold condition.

[0010] Preferably, the step of constructing the working condition state transition model based on the matching relationship set includes: statistically analyzing the state transition frequency based on the working condition state sequence in historical data, determining the state transition probability between the working condition state nodes, and using the state transition probability as a parameter of the working condition state transition model.

[0011] Preferably, generating a unified time reference based on the matching relationship set and the working condition state transition model includes: aggregating working condition change events across data sources into event clusters based on the matching relationship set; sorting the event clusters according to the constraints of the working condition state transition model; and determining the order of the logical time intervals based on the sorting results, thereby generating the unified time reference.

[0012] Preferably, the mapping conflict includes candidate mappings for the same data record corresponding to multiple logical time intervals. Adjusting the mapping conflict includes: assigning the data record to the logical time interval corresponding to the candidate mapping with the highest matching confidence, or weighting the candidate mappings according to the matching confidence.

[0013] Preferably, it also includes dynamic updates, specifically, when a preset trigger condition is met, updating at least one of the matching relationship set or the working condition state transition model, and updating the unified time base and the mapping relationship accordingly; wherein, the preset trigger condition includes the switching frequency of the working condition change event meeting a threshold condition, the matching confidence meeting a threshold condition, or the latency anomaly of at least one data source meeting a threshold condition.

[0014] In another aspect, this application also provides a system for acquiring and analyzing multi-source heterogeneous industrial data, comprising: An industrial data acquisition module is used to acquire industrial data from multiple data sources. The industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and forms a data sequence corresponding to each data source based on the data source identifier. The working condition behavior feature extraction module is used to extract working condition behavior features that characterize working condition changes for each of the data sequences and generate working condition behavior feature sequences corresponding to each data source. The working condition behavior features include working condition change events, and the working condition change events include at least feature types and corresponding original timestamps. The matching module is used to perform feature similarity matching across data sources based on the sequence of behavioral characteristics of each working condition, to obtain a set of matching relationships for working condition change events across data sources, and to determine the matching confidence for the matching relationships in the set of matching relationships; The model building module is used to build a working condition state transition model based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. The time reference generation module is used to generate a unified time reference based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events and does not use the original timestamp of any single data source as the reference for cross-data source alignment. The mapping and alignment module is used to map the data sequences of each data source to the logical time interval on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when there is a mapping conflict, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

[0015] This application also provides a computer device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the industrial multi-source heterogeneous data acquisition and analysis method as described above.

[0016] This invention achieves precise alignment of multi-source data through a unified time benchmark, effectively solving the time alignment problem between different data sources. Specifically, by using cross-data source feature similarity matching and dynamic time warping, it overcomes the shortcomings of traditional methods in handling data source heterogeneity and timestamp discrepancies, thereby improving the accuracy and reliability of data fusion. Furthermore, by extracting operational behavior features and constructing operational state transition models, this invention enables multi-source data analysis to go beyond simple data merging, uncovering key features in the production process from the perspective of operational changes, thus optimizing the quality of data analysis. A dynamic update mechanism ensures that the method can adaptively handle new changes occurring in the production process, further enhancing the flexibility and real-time performance of the data analysis system. Overall, the method of this invention has significant advantages in accuracy, efficiency, and adaptability in processing multi-source heterogeneous data, and can be widely applied in fields such as intelligent manufacturing and industrial big data analysis, promoting the improvement of data intelligence and automation levels in industrial production. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure.

[0019] Figure 2 A flowchart of a method for collecting and analyzing multi-source heterogeneous industrial data provided in this embodiment of the disclosure.

[0020] Figure 3 A schematic diagram of the time base generation process provided in the embodiments of this disclosure.

[0021] Figure 4 This is a schematic diagram of an industrial multi-source heterogeneous data acquisition and analysis system provided in an embodiment of this disclosure.

[0022] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] It should be noted that all user information (including but not limited to user device information, user personal information, object information corresponding to device usage data, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, device usage data, etc.) involved in all embodiments of this disclosure are information and data authorized by the user or fully authorized by all parties.

[0025] This embodiment discloses a method for acquiring and analyzing multi-source heterogeneous industrial data. This method can be executed by an industrial data acquisition and analysis system and is applicable to industrial settings where at least two data sources operate in parallel. Please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this disclosure. The data sources include, for example, control systems, production management systems, and energy consumption or auxiliary systems. The original timestamps of the industrial data output from each data source can originate from the device's local clock, server system clock, host computer or database server system clock, and manual data entry time, and can be affected by clock drift, network jitter, and cache latency.

[0026] In this embodiment, the data source is a system or subsystem that generates and outputs industrial data, distinguished by a data source identifier. Industrial data refers to records collected by the data source, including at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier. The data identifier is an identifier field used to identify a location, variable, or event type. The data source identifier is an identifier field used to identify the data source to which the data belongs. A data sequence is a collection of industrial data organized by original timestamps within the same data source. An operating condition is a collection of the operating states of an industrial process over a period of time. Operating condition behavior characteristics are a set of feature types used to characterize changes in operating conditions. An operating condition change event is an event entity extracted from the operating condition behavior characteristics, including at least a feature type and a corresponding original timestamp; in some embodiments, it also includes a feature intensity or amplitude field.

[0027] A working condition behavior feature sequence is a sequence of working condition change events within the same data source. A matching relationship is the correspondence formed when two working condition change events across data sources are determined to be different observations of the same actual working condition event. A matching relationship set is a collection of multiple matching relationships. Matching confidence is a numerical field used to characterize the reliability of a matching relationship. A working condition state node is a node used to represent a state in a working condition state transition model. A state transition relationship is the adjacent transition connection relationship between working condition state nodes. A working condition state transition model is a model that includes working condition state nodes and state transition relationships, and in some embodiments includes state transition probabilities. An event cluster is a set of working condition change events across data sources obtained by aggregating the matching relationship set. A unified time base is a logical time reference determined by the sorting results of working condition change events or event clusters, and does not use the original timestamp of any single data source as the cross-data source alignment base. A logical time interval is an interval unit on a unified time base. A mapping conflict occurs when industrial data records are mapped to a logical time interval and multiple candidate assignments or inconsistent assignments occur. Dynamic updates are the process of updating at least one of the matching relationship set or the working condition state transition model when a preset trigger condition is met, and updating the unified time base and mapping relationship accordingly.

[0028] like Figure 2 As shown in the figure, this invention discloses a method for acquiring and analyzing multi-source heterogeneous industrial data, including the following steps: S1, acquire industrial data from multiple data sources, wherein the industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and form a data sequence corresponding to each data source based on the data source identifier; S2, for each of the data sequences, extract the operating condition behavior features that characterize the changes in operating conditions, and generate the operating condition behavior feature sequence corresponding to each data source, wherein the operating condition behavior features include operating condition change events, and the operating condition change events include at least feature types and corresponding original timestamps. S3, based on the characteristic sequences of each working condition, perform feature similarity matching across data sources to obtain a set of matching relationships for working condition change events across data sources, and determine the matching confidence for the matching relationships in the set of matching relationships; S4. Construct a working condition state transition model based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. S5. A unified time reference is generated based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events and does not use the original timestamp of any single data source as the reference for cross-data source alignment. S6, map the data sequences from each data source to the logical time interval on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when there is a mapping conflict, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

[0029] In some embodiments, for step S1, the industrial data acquisition module acquires industrial data from at least two data sources. Each piece of industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier. The industrial data is grouped according to the data source identifier to form a data sequence corresponding to each data source. Within the same data source, the data sequence is sorted by the original timestamp to determine the processing order within the same source; this sorting is not used as a cross-data source alignment benchmark. Optionally, the industrial data acquisition module configures different acquisition interfaces and parsing rules for different data sources to ensure consistency of the data sequences at the structural field level.

[0030] This step involves collecting industrial data from multiple data sources and forming a data sequence. This ensures that the input data from each data source is clearly distinguishable in terms of time and identifier, facilitating subsequent data alignment and analysis, and providing a unified raw data structure for subsequent processing.

[0031] In some embodiments, for step S2, operating condition behavior features are extracted from the data sequences of each data source, and an operating condition behavior feature sequence is generated. Exemplarily, the operating condition behavior features include at least one of the following: changes in equipment start / stop status, step changes in valve opening or actuator position, sudden changes in power, current, steam, or energy consumption, changes in process status flags, and operating mode switching signals. Operating condition change events are output in the form of event records, which at least include a feature type field and a corresponding original timestamp field; in some embodiments, the event record also includes a feature intensity or amplitude field and an event direction field.

[0032] In this context, characteristic intensity or amplitude refers to a numerical indicator used to characterize the degree or intensity of change in operating condition events. For continuous numerical data, characteristic intensity typically reflects the magnitude of data change, such as the rate of change or the amount of change, representing the magnitude of change of a variable within a specific time period. For example, a temperature increase from 30°C to 50°C represents a change of 20°C. For discrete-state data, characteristic intensity may be related to the frequency or magnitude of state transitions, representing the intensity of change from one state to another. The magnitude of characteristic intensity or amplitude helps determine the significance of operating condition change events and their impact on the entire system.

[0033] By introducing characteristic intensity or amplitude fields, the intensity or significance of changes in operating conditions can be quantified, enabling precise measurement of the degree of change in operating conditions, improving the descriptive ability of event characteristics, and facilitating subsequent analysis and processing.

[0034] In this context, event direction refers to the trend of data or status changes in certain types of operating condition change events. It describes the nature of the change in operating conditions, such as whether it is an increase or decrease, or the direction of a switch to a specific state. Common event directions include: **Ascending direction:** When a data or status value changes from a lower value to a higher value, the event direction is ascending. For example, if the operating temperature of a device increases from 30°C to 50°C, the temperature event direction is ascending. **Descending direction:** When a data or status value decreases from a higher value to a lower value, the event direction is descending. For example, if the pressure of a device decreases from 100Pa to 50Pa, the pressure event direction is descending. **Stable direction:** When a data or status value remains unchanged for a certain period of time, the event direction can be marked as stable (or without direction).

[0035] In one embodiment, for continuous numerical data, the rate of change is calculated, and the operating condition change event is determined accordingly. Let the adjacent sampling point numbers corresponding to the same data identifier within the same data source be... and The corresponding values ​​are respectively and The corresponding original timestamps are respectively and Time difference is Rate of change Defined as:

[0036] in, and The values ​​are those of adjacent sampling points; and The original timestamps of adjacent sampling points; The time difference between adjacent sampling points; For the rate of change. Optionally, in If the sample point pair is equal to zero or less than the time resolution threshold, it is considered invalid and skipped. When the condition is met... At that time, determine the sampling point There are operating condition change events at the corresponding locations, among which This is the preset rate of change threshold parameter.

[0037] Optionally, the rate of change threshold condition can be used in combination with the magnitude of change threshold condition. (Magnitude of change) Defined as:

[0038] in, This represents the variation amplitude between adjacent sampling points. When the following conditions are met... and At that time, identify the operating condition change event, among which This is a preset amplitude threshold parameter. Optionally, it can be set to... The symbol is used as the event direction field. The The terms "rise" and "fall" are used to distinguish between sudden increases and sudden decreases.

[0039] In one embodiment, for status-based or flag-based data, state value transitions are detected, and operating condition change events are determined accordingly. Let the state values ​​of adjacent sampling points with the same data identifier within the same data source be... and ,in and These are discrete state values. When the following conditions are met... At that time, determine the sampling point A condition change event exists at the corresponding location, and the characteristic type is marked as the corresponding type among start / stop status change, mode switch, or status flag change; the status values ​​before and after the status switch are used as event attribute fields. and .

[0040] The preset threshold can be set using either a statistical distribution based on historical operating data or a range of process parameters. Optionally, the statistical distribution setting includes: calculating from historical data... The distribution and, for example, a certain quantile as ; Calculation of historical data The distribution and, for example, a certain quantile as Optionally, the process range limitation includes: determining it based on the upper limit of the allowable rate of change of the process. Determined based on the upper limit of the allowable jump range of the process. Configure different data identifiers separately. and The threshold is configured as a parameter and is not limited to a specific value.

[0041] Characteristic intensity or amplitude field of operating condition change events Optional , or One or a combination thereof; when a combination is used, Configurable as a weighted sum

[0042] in The weighting parameters are integrated to determine the rate of change magnitude. Weighting parameters are incorporated to account for the magnitude of change. The sequence of operating condition behavior features is sorted within the same source by the original timestamp field of the events, and used to form the sequence input.

[0043] By calculating the rate of change or performing differential analysis on continuous numerical data, and detecting the switching of state-type data, we can effectively identify events that change operating conditions, ensuring that we can capture subtle changes in industrial processes and accurately extract the features of events with significant impact.

[0044] Therefore, by extracting the operating condition behavior features from each data source and generating the corresponding operating condition behavior feature sequence, the state changes of the industrial process can be accurately reflected, ensuring that the core features of the data are effectively extracted, and providing high-quality input for subsequent matching and model building.

[0045] In some embodiments, for step S3, similarity matching is performed on the operating condition behavior feature sequences from different data sources to form a set of matching relationships, and a matching confidence level is determined for each matching relationship. The similarity matching uses dynamic time warping to non-linearly align the operating condition behavior feature sequences from different data sources, calculates the similarity based on the alignment path, and generates a matching relationship when the similarity meets a preset matching threshold.

[0046] Let the operating condition behavior feature sequence of data source A be as follows: The operating condition behavior characteristic sequence of data source B is as follows: .in For sequence The number of events, For sequence The number of events; For the first in data source A Characteristics corresponding to each operating condition change event For the first in data source B Features corresponding to each operating condition change event. Each feature must contain at least a feature type. In some embodiments, it also includes characteristic intensity or amplitude. With direction .

[0047] Single-point distance function It can be defined as:

[0048] in for Event index in for Event index in; For feature type; Characteristic intensity or amplitude; For event direction field; This is an indicator function; it takes the value 1 if the condition is true, and 0 otherwise. Weighting coefficients for inconsistent feature types; The weighting coefficient for the amplitude difference term; This represents the weighting coefficient for the direction inconsistency term. For terms not included... An example of the field will be Configured to 0; for those not included An example of the field will be Configured to 0.

[0049] Dynamic Time Warping (DTW) Cumulative Distance Defined as:

[0050] in The alignment path consists of several alignment index pairs. The structure consists of elements that satisfy both the index monotonicity constraint and the step size constraint. The index monotonicity constraint is: if... and If they are two adjacent pairs on the path, then and The step size constraint is: the increments of two adjacent pairs belong to... , , At least one of the following is used to ensure path continuity. To align the summation operator on the path; This indicates selecting the alignment path with the smallest cumulative distance from the set of candidate alignment paths; Let be the dynamic time-normalized distance between the two sequences.

[0051] Similarity Depend on The mapping yields:

[0052] in It is an exponential function; For similarity scale parameters; This represents the similarity score. Let the preset matching threshold be... , This is the similarity threshold parameter used to determine matching relationships. When the following conditions are met... At that time, the corresponding matching relationship is generated and included in the matching relationship set. Optionally, for events in the same data source A... With multiple candidate events in data source B When all conditions are met, the similarity will be... The largest candidate is used as the primary match, and the remaining candidates are used as alternative matches with their confidence levels preserved.

[0053] By using the Dynamic Time Warping (DTW) method to nonlinearly align the operating condition feature sequences from different data sources, this step can effectively eliminate alignment errors caused by different time scales, accurately match operating condition change events across data sources, and improve the accuracy and robustness of event matching.

[0054] Matching confidence Determined by similarity. Let candidate matching pairs be... The similarity is The normalized upper bound is ,in It can take the maximum similarity value within the current matching window or a preset upper bound, and Matching confidence Defined as:

[0055] in The similarity of candidate matching pairs; This represents the matching confidence score for candidate matching pairs. Optionally, a data source reliability weight can be introduced. and And form:

[0056] in For the reliability weight parameter of data source A, For data source B, use reliability weight parameters. The mapping function from similarity to confidence is... It can be configured as a linear mapping or a piecewise linear mapping; For the reliability combination function, the Configurable to or . and It can be configured to use weights based on data source missing rate, duplication rate, delay anomaly indicators, or historical matching confidence statistics.

[0057] Therefore, by matching features across data sources, event alignment between different data sources can be achieved, eliminating the impact of inconsistent data timestamps, ensuring the effective fusion of multi-source data, and improving data consistency and accuracy.

[0058] In some embodiments, for step S4, a working condition state transition model is constructed based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes; in some embodiments, the state transition frequency is statistically analyzed and the state transition probability is determined based on the working condition state sequence in historical data, and these probabilities are used as parameters of the working condition state transition model.

[0059] The determination of operating condition status nodes may optionally include: dividing the data segment based on the process operation mode field, start / stop field, key valve position range, or energy consumption range, and taking each division result as an operating condition status node; or clustering based on the characteristic distribution of operating condition change events within the window, and taking the cluster category as the operating condition status node. State transition relationships are formed by state changes in adjacent time steps, and these adjacent time steps are determined by the event cluster sorting result or the adjacent symbols of the observation sequence.

[0060] In some embodiments, the operating condition state transition model is configured as a Hidden Markov Model. The set of hidden states is... ,in For the first Each working condition status node The number of state nodes. The initial state distribution is as follows: ,in The state transition probability matrix is: ,in The observation probability distribution is as follows: Used to describe the hidden state The observed symbol was observed below. The probability of . The observation sequence is ,in The number of observation time steps. For the first Each observation symbol is determined by the characteristic type of the operating condition change event. With amplitude discretization results The combination of is used as the encoding, the To make the amplitude field The intervals are numbered according to the discrete intervals.

[0061] The state transition probability is obtained based on historical data statistics. Let's assume the state transition probability starts from... to state The number of adjacent transitions is The smoothing parameter is ,in If the parameter is non-negative, then:

[0062] in Index for target state; To sum over all target states; This represents the normalized transition probability. Optionally, the initial state distribution... This is obtained from the frequency statistics of the starting states in the historical sequence. For examples lacking annotations, , and The estimation is obtained through expectation-maximization iterative estimation, which includes: calculating and updating the posterior probability of the hidden state at each time step given parameters. and The process continues until the log-likelihood increment is less than the convergence threshold.

[0063] Therefore, constructing a state transition model based on the matching relationship of operating condition behavior characteristics helps to analyze the change patterns between different operating conditions, provides accurate state transition prediction capabilities, and provides constraints and support for the subsequent generation of a unified time benchmark and alignment.

[0064] In some embodiments, for step S5, a unified time reference is generated based on the matching relationship set and the operating condition transition model. The unified time reference is determined by the sorting result of operating condition change events or event clusters, the sorting satisfying the constraints of the operating condition transition model, and not using the original timestamp of any single data source as the cross-data source alignment reference.

[0065] Please see Figure 3 , Figure 3 This is a schematic diagram of the time base generation process provided in this embodiment of the disclosure. Specifically, in S301, event cluster aggregation occurs. This includes: constructing an event graph from the set of matching relationships, where the nodes of the event graph are operating condition change events and the edges of the event graph are matching relationships; and aggregating events that satisfy the matching relationships into the same event cluster. The event cluster is denoted as... , It contains a set of operational condition change events from different data sources and associated through matching relationships. Optionally, if alternative matching relationships exist in the event graph, the matching confidence level is used as the edge weight, and edges with weights lower than the aggregation threshold are not included in the aggregation. The aggregation threshold is... ,in This is the confidence gating parameter used for event cluster aggregation.

[0066] In S302, event cluster sorting is performed. Specifically, this includes constructing a candidate sorting set and selecting the sorting result based on constraints. Evaluation terms for candidate sorting include consistency of origin order, consistency of matching confidence, and consistency of state transition. The consistency of origin order includes: for any two events within a data source... and If its original timestamp satisfies Then it is required that the event cluster to which it belongs satisfies the following in the sorting: exist Previously, when this constraint conflicted with a cross-source high-confidence matching relationship, the conflict was preserved and decided by the state transition consistency term. The matching confidence consistency term includes: for cross-source matching relationships... ,like A higher probability requires that the two corresponding events belong to the same event cluster or adjacent event clusters in the sorting. The state transition consistency term includes: calculating the probability of the state sequence inferred from the candidate event cluster sequence and selecting the sequence with the higher probability. For the Hidden Markov Model embodiment, the most probable state sequence corresponding to the candidate event cluster sequence is obtained by the Viterbi algorithm, and the Viterbi recursion uses... and Calculate the path probability and output the most likely path.

[0067] By aggregating operating condition change events across data sources into event clusters and sorting these event clusters according to the operating condition state transition model, this step ensures that the sequence of events between the event clusters has a reasonable logical order, providing the necessary structured support for the generation of a unified time base and enhancing the accuracy and reliability of time alignment.

[0068] In S303, a unified time base is established using a logical time indexing system. A logical time index is assigned to each event cluster. The The sequence number of the event cluster in the sorted sequence; indexed by logical time. The corresponding interval is used as the logical time interval. The A unified time base is established. The boundaries of logical time intervals are determined by the sorting positions of adjacent event clusters. Optionally, the length of a logical time interval is measured by the number of data points within the interval or the span of sampling sequence numbers within the interval, rather than using the value of the original timestamp from a specific data source as the interval length.

[0069] Relative position parameters are used within the logical time interval. Indicates the order within the interval. The range of values ​​is Suppose there is a data source. In event clusters There are a total of within the corresponding interval The data points, It is a positive integer; the first integer in this data source. The index of each data point within the interval is and Then the data point It can be configured as follows:

[0070] in This refers to the relative position parameter within the interval; The index is the serial number within the interval; This represents the number of data points from this data source within this logical time interval. When it equals 1, The configuration is set to 0. The aforementioned... Used for sorting within a range, not as a benchmark for cross-data source alignment.

[0071] Therefore, by sorting the events of changing operating conditions to generate a unified time base, without relying on the original timestamps of a single data source, the problem of time inconsistency between data sources is eliminated, and a unified time frame for multi-source data is achieved, which is helpful for data integration and analysis from different data sources.

[0072] In some embodiments, for step S6, the data sequences of each data source are mapped to a logical time interval on a unified time base to obtain multi-source aligned data under a unified time base; when there is a mapping conflict, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

[0073] In one embodiment, the rule for assigning data points to logical time intervals is as follows: if a data point corresponds to the detection trigger point of a working condition change event, then the data point is assigned to the event cluster to which the corresponding working condition change event belongs. And mapped to logical time intervals If a data point does not correspond to a detection trigger point, it is assigned to a stable segment between two adjacent operating condition change events within the same source, and mapped to a logical time interval between adjacent event clusters. If missing events cause the boundary of the stable segment to be unclear, the most recent valid operating condition change event is used as the boundary, and the stable segment is mapped to the logical time interval corresponding to that event cluster. When using a sliding time window to extract operating condition change events, data points within the window are assigned to the corresponding event cluster of the window and mapped to the corresponding logical time interval.

[0074] Mapping is represented as a function For any industrial data record , or ,in For logical time index, This refers to the relative position parameter within the interval. Multi-source aligned data under a unified time base must contain at least a logical time index. Relative position parameters within the interval The data source identifier, data identifier and numerical or status value; optionally, it also includes a matching confidence field or mapping weight field associated with the data point.

[0075] In one embodiment, mapping conflicts include at least candidate mappings for multiple logical time intervals corresponding to the same industrial data record. Let the industrial data record... The candidate logical time index set is ,in The number of candidates; the matching confidence level corresponding to the candidate mapping is The To be Belong to The confidence level.

[0076] Maximum confidence attribution includes, selection satisfy:

[0077] Data points Belong to This refers to the corresponding logical time interval. Here... Indicates data points Assigned to candidate mapping confidence level It is a set of candidate logical time intervals. It is the candidate interval with the highest confidence. This is the final logical time interval selected.

[0078] Confidence-weighted attribution includes calculating normalized weights. And allocated according to weights, the normalized weights are defined as follows:

[0079] in For normalized weights; To sum the candidate confidence scores; when the denominator is 0, all candidate weights are set to be equal. Confidence-gated elimination includes: when... Candidates are removed at the same time ,in This is the confidence level gating threshold parameter. State transition consistency backoff includes: when the probability of the state sequence corresponding to the conflict-adjusted event cluster sequence under the operating condition state transition model falls below the threshold. At that time, the event cluster sorting and data mapping for the conflict-occurring interval are re-executed, where This is the probability threshold parameter for the state sequence.

[0080] This step ensures that multi-source data can be aligned according to a unified time frame, resolves timestamp conflicts across data sources, and further improves the accuracy and reliability of data fusion by adjusting mapping conflicts.

[0081] In some embodiments, a dynamic update step is also included. The triggering conditions for dynamic update include the switching frequency of operating condition change events meeting a threshold condition, the matching confidence meeting a threshold condition, or the latency anomaly of at least one data source meeting a threshold condition; when dynamic update is triggered, at least one of the matching relationship set or the operating condition state transition model is updated, and the unified time base and mapping relationship are updated accordingly.

[0082] The switching frequency of operating condition change events is expressed as the ratio of the number of operating condition change events within the monitoring window to the span of the monitoring window. Let the number of operating condition change events within the monitoring window be... The monitoring window time span is ,in Positive number; switching frequency Defined as:

[0083] in To monitor the number of events involving changes in operating conditions within the monitoring window; To monitor the time span of the monitoring window; This is the operating condition switching frequency. When the following conditions are met... Dynamic updates are triggered at specific times, where To switch frequency threshold parameters. Monitoring window span. Optionally, the equivalent time length can be calculated based on the time length or the number of logical time intervals covered.

[0084] The matching confidence threshold is determined by the overall confidence level of the matching relationship set within the monitoring window. Let the matching relationship set within the monitoring window be... The size of the set is ;No. The matching confidence of each matching relation is Average confidence level Defined as:

[0085] in This represents the average match confidence level. To monitor the number of matching relationships within the window; This is a built-in confidence summation operator for a set. When the following conditions are met... Dynamic updates are triggered at specific times, where This is the average confidence threshold parameter. Optionally, it represents the proportion of low confidence levels. Defined as:

[0086] in The proportion of low confidence levels; Below the threshold The number of matching relationships. When satisfied Dynamic updates are triggered at specific times, where This is the low confidence level threshold parameter.

[0087] The data source latency anomaly threshold condition is determined using cross-source matching residual discrimination. Let the data source be... The residual index is Residual index Optionally, one or a combination of the following can be selected: the average index offset of the dynamic time-warped alignment path, the average absolute value of the event cluster index difference in the matching relationship, and the magnitude of the decrease in the mean similarity within the monitoring window. Let the residual threshold be... When satisfied Dynamic updates are triggered at specific times, where This is the residual threshold parameter.

[0088] The dynamically updated objects include at least one of the matching relationship set or the working condition state transition model. Updating the matching relationship set includes: re-performing dynamic time warping similarity matching and updating the similarity, matching relationships, and matching confidence; or recalculating based on the existing matching relationship set according to the events within the latest monitoring window. And update The update of the operating condition state transition model includes: statistically analyzing the incremental frequency of state transitions for newly added historical data and updating the state transition probabilities; for the Hidden Markov Model implementation, the update also includes updating the initial state distribution. State transition matrix or observed probability distribution At least one of them should be re-estimated.

[0089] In dynamic updates, weight adjustments may be performed optionally. Weight adjustments include: adjusting the weights in the distance function. , and Adjustments were made to the reliability weights. and Adjustments were made to the confidence level gating threshold. With aggregated gate threshold Adjustments were made. The reliability weight of data sources that trigger latency anomalies was reduced, the corresponding weight coefficient of noise-sensitive feature types was reduced, and the gating threshold of monitoring windows with high switching frequency was increased to eliminate low-confidence candidate mappings.

[0090] In some embodiments, step S2 extracts operating condition change events based on a sliding time window. Let the window length be... ,in The sliding time window length parameter. Calculate the local rate of change statistic within the window, either by time length or number of sampling points. Trend slope or variance And when a threshold condition is met, a window-level operating condition change event is generated. Local rate of change statistics. Configured as the average rate of change within the window; trend slope The slope is obtained by linear fitting of the sampling points within the window; variance This represents the variance of values ​​within the window. The original timestamp field for window-level condition change events is taken from the original timestamp corresponding to the center of the window or the original timestamp corresponding to the window boundary.

[0091] In some embodiments, step S3 achieves similarity matching based on feature distribution similarity. This involves analyzing the feature type distribution vector of events within the monitoring window. Amplitude distribution vector or event interval distribution vector Construction distribution distance And mapped to similarity Distribution distance Configure it as L1 distance or KL divergence. Taking L1 distance as an example, let the distribution vectors corresponding to the two data sources be respectively... and The and If the probability vectors are of the same dimension and all their components are non-negative and sum to 1, then:

[0092] in Indexed by vector dimension. and For the first Dimensional components. Distribution similarity configuration is as follows:

[0093] in This is the distribution similarity scale parameter. Matching threshold The comparisons are used to generate a set of matching relationships, and the matching confidence is determined accordingly.

[0094] In some embodiments, step S5 generates a unified time base for each production batch. Industrial data is segmented based on batch number, work order number, or production segment. For each segment, steps S2 to S6 are executed to generate a logical time index system, and the logical time indexes are concatenated in segment order. Optionally, inter-segment connection intervals are set between segments to carry inter-segment transition data.

[0095] In some embodiments, step S5 generates a locally unified time base according to a specific analysis task. Let the target analysis interval fall within the original timestamp range of a certain data source. ,in and This represents the original timestamp boundary; in one example, the buffer length is configured as follows: And expand the scope of analysis to ,in The buffer time length parameter is used; steps S2 to S6 are performed only on data within the extended range to generate a local logical time index and output local multi-source aligned data.

[0096] Multi-source aligned data is organized by logical time index and contains at least a logical time index. Relative position parameters within the interval Data source identifier, data identifier, numerical or status value; optionally includes a matching confidence field. Or mapping weight field Based on the multi-source aligned data, a consistent logical alignment result across data sources can be generated under conditions of frequent operating condition switching or inconsistent data source timestamps, and can be used for joint analysis and playback processing after alignment.

[0097] This invention achieves precise alignment of multi-source data through a unified time benchmark, effectively solving the time alignment problem between different data sources. Specifically, by using cross-data source feature similarity matching and dynamic time warping, it overcomes the shortcomings of traditional methods in handling data source heterogeneity and timestamp discrepancies, thereby improving the accuracy and reliability of data fusion. Furthermore, by extracting operational behavior features and constructing operational state transition models, this invention enables multi-source data analysis to go beyond simple data merging, uncovering key features in the production process from the perspective of operational changes, thus optimizing the quality of data analysis. A dynamic update mechanism ensures that the method can adaptively handle new changes occurring in the production process, further enhancing the flexibility and real-time performance of the data analysis system. Overall, the method of this invention has significant advantages in accuracy, efficiency, and adaptability in processing multi-source heterogeneous data, and can be widely applied in fields such as intelligent manufacturing and industrial big data analysis, promoting the improvement of data intelligence and automation levels in industrial production.

[0098] Figure 4 An industrial multi-source heterogeneous data acquisition and analysis system 400 is shown. The system embodiment is similar to... Figure 1 Corresponding to the illustrated method embodiments, the specific methods include: The industrial data acquisition module 401 is used to acquire industrial data from multiple data sources. The industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and forms a data sequence corresponding to each data source based on the data source identifier. The working condition behavior feature extraction module 402 is used to extract working condition behavior features that characterize working condition changes for each of the data sequences and generate working condition behavior feature sequences corresponding to each data source. The working condition behavior features include working condition change events, and the working condition change events include at least feature types and corresponding original timestamps. Matching module 403 is used to perform feature similarity matching across data sources based on the characteristic sequences of each working condition, to obtain a set of matching relationships for working condition change events across data sources, and to determine the matching confidence for the matching relationships in the set of matching relationships; The model building module 404 is used to build a working condition state transition model based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. The time reference generation module 405 is used to generate a unified time reference based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events and does not use the original timestamp of any single data source as the reference for cross-data source alignment. The mapping alignment module 406 is used to map the data sequences of each data source to the logical time interval on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when there is a mapping conflict, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

[0099] Based on the same inventive concept, this application also provides a computer device, the method corresponding to which can be the method in the foregoing embodiments, and the principle of solving the problem is similar to that method. The computer device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the foregoing embodiments of this application.

[0100] The computer device can be a user device, or a device formed by integrating user devices and network devices through a network, or it can be an application running on the aforementioned devices. The user device includes, but is not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and smart bands. The network device includes, but is not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, and can be used to implement some processing functions when setting an alarm clock. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.

[0101] Figure 5The diagram illustrates the structure of an apparatus suitable for implementing the methods and / or technical solutions in the embodiments of this application. The apparatus 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0102] The following components are connected to I / O interface 505: input section 506 including keyboard, mouse, touch screen, microphone, infrared sensor, etc.; output section 507 including cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, etc., and speakers, etc.; storage section 508 including one or more computer-readable media such as hard disk, optical disk, magnetic disk, semiconductor memory, etc.; and communication section 509 including network interface card such as LAN (local area network) card, modem, etc. Communication section 509 performs communication processing via a network such as the Internet.

[0103] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, it performs the functions defined in the methods of this application.

[0104] Another embodiment of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.

[0105] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0106] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0107] Furthermore, the inclusion of a single word does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any particular order.

[0108] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0109] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for acquiring and analyzing multi-source heterogeneous industrial data, characterized in that, Includes the following steps: Acquire industrial data from multiple data sources, wherein the industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and form a data sequence corresponding to each data source based on the data source identifier; For each of the data sequences, extract the operating condition behavior features that characterize the changes in operating conditions, and generate the operating condition behavior feature sequence corresponding to each data source. The operating condition behavior features include operating condition change events, and the operating condition change events include at least feature types and corresponding original timestamps. Based on the characteristic sequences of each working condition, feature similarity matching across data sources is performed to obtain a set of matching relationships for working condition change events across data sources, and a matching confidence level is determined for the matching relationships in the set of matching relationships. A working condition state transition model is constructed based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. A unified time reference is generated based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events, and does not use the original timestamp of any single data source as the reference for cross-data source alignment. Data sequences from various data sources are mapped to logical time intervals on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when mapping conflicts exist, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

2. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, The extraction of operating condition behavior features representing changes in operating conditions for each of the data sequences includes, Calculate the rate of change or difference between adjacent sampling points for continuous numerical data, and determine the corresponding working condition change event when the rate of change or difference meets a preset threshold condition. The system detects state value transitions for status-type or flag-type data and determines the corresponding operating condition change event when the state value transition is detected.

3. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 2, characterized in that, The operating condition change event also includes characteristic intensity or amplitude, wherein the characteristic intensity or amplitude is determined by the change amplitude, rate of change or step amplitude of the corresponding continuous numerical data.

4. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, The step of performing cross-data source feature similarity matching based on the various working condition behavior feature sequences includes: using dynamic time warping to non-linearly align the working condition behavior feature sequences from different data sources, calculating the similarity based on the alignment path, and generating the matching relationship in the matching relationship set when the similarity meets the preset matching threshold condition.

5. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, The step of constructing the working condition state transition model based on the matching relationship set includes determining the state transition probability between the working condition state nodes by statistically analyzing the state transition frequency based on the working condition state sequence in historical data. The state transition probability is then used as a parameter of the operating condition state transition model.

6. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, The step of generating a unified time reference based on the matching relationship set and the working condition state transition model includes: aggregating working condition change events across data sources into event clusters based on the matching relationship set; sorting the event clusters according to the constraints of the working condition state transition model; determining the order of the logical time intervals based on the sorting results; and thereby generating the unified time reference.

7. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, The mapping conflict includes candidate mappings for the same data record corresponding to multiple logical time intervals. Adjusting the mapping conflict includes: assigning the data record to the logical time interval corresponding to the candidate mapping with the highest matching confidence, or assigning candidate mappings with weighted weights based on matching confidence.

8. The method for acquiring and analyzing multi-source heterogeneous industrial data according to claim 1, characterized in that, It also includes dynamic updates, specifically, when a preset trigger condition is met, updating at least one of the matching relationship set or the working condition state transition model, and updating the unified time base and the mapping relationship accordingly; wherein, the preset trigger condition includes the switching frequency of the working condition change event meeting a threshold condition, the matching confidence meeting a threshold condition, or the latency anomaly of at least one data source meeting a threshold condition.

9. A system for acquiring and analyzing multi-source heterogeneous industrial data, characterized in that, An industrial data acquisition module is used to acquire industrial data from multiple data sources. The industrial data includes at least a data identifier, a numerical or status value, an original timestamp, and a data source identifier, and forms a data sequence corresponding to each data source based on the data source identifier. The working condition behavior feature extraction module is used to extract working condition behavior features that characterize working condition changes for each of the data sequences and generate working condition behavior feature sequences corresponding to each data source. The working condition behavior features include working condition change events, and the working condition change events include at least feature types and corresponding original timestamps. The matching module is used to perform feature similarity matching across data sources based on the sequence of behavioral characteristics of each working condition, to obtain a set of matching relationships for working condition change events across data sources, and to determine the matching confidence for the matching relationships in the set of matching relationships; The model building module is used to build a working condition state transition model based on the matching relationship set. The working condition state transition model includes working condition state nodes and the state transition relationships between working condition state nodes. The time reference generation module is used to generate a unified time reference based on the matching relationship set and the working condition state transition model. The unified time reference is determined by the sorting result of the working condition change events and does not use the original timestamp of any single data source as the reference for cross-data source alignment. The mapping and alignment module is used to map the data sequences of each data source to the logical time interval on the unified time base to obtain multi-source aligned data under the unified time base; wherein, when there is a mapping conflict, the mapping relationship of the conflicting data is adjusted according to the matching confidence.

10. A computer device, wherein the computer device is characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.