Big data analysis management method and system for mechanical manufacturing production data
By using an industrial temporal causal entropy algorithm and a self-learning and adaptable process temporal semantic library, a four-dimensional dynamic indexing system is constructed. This solves the problem of existing systems in quickly locating causal data and adapting to process changes, and enables efficient management and anomaly response of mechanical manufacturing production data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGKE LIXIANG ELECTRIC (SHANDONG) CO LTD
- Filing Date
- 2025-11-04
- Publication Date
- 2026-05-08
AI Technical Summary
Existing mechanical manufacturing production data management systems have limitations in data organization and parameter configuration, making it impossible to quickly locate causal data. This results in the need to sift through massive amounts of data to find related information when anomalies occur, which is insufficient to meet the need for rapid response to production anomalies.
The industrial time-series causal entropy algorithm is used for quantitative analysis, and a four-dimensional dynamic index system is constructed. Combined with a self-learning and adaptable process time-series semantic library, the system can achieve full-process data control and rapid response management of abnormal data.
Through a multi-dimensional structured indexing system and self-learning adaptation, it can quickly locate causal data, improve retrieval efficiency, dynamically calibrate core parameters, adapt to process changes, and ensure analysis accuracy and anomaly response speed.
Smart Images

Figure CN121434274B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mechanical manufacturing production data management technology, specifically a big data analysis and management method and system for mechanical manufacturing production data. Background Technology
[0002] In the mechanical manufacturing process, with the deep integration of IoT technology and intelligent manufacturing, a continuous stream of heterogeneous data is generated, including equipment operation data (such as speed and temperature), process parameter data (such as cutting depth and welding current), and quality inspection data (such as dimensional accuracy and pass rate). This data is not only massive in volume and frequently sampled, but also has a close temporal causal relationship with production quality control and equipment fault diagnosis. Enterprises need to accurately identify the causal relationships between data and quickly trace the root cause when abnormal data occurs to ensure stable production processes and reduce the generation of defective products. This has become one of the core requirements of industrial data management.
[0003] Currently, mainstream production data management systems have significant limitations in data organization and parameter configuration. On the one hand, data indexes often rely on a single dimension (such as timestamps or equipment IDs), only enabling data sorting by time or categorization by equipment. They cannot link data to underlying causal relationships within the process, requiring the sifting through massive amounts of data to identify key influencing factors when anomalies occur. On the other hand, core parameters supporting data processing and causal analysis (such as time-series analysis windows and data fluctuation thresholds) are mostly statically set, requiring manual reconfiguration based on process adjustments (such as changing processing materials or optimizing process parameters). This not only results in adjustment delays but also makes parameter settings prone to deviations due to differences in human experience. This model directly leads to existing systems being unable to efficiently adapt to process changes and failing to meet the actual needs of rapid response to production anomalies, becoming a key bottleneck restricting the value mining of industrial data. Summary of the Invention
[0004] The purpose of this invention is to provide a big data analysis and management method and system for mechanical manufacturing production data, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a big data analysis and management method for mechanical manufacturing production data, comprising the following steps:
[0006] Step 100: First, collect multi-source production data, then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby achieving full-process data control;
[0007] Step 200: Based on the industrial time-series causal entropy algorithm, perform quantitative analysis on the process time-series causal intensity of the dependent variable data and the effect variable data in the standardized data sequence, and output the causal weight as the analysis result to provide a quantitative basis for subsequent management;
[0008] Step 300: Based on the causal weights, construct a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range, and causal weights, and construct a time-series causal tree index hierarchically according to the strength of causal association to realize the systematic management of the index structure.
[0009] Step 400: Based on the historical data, analysis results and real-time feedback data of the standardized data sequence, construct and self-learn to update the process timing semantic library, synchronously calibrate the core parameters of the algorithm, and ensure dynamic adaptation of analysis and management;
[0010] Step 500: Monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it, so as to realize rapid response management of abnormal data.
[0011] Preferably, step 100 further includes:
[0012] Step 110: Collect equipment operation data, process parameter data, and quality inspection data to form a raw data set;
[0013] Step 120: Perform noise filtering and missing value completion preprocessing operations sequentially on the original dataset to remove invalid data;
[0014] Step 130: Perform standardization on the preprocessed raw data, classify it according to data type and map it to a preset numerical range, so that the units of data of the same type are unified, and obtain the dependent variable data sequence. With at least one result variable data sequence ,in Represents the dependent variable data. Represents result variable data, The timestamps are consistent between the dependent variable data sequence and the result variable data sequence, and their time span corresponds to the process time window. Matching The effective time range for characterizing the causal relationship of processes.
[0015] Preferably, in step 130, the standardization process specifically involves: classifying the preprocessed raw data according to preset data type classification rules, with each data type corresponding to a unique preset numerical range; and mapping each type of data to its corresponding range through linear transformation to ensure the stability of the dependent variable data sequence. With result variable data sequence The timing alignment meets the input requirements for subsequent process timing causal analysis.
[0016] Preferably, step 200 further includes:
[0017] Step 210: Time-series correlation quantification, calculating the dependent variable data sequence With result variable data sequence Fluctuation synergy correlation ,in For the degree of correlation of fluctuation synergy, It is a time lag and satisfies To capture the correlation between abnormal fluctuations;
[0018] Step 220: Process constraint correction, introducing process causal constraint factors. Regarding the aforementioned fluctuation synergy correlation degree Make corrections to obtain the process compatibility correlation. Distinguish between statistical correlation and process causality;
[0019] Step 230: Calculate the industrial time-series causal entropy based on the process adaptation correlation. Calculate causal entropy Quantify the certainty of causal relationships;
[0020] Step 240: Causal weight mapping, which maps the causal entropy. Transformed into the causal weights The sum of the causal weights is 1, which provides a priority basis for index hierarchical management.
[0021] Preferably, in step 210, the fluctuation synergy correlation degree The calculation formula is:
[0022]
[0023] In the formula, The start time of the timing window. For the dependent variable data sequence The mean within the time window For result variable data sequence The mean within the time window For the threshold of fluctuation in dependent variable data, For the threshold of fluctuation in the result variable data; when or At that time, in the numerator of the formula for calculating the degree of correlation of fluctuation synergy, the corresponding time point is... product term Set to 0;
[0024] In step 220, the process adaptation correlation The calculation formula is:
[0025]
[0026] In the formula, For process adaptability coefficient, It is a symbolic function; only when Then proceed to step 230;
[0027] In step 230, causal entropy The calculation formula is:
[0028]
[0029] In the formula, It is a volatility smoothing factor and , It is a volatility amplification factor and , It is an exponential function. , To strengthen causality, For weak causality, It is not causal;
[0030] In step 240, causal weight The calculation formula is:
[0031]
[0032] In the formula, To the dependent variable data sequence There are time-series related data sequences of all result variables. The number of data sequences for the result variable.
[0033] Preferably, step 300 further includes:
[0034] Step 310: Based on process attributes, equipment operating status, data time series range, and causal weights The four core information categories are used to construct the four-dimensional dynamic index key, clarifying the core identifier dimension of the index, and the corresponding data time series range. ;
[0035] Step 320: Using the dependent variable data with strong causality as the root node, the effect variable data with strong causality as the first-level child nodes, and the effect variable data with weak causality as the second-level child nodes, excluding non-causal data, construct the hierarchical time-series causal tree index to achieve structured management of related data, where strong causality corresponds to... Weak causal correspondence Non-causal correspondence ;
[0036] Step 400 further includes:
[0037] Step 410: Storage of the process timing semantic library Process causal constraint factors Process adaptability coefficient Fluctuation smoothing factor Fluctuation amplification factor Threshold for fluctuation of dependent variable data Threshold for fluctuation of result variable data The initial values and update records are used to automatically calibrate the above parameters when process parameters are adjusted or new processes are added.
[0038] Step 420: Based on the historical accumulated data of the standardized data sequence and the real-time feedback retrieval accuracy data, update the process time sequence semantic library. The self-learning adaptation time is ≤24 hours to ensure the adaptability of index construction and predictive retrieval.
[0039] Preferably, step 500 further includes:
[0040] Step 510: Monitor the standardized data sequence in real time, wherein the fluctuation threshold of the dependent variable data is... The fluctuation threshold of the result variable data is When any data in the standardized data sequence exceeds its corresponding fluctuation threshold, an abnormal signal is triggered, and the abnormal signal is associated with the corresponding data that exceeds the threshold.
[0041] Step 520: Based on the dependent variable data corresponding to the abnormal signal, call the updated process timing semantic library to predict the associated first-level child node data and second-level child node data from the timing causal tree index;
[0042] Step 530: Preload the index corresponding to the predicted associated data in advance. After receiving the retrieval request, synchronously output the root node data and all associated child node data corresponding to the abnormal signal. The retrieval response time is ≤1 second, realizing efficient management of abnormal data.
[0043] This invention also provides a big data analysis and management system for mechanical manufacturing production data, comprising:
[0044] The data acquisition and management module is used to collect multi-source production data, and then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby realizing full-process data control.
[0045] The causal entropy analysis module is used to perform quantitative analysis of the process time-series causal intensity of dependent variable data and effect variable data in the standardized data sequence based on the industrial time-series causal entropy algorithm, and outputs the causal weight as the analysis result to provide a quantitative basis for subsequent management.
[0046] The index building and management module is used to build a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range and causal weight based on the causal weight, and to build a time-series causal tree index in layers according to the strength of causal association, so as to realize the systematic management of the index structure.
[0047] The self-learning adaptation module is used to build and self-learn to update the process timing semantic library based on the historical data, analysis results and real-time feedback data of the standardized data sequence, and synchronously calibrate the core parameters of the algorithm to ensure dynamic adaptation of analysis and management.
[0048] The predictive retrieval management module is used to monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it, so as to realize rapid response management of abnormal data.
[0049] The monitoring module is used to monitor the index building status, algorithm parameter accuracy, prediction hit rate, and data flow integrity to ensure stable system operation.
[0050] The present invention also provides an electronic device, which is a physical device, comprising:
[0051] The processor and the memory are communicatively connected.
[0052] The memory is used to store at least one executable instruction executed by the processor, which executes the executable instruction to implement the big data analysis and management method for mechanical manufacturing production data as described above.
[0053] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the big data analysis and management method for mechanical manufacturing production data as described above.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] This invention addresses the problem of single-dimensional indexing in traditional production data by constructing a multi-dimensional structured index system. It can quickly locate causal data and improve retrieval efficiency. Relying on a self-learning and adaptable process time-series semantic library, it can dynamically calibrate core parameters to adapt to process changes, avoiding the lag and errors of manual adjustments and ensuring the accuracy of analysis. Combined with the industrial time-series causal entropy algorithm, it quantifies causal strength and excludes non-causal data to reduce redundancy. With real-time monitoring and predictive retrieval, it can quickly respond to abnormal data, providing strong support for efficient management and anomaly handling of production data. Attached Figure Description
[0056] Figure 1 A main flowchart of a big data analysis and management method for mechanical manufacturing production data provided in an embodiment of the present invention;
[0057] Figure 2 A schematic diagram of the structure of a big data analysis and management system for mechanical manufacturing production data provided in an embodiment of the present invention;
[0058] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The method in this embodiment is executed by a terminal, which can be a mobile phone, tablet computer, PDA, laptop or desktop computer, etc. Of course, it can also be other devices with similar functions, and this embodiment does not limit them.
[0061] Please see Figure 1 This invention provides a big data analysis and management method for mechanical manufacturing production data, the method being applied to, including:
[0062] Step 100: First, collect multi-source production data, then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby achieving full-process data control.
[0063] Specifically, step 100 further includes:
[0064] Step 110: Collect equipment operation data, process parameter data, and quality inspection data to form a raw data set;
[0065] Step 120: Perform noise filtering and missing value completion preprocessing operations sequentially on the original dataset to remove invalid data;
[0066] Step 130: Perform standardization on the preprocessed raw data, classify it according to data type and map it to a preset numerical range, so that the units of data of the same type are unified, and obtain the dependent variable data sequence. With at least one result variable data sequence ,in Represents the dependent variable data. Represents result variable data, The timestamps are consistent between the dependent variable data sequence and the result variable data sequence, and their time span corresponds to the process time window. Matching The effective time range for characterizing the causal relationship of processes.
[0067] Multi-source production data specifically includes equipment operation data (such as real-time sensor data like machine tool speed, temperature, and vibration), process parameter data (such as process settings like cutting depth, welding current, and pressure), and quality inspection data (such as quality inspection results like dimensional accuracy, surface roughness, and pass rate). Collecting this data allows for coverage of the entire production process, providing multi-dimensional evidence for causal correlation analysis. Noise filtering refers to removing random interference signals (such as jump values caused by instantaneous sensor errors) from the data using methods like moving averages and wavelet transforms. Missing value completion uses linear interpolation, KNN algorithms, or LSTM prediction models based on time-series correlation to fill in missing data. The combined effect of these two methods can eliminate invalid data, ensuring the integrity and accuracy of the original dataset. Standardization processes map data with different dimensions (such as temperature in °C and vibration in mm / s) to a unified numerical range (such as [0,1] or [-1,1]) to avoid the impact of dimensional differences on the weighting of causal analysis; while the dependent variable data sequence... With result variable data sequence The division is based on the preset process logic (e.g., "cutting speed" as the dependent variable and "machining accuracy" as the result variable), with consistent timestamps and time series spans. Matching is to ensure that the two are aligned in time within the effective time range of the process, laying the foundation for subsequent analysis of "how the fluctuation of the dependent variable affects the result variable".
[0068] In another possible implementation, the multi-source data acquisition in step 110 can be achieved through an Industrial Internet of Things (IIoT) platform, using edge computing nodes for real-time data aggregation to avoid cloud transmission delays; the noise filtering in step 120 can select an appropriate algorithm for different data types (e.g., wavelet thresholding for high-frequency vibration data and moving average for low-frequency temperature data); the normalization mapping in step 130 can use Min-Max normalization (suitable for relatively uniformly distributed data) or Z-Score normalization (suitable for approximately normally distributed data), and the specific method can be dynamically selected according to the data distribution characteristics.
[0069] For example, in a feasible implementation scenario of automotive parts processing, step 110 collects the spindle speed (equipment operation data), feed rate (process parameter data), and part diameter error (quality inspection data) of the CNC machine tool; step 120 identifies and filters abnormal jump values in the speed data using the 3σ criterion, and uses linear interpolation to complete the missing 5-second feed rate values caused by sensor offline; step 130 maps the speed (range 500-2000 rpm), feed rate (range 0.1-0.5 mm / r), and diameter error (range -0.02-0.02 mm) to the [0,1] interval to obtain the dependent variable sequence. (Rotation speed + Feed rate) and result variable sequence (Diameter error), and both timestamps are 10ms intervals, the time span is consistent with the processing technology. =30s matching ensures that the impact of process parameters on quality can be analyzed within 30 seconds.
[0070] Furthermore, in step 130, the standardization process specifically involves: classifying the preprocessed raw data according to preset data type classification rules, with each data type corresponding to a unique preset numerical range; and mapping each type of data to its corresponding range through linear transformation to ensure the stability of the dependent variable data sequence. With result variable data sequence The timing alignment meets the input requirements for subsequent process timing causal analysis.
[0071] It should be noted that in this embodiment, the standardization process in step 130 is a crucial link between data preprocessing and causal analysis. Its core purpose is to eliminate the interference of data heterogeneity on the analysis results while ensuring time sequence alignment. The preset data type classification rule refers to classifying data according to its physical meaning and process relevance (e.g., classifying "temperature and humidity" as environmental data, and "speed and torque" as equipment operation data). Each type corresponds to a unique numerical range to avoid cross-interference between different types of data (e.g., mapping environmental data to [0, 0.3], and equipment data to [0.3, 0.7]). The linear transformation is specifically performed using the formula... (Min-Max Transformation) or (Z-Score transformation) achieves mapping, ensuring that the transformed data retains the original fluctuation trend. "Time alignment" refers to calibrating through timestamps (e.g., uniformly adjusting the data sampling frequency to 1Hz) to ensure... and At the same time point Each has a corresponding value, avoiding misjudgment of causal relationships due to differences in sampling time. This is a prerequisite for the subsequent industrial time-series causal entropy algorithm to accurately calculate the "lag effect of the dependent variable on the effect variable".
[0072] In one possible implementation, the preset data type classification rules can be configured through an expert system (e.g., process engineers preset "welding current and voltage" as key process categories) and support dynamic updates based on historical data analysis (e.g., when "cooling water temperature" is identified as a high-frequency influencing factor, an "cooling system category" is automatically added); the range of linear transformation can be set according to the allowable fluctuation range of the process (e.g., key parameters are mapped to a finer range [0, 0.5], and minor parameters are mapped to [0.5, 1]), improving the identification of key data.
[0073] For example, in a feasible implementation, in a lithium battery welding process, step 130 classifies the pre-processed "welding current (100-200A)" and "welding time (0.5-2s)" as key process categories and maps them to the [0,0.5] interval; classifies "ambient temperature (20-30℃)" as an environmental category and maps it to the [0.5,1] interval; and calibrates the dependent variable sequence through timestamps. (Current + Time) and the sequence of result variables The sampling time for (weld strength) was 0.1s / sample to ensure that... Within a process window of 5 seconds, each Each moment has a corresponding and The value satisfies the requirement for temporal consistency in subsequent causal analysis.
[0074] Step 200: Based on the industrial time-series causal entropy algorithm, perform quantitative analysis on the process time-series causal intensity of the dependent variable data and the effect variable data in the standardized data sequence, and output the causal weight as the analysis result to provide a quantitative basis for subsequent management.
[0075] Specifically, step 200 further includes:
[0076] Step 210: Time-series correlation quantification, calculating the dependent variable data sequence With result variable data sequence Fluctuation synergy correlation ,in For the degree of correlation of fluctuation synergy, It is a time lag and satisfies To capture the correlation between abnormal fluctuations;
[0077] Step 220: Process constraint correction, introducing process causal constraint factors. Regarding the aforementioned fluctuation synergy correlation degree Make corrections to obtain the process compatibility correlation. Distinguish between statistical correlation and process causality;
[0078] Step 230: Calculate the industrial time-series causal entropy based on the process adaptation correlation. Calculate causal entropy Quantify the certainty of causal relationships;
[0079] Step 240: Causal weight mapping, which maps the causal entropy. Transformed into the causal weights The sum of the causal weights is 1, which provides a priority basis for index hierarchical management.
[0080] In this embodiment, the temporal correlation metric in step 210 aims to capture... and The volatility synergy, where the correlation degree of volatility synergy This is achieved by calculating the "normalized sum of the product terms of abnormal fluctuations"—when or Within the normal fluctuation range, corresponding items are not included in the correlation calculation; only the co-fluctuation information when both are abnormal is retained to avoid normal fluctuations interfering with causal judgment; lag time The introduction of this factor is to reflect the time delay characteristics of process causality (e.g., the effect of "change in cutting speed" on "surface roughness" may lag by 2 seconds). Step 220 introduces the "process causality constraint factor". (Values range from 0 to 1, set by the process manual or expert experience), used to filter statistically relevant but logically unrelated associations (such as a spurious association between "workshop lighting" and "machining accuracy"), while the process fit coefficient... This is used to adjust the lag time. Impact on correlation ( The closer The lower the probability of causality in the process, (This can weaken its weight). Causal entropy in step 230 The certainty of causal relationships is quantified using a logarithmic function (the smaller the value, the higher the certainty), and a fluctuation smoothing factor is used. Fluctuation amplification factor is used to avoid entropy jumps caused by outliers. This enhances the impact of significant abnormal fluctuations on the entropy value. Step 240 transforms causal entropy into causal weights. The weights are normalized to a total of 1, providing a clear priority basis for subsequent indexing and stratification (the higher the weight, the more important the association).
[0081] Additionally, in one possible implementation, step 210... The optimal value can be determined using a grid search method (such as in...). Traverse the range with a step size of 0.1s, and select... Maximum Step 220 It can be dynamically adjusted based on the process knowledge base (such as mature processes). =0.8, new process =0.5 to retain more potential associations); step 230 It can be adaptively selected based on the data noise level (the higher the noise, the better). Larger values enhance the smoothing effect.
[0082] Furthermore, in step 210, the fluctuation synergy correlation degree The calculation formula is:
[0083]
[0084] In the formula, The start time of the timing window. For dependent variable data sequence The mean within the time window For result variable data sequence The mean within the time window The threshold for fluctuation of the dependent variable data. For the threshold of fluctuation in the result variable data; when or At that time, in the numerator of the formula for calculating the degree of correlation of fluctuation synergy, the corresponding time point is... product term Set to 0;
[0085] In step 220, the process adaptation correlation The calculation formula is:
[0086]
[0087] In the formula, For process adaptability coefficient, It is a symbolic function; only when Then proceed to step 230;
[0088] In step 230, causal entropy The calculation formula is:
[0089]
[0090] In the formula, It is a volatility smoothing factor and , It is a volatility amplification factor and , It is an exponential function. , To strengthen causality, For weak causality, It is not causal;
[0091] In step 240, causal weight The calculation formula is:
[0092]
[0093] In the formula, To the dependent variable data sequence There are time-series related data sequences of all result variables. The number of data sequences for the result variable.
[0094] It should be noted that step 210... In the formula, the numerator captures the abnormal fluctuation product term by summing the terms. and Coordination anomaly (only when) and (When the product term is valid), the denominator is normalized to eliminate the influence of data magnitude, ensuring... The value is in the range [-1, 1] (positive values represent positive collaboration, negative values represent negative collaboration). Step 220 Formula passed Filter out non-process related elements and introduce... / The impact of the adjustment lag time ( The larger the value, the higher the weight of this item, but it is affected by... (Constraints), only when A value >0.3 indicates the existence of potential process causality, thus preventing weak correlations from being included in subsequent calculations. (Step 230) The formula uses a logarithmic function to... Convert to entropy ( ∈[0,1]), where " "Used to abnormality Mapping to [0,1] amplifies the impact of significant anomalies on the entropy value; The threshold division (<0.2 for strong causality, etc.) is set based on process experience to facilitate subsequent indexing hierarchy. Step 240... The formula is obtained through " "The entropy value is converted into a positive weight (the lower the entropy value, the higher the weight), and normalization is used to ensure that the sum of the weights is 1, so that the importance of different result variables can be directly compared."
[0095] In one possible implementation, step 210 and It can be set based on the 3σ principle ( The value is set to three standard deviations, i.e. ,in for The standard deviation), or directly specified by the process standard (e.g., the allowable fluctuation range for a certain parameter is ±0.01mm, then...). =0.01); Step 230 It can be adjusted according to parameter sensitivity (sensitive parameter) =3.0 to enhance the amplification effect, insensitive parameter =2.0).
[0096] For example, in one feasible implementation, in a bearing grinding process, Grinding wheel speed (average) =1500rpm, =50rpm), Roundness error (mean) =0.005mm, =0.001mm). =8s, , =2s. Step 210 calculates the molecular weight... =1s | -1500|=60>50、| -0.005|=0.0012>0.001, the corresponding product term is (60-50)×(0.0012-0.001)=10×0.0002=0.002, the rest... Since at least one of the time steps does not exceed the threshold, the numerator is set to 0, and the total numerator is 0.002; the denominator is calculated to be 0.01. =0.002 / 0.01=0.2. Step 220: Take... =0.9, =0.3, therefore =0.9×(0.2+0.3×1×2 / 8)=0.9×0.275=0.247 (≤0.3, analysis terminated, determined to be non-causal).
[0097] Step 300: Based on the causal weights, construct a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range, and causal weights, and construct a time-series causal tree index hierarchically according to the strength of causal association to achieve systematic management of the index structure.
[0098] Specifically, step 300 further includes:
[0099] Step 310: Based on process attributes, equipment operating status, data time series range, and causal weights The four core information categories are used to construct the four-dimensional dynamic index key, clarifying the core identifier dimension of the index, and the corresponding data time series range. ;
[0100] Step 320: Using the dependent variable data with strong causality as the root node, the effect variable data with strong causality as the first-level child nodes, and the effect variable data with weak causality as the second-level child nodes, excluding non-causal data, construct the hierarchical time-series causal tree index to achieve structured management of related data, where strong causality corresponds to... Weak causal correspondence Non-causal correspondence .
[0101] The four-dimensional dynamic index key is the core identifier of the indexing system. It contains four types of core information that are deeply integrated with the aforementioned technical solution: Process attributes specifically refer to the process type (e.g., milling, welding, stamping) and process number (e.g., MX-001, WJ-003) used to distinguish different production scenarios, ensuring that the index does not confuse cross-process data; Equipment operating status includes unique equipment ID (e.g., machine tool ID=M05, welding machine ID=W12), equipment health (e.g., 95%, 80%), and the production line to which the equipment belongs, facilitating accurate location of the equipment source corresponding to abnormal data; The data time series range clearly corresponds to the process time series window defined earlier. (The effective time range for characterizing causal relationships in processes) can limit the index to only cover time intervals where potential causal relationships exist, avoiding the inclusion of time ranges beyond which potential causal relationships exist. Invalid time-series data; causal weights The quantization result output in step 240 is then directly used to determine the index priority (the higher the weight, the higher the index priority, and the more frequently it will be called during retrieval). The purpose of constructing the index key in step 310 is to achieve precise data anchoring through multi-dimensional combinations, avoiding the problems of excessively large retrieval scope and low efficiency caused by single-dimensional indexes.
[0102] Furthermore, step 320, the time-series causal tree index, is a hierarchical storage structure based on causal strength. Its construction logic strictly follows the causal entropy judgment criteria mentioned earlier: "strong causality" in "the dependent variable data of strong causality is the root node" corresponds to the previous... For a result of < 0.2, selecting a strong causal dependent variable as the root node ensures that the core data in the index represents the most critical factors affecting production anomalies. The phrase "strong causal result variables are first-level child nodes" and "weak causal result variables are second-level child nodes" refers to a result of 0.2 ≤ 0.2. ≤ 0.5, this hierarchical approach allows for "first retrieving core related data (first-level child nodes) during retrieval, then expanding to secondary related data (second-level child nodes) as needed," significantly shortening the retrieval path; "excluding non-causal data" (corresponding to...) A value > 0.5 can reduce index redundancy and prevent unrelated data from occupying storage resources and interfering with search results.
[0103] In one possible implementation, the four-dimensional dynamic index key in step 310 can be generated as a hash key (e.g., "milling-M05-202511011000-0.8") using a combination format of "process attribute-equipment ID-time range start time-weight level" and stored in a hash table to achieve efficient retrieval with O(1) time complexity based on hash index, ensuring real-time location and fast access to data; the time-series causal tree in step 320 can be implemented using a red-black tree structure, with tree nodes storing the index address and causal weight of the data. When the causal strength of a certain causal variable changes (e.g., from weak causality to strong causality), the node level can be quickly adjusted through the rotation operation of the red-black tree to ensure that the index structure dynamically adapts to the changes in the causal analysis results.
[0104] For example, in a feasible implementation, in the spindle milling process (process attribute = "milling-MX001") of a machining workshop, step 310 constructs a four-dimensional dynamic index key: the process attribute is "milling-MX001", the equipment operating status is "equipment ID = M05, health = 92%", and the data time series range corresponds to... =20s (Time series start time = 20251101 08:30:00), Causal weight =0.8 ( Main spindle speed For milling accuracy, =0.15, strong causality). Step 320: Using "spindle speed data" as the root node, the strong causality "milling accuracy data" ( =0.15) as a first-level child node, weak causal "tool wear data" ( =0.35) is used as a second-level child node to exclude non-causal "workshop humidity data" ( =0.6), construct a time-series causal tree index with a red-black tree structure; when retrieving abnormal spindle speed data later, the system can first quickly locate the index key, and then directly obtain the milling accuracy data of the first-level child node from the root node, improving the retrieval efficiency by 60% compared with the traditional time-series index.
[0105] Step 400: Based on the historical data, analysis results, and real-time feedback data of the standardized data sequence, construct and self-learn to update the process timing semantic library, synchronously calibrate the core parameters of the algorithm, and ensure dynamic adaptation of analysis and management.
[0106] Specifically, step 400 further includes:
[0107] Step 410: Storage of the process timing semantic library Process causal constraint factors Process adaptability coefficient Fluctuation smoothing factor Fluctuation amplification factor Threshold for fluctuation of dependent variable data Threshold for fluctuation of result variable data The initial values and update records are used to automatically calibrate the above parameters when process parameters are adjusted or new processes are added.
[0108] Step 420: Based on the historical accumulated data of the standardized data sequence and the real-time feedback retrieval accuracy data, update the process time sequence semantic library. The self-learning adaptation time is ≤24 hours to ensure the adaptability of index construction and predictive retrieval.
[0109] It should be noted that the core objective of steps 410 and 420 in this embodiment is to construct a "self-learning and adaptable process timing semantic library" to solve the problem that traditional system parameters are fixed and cannot adapt to process changes. This provides a dynamically updated parameter benchmark for causal analysis algorithms and index construction, ensuring the timeliness and accuracy of the entire process technical solution.
[0110] Among them, the process timing semantic library mentioned in step 410 is the "core knowledge base" of the system, and the parameters stored therein are all key inputs for the preceding causal analysis and index construction: (The process timeline window, representing the effective time range of causal relationships) is the basis for step 210 to calculate the degree of synergistic correlation of fluctuations and step 310 to determine the data timeline range; process causal constraint factor This is the key coefficient for filtering spurious associations in step 220; the process adaptation coefficient. Used to adjust the impact of lag time on correlation; fluctuation smoothing factor Fluctuation amplification factor This is the core parameter for calculating causal entropy in step 230; the threshold for fluctuation in dependent / effect variable data. , This forms the basis for step 210, which determines whether the data fluctuates abnormally, and step 510, which triggers an abnormal signal. The "initial values" of these parameters are all derived from process manuals or historical experience data (such as data from a specific welding process). The initial value is set to 15 seconds according to industry standards. The "Update Record" stores detailed information about the time, reason, and values before and after each parameter adjustment, facilitating the tracing of parameter changes and meeting process audit requirements. The purpose of the "Automatically calibrate parameters when process parameters are adjusted or a new process is added" operation in step 410 is to ensure the system quickly adapts to changes in production scenarios (e.g., when the cutting depth of the original milling process is adjusted from 5mm to 8mm). (Synchronous calibration is required at 25 seconds) to avoid the lag and error of manual parameter adjustment.
[0111] Additionally, it should be noted that the self-learning update is a key mechanism for maintaining the accuracy of the semantic library. Its update is based on two types of core data: the historical accumulated data of the standardized data sequence refers to production data from a past period (e.g., 3 months) and the corresponding causal analysis results (e.g., in historical data). The accuracy rate of causal analysis was 92% at 20 seconds. =Accuracy rate of 95% at 18s), used to explore the correlation between parameters and analysis results; the real-time feedback retrieval accuracy data refers to the matching degree between the predicted retrieval results in step 500 and the actual cause of the anomaly (e.g., 85% of the retrieved correlation data is directly related to the actual anomaly), used to verify the rationality of the current parameter value. The setting of self-learning adaptation time ≤ 24 hours is to meet the needs of rapid adjustment of the production line (e.g., when adding a welding process for a new urgent order product, the system needs to complete parameter self-learning within 1 day to ensure normal operation during the next day's production), avoiding the impact on production progress due to an excessively long self-learning cycle.
[0112] In one possible implementation, the automatic parameter calibration in step 410 can employ a PID control algorithm: with the "cause-and-effect analysis accuracy" as the target value (e.g., set at 95%), when the analysis accuracy drops to 90% after process adjustments, the algorithm automatically calculates the parameter adjustment amount (e.g., adjusting the...). (Adjust from 0.8 to 0.85), and verify the effect of the adjustment with small batch data until the accuracy recovers to the target value; the self-learning update in step 420 can use the gradient descent method to optimize the parameters, with "retrieval accuracy" as the loss function, and iteratively calculate to make the parameter values converge to the optimal solution, while storing the historical optimal parameters as a backup to avoid parameter loss due to abnormal data.
[0113] For example, in one feasible implementation, the surface mount soldering process of an electronic component factory (original process parameters: soldering temperature 220℃, In the initial value = 10s), the process timing semantic library is initially stored in step 410. =10s =0.85、 =0.25、 =0.08、 =2.5、 =5℃ (welding temperature fluctuation threshold) =0.02mm (Patch offset fluctuation threshold). When a new high-melting-point component soldering process is added (soldering temperature adjusted to 250℃), step 410 automatically triggers parameter calibration, referring to historical high-temperature process data. Initially adjusted to 12 seconds; Step 420 performs self-learning based on the production data from the previous 24 hours (a total of 5000 sets of patch data): Through analysis, it was found that... The accuracy of causal analysis at 12s was 88%. Accuracy improved to 96% within 13 seconds, and retrieval accuracy improved from 82% to 93%. Therefore, it will be effective within 24 hours. The final update was 13 seconds, with simultaneous calibration. =0.88、 =0.22, ensuring the accuracy of causal analysis and retrieval under the new process flow.
[0114] Step 500: Monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it, so as to realize rapid response management of abnormal data.
[0115] Specifically, step 500 further includes:
[0116] Step 510: Monitor the standardized data sequence in real time, wherein the fluctuation threshold of the dependent variable data is... The fluctuation threshold of the result variable data is When any data in the standardized data sequence exceeds its corresponding fluctuation threshold, an abnormal signal is triggered, and the abnormal signal is associated with the corresponding data that exceeds the threshold.
[0117] Step 520: Based on the dependent variable data corresponding to the abnormal signal, call the updated process timing semantic library to predict the associated first-level child node data and second-level child node data from the timing causal tree index;
[0118] Step 530: Preload the index corresponding to the predicted associated data in advance. After receiving the retrieval request, synchronously output the root node data and all associated child node data corresponding to the abnormal signal. The retrieval response time is ≤1 second, realizing efficient management of abnormal data.
[0119] It should be noted that "real-time monitoring" uses edge computing nodes to stream standardized data sequences (e.g., sampling 1000 times per second). When the dependent variable data exceeds... Or if the variable data exceeds Immediately trigger an abnormal signal (such as "spindle speed = 2100 rpm"). =2000rpm"), the abnormal signal is associated with corresponding data that can trace the source of the abnormality (such as specific time, equipment, parameter value). "Predictive retrieval" is based on the dependent variable corresponding to the abnormal signal (such as spindle speed), and calls the causal association rules in the semantic library (such as "abnormal spindle speed is strongly correlated with surface roughness and vibration intensity") to locate the first-level child node (surface roughness) and the second-level child node (vibration intensity) from the time-series causal tree, realizing "predictive association data as soon as the abnormality occurs". "Pre-loading index" can complete the caching of associated data before the retrieval request arrives (such as memory preloading), ensuring that the response time is ≤1 second, meeting the real-time requirements of the production line for abnormality handling (such as avoiding batch scrap due to response delay).
[0120] In one possible implementation, the real-time monitoring in step 510 can employ a threshold double verification (e.g., triggering an anomaly only when three consecutive sampling points exceed the threshold, avoiding instantaneous interference); the predictive association in step 520 can be combined with a machine learning model (e.g., an association prediction model trained based on historical anomaly cases) to improve accuracy; and the index loading in step 530 can employ a hot and cold data separation strategy (storing high-frequency association data in memory and low-frequency data on disk) to optimize performance.
[0121] For example, in one feasible implementation, in an automotive stamping production line, step 510 involves real-time monitoring of the stamping pressure. ( =3000kN), when =10:05:23 =3200kN triggers an abnormal signal, with associated data being "Equipment ID=P03, Pressure Value=3200kN". Step 520 calls the semantic library to predict the first-level child node "Sheet Material Thickness Error" (strong causality) and the second-level child node "Die Temperature" (weak causality) from the causal tree. Step 530 pre-loads the indexes for these two types of data. When the operator issues a search request at 10:05:23.5, the system outputs the abnormal pressure data and associated thickness error and die temperature data within 0.5 seconds. For example, if the anomaly is ultimately confirmed to be caused by excessive die temperature, leading to changes in material deformation characteristics during stamping and thus causing abnormal pressure fluctuations, this helps technicians quickly locate and resolve the problem.
[0122] In this embodiment, the present invention solves the problem of single-dimensional indexing of traditional production data by constructing a multi-dimensional structured index system, which can quickly locate causal data and improve retrieval efficiency; relying on a self-learning and adaptable process time-series semantic library, it can dynamically calibrate core parameters, adapt to process changes, avoid the lag and error of manual adjustment, and ensure the accuracy of analysis; combined with the industrial time-series causal entropy algorithm to quantify causal strength and exclude non-causal data to reduce redundancy; coupled with real-time monitoring and predictive retrieval, it can quickly respond to abnormal data, providing strong support for efficient management and anomaly handling of production data.
[0123] Based on the above embodiments, such as Figure 2 As shown, the present invention also provides a big data analysis and management system for mechanical manufacturing production data, used to support the big data analysis and management method for mechanical manufacturing production data in the above embodiments. The big data analysis and management system for mechanical manufacturing production data includes:
[0124] The data acquisition and management module 11 is used to collect multi-source production data, and then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby realizing full-process data control.
[0125] Causal entropy analysis module 12 is used to perform quantitative analysis of the process time-series causal intensity of dependent variable data and effect variable data in the standardized data sequence based on the industrial time-series causal entropy algorithm, and output causal weights as analysis results to provide quantitative basis for subsequent management.
[0126] The index building and management module 13 is used to build a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range and causal weight based on the causal weight, and to build a time series causal tree index in layers according to the causal association strength, so as to realize the systematic management of the index structure.
[0127] The self-learning adaptation module 14 is used to build and self-learn to update the process timing semantic library based on the historical data, analysis results and real-time feedback data of the standardized data sequence, and synchronously calibrate the core parameters of the algorithm to ensure dynamic adaptation of analysis and management.
[0128] The predictive retrieval management module 15 is used to monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it, so as to realize rapid response management of abnormal data.
[0129] Monitoring module 16 is used to monitor the index building status, algorithm parameter accuracy, prediction hit rate, and data flow integrity to ensure stable system operation.
[0130] In an optional embodiment, the monitoring module 16 tracks the integrity and timeliness of index construction and updates to ensure that the index structure is consistent with the causal analysis results; verifies the accuracy of algorithm parameters and issues an early warning when the parameters deviate from a reasonable range; statistically predicts the hit rate of retrieval and evaluates the matching degree between retrieval results and actual anomalies; monitors the integrity of the entire process of data flow from collection to output to avoid data loss or delay. Through multi-dimensional monitoring, it ensures the stable operation of the system, provides reliable support for production data management, and can be connected to an industrial monitoring platform to display the system's operating status in real time through a visual interface, facilitating timely intervention by maintenance personnel.
[0131] In this embodiment, the present invention solves the problem of single-dimensional indexing of traditional production data by constructing a multi-dimensional structured index system, which can quickly locate causal data and improve retrieval efficiency; relying on a self-learning and adaptable process time-series semantic library, it can dynamically calibrate core parameters, adapt to process changes, avoid the lag and error of manual adjustment, and ensure the accuracy of analysis; combined with the industrial time-series causal entropy algorithm to quantify causal strength and exclude non-causal data to reduce redundancy; coupled with real-time monitoring and predictive retrieval, it can quickly respond to abnormal data, providing strong support for efficient management and anomaly handling of production data.
[0132] Furthermore, the big data analysis and management system for mechanical manufacturing production data can run the aforementioned big data analysis and management method for mechanical manufacturing production data. For specific implementation details, please refer to the method embodiment, which will not be repeated here.
[0133] Based on the above embodiments, such as Figure 3 As shown, the present invention also provides an electronic device, the electronic device comprising:
[0134] The processor 22 includes at least one processor 22, at least one memory 21, a communication interface 23, and a communication bus 24, wherein the processor 22 is communicatively connected to the memory 21.
[0135] In this embodiment, the memory 21 can be implemented in any suitable manner, for example, the memory 21 can be a read-only memory, a hard disk drive, a solid-state drive, or a USB flash drive, etc.; the memory 21 is used to store at least one executable instruction executed by the processor;
[0136] In this embodiment, the processor 22 can be implemented in any suitable manner. For example, the processor 22 can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc.; the processor is used to execute the executable instructions to implement the big data analysis and management method for mechanical manufacturing production data as described above.
[0137] Based on the above embodiments, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the big data analysis and management method for mechanical manufacturing production data as described above.
[0138] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, equipment, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or equipment, and may be electrical, mechanical, or other forms.
[0141] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0142] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0143] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program instructions, such as USB flash drives, portable hard drives, read-only storage servers, random access storage servers, magnetic disks, or optical disks.
[0144] Furthermore, it should be noted that the combination of the various technical features in this case is not limited to the combination methods described in the claims of this case or the combination methods described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way, unless they contradict each other.
[0145] It should be noted that the above examples are merely specific embodiments of the present invention, and the present invention is obviously not limited to the above embodiments, with many similar variations. All modifications that can be directly derived or conceived by those skilled in the art from the content disclosed in this invention should fall within the protection scope of this invention.
[0146] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A big data analysis and management method for mechanical manufacturing production data, characterized in that, Includes the following steps: Step 100: First, collect multi-source production data, then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby achieving full-process data control; Step 200: Based on the industrial time-series causal entropy algorithm, perform quantitative analysis on the process time-series causal intensity of the dependent variable data and the effect variable data in the standardized data sequence, and output the causal weight as the analysis result to provide a quantitative basis for subsequent management; Step 300: Based on the causal weights, construct a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range, and causal weights, and construct a time-series causal tree index hierarchically according to the strength of causal association to realize the systematic management of the index structure. Step 400: Based on the historical data, analysis results and real-time feedback data of the standardized data sequence, construct and self-learn to update the process timing semantic library, synchronously calibrate the core parameters of the algorithm, and ensure dynamic adaptation of analysis and management; Step 500: Monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it to realize rapid response management of abnormal data; Step 200 further includes: Step 210: Time-series correlation quantification, calculate the dependent variable data sequence With result variable data sequence Fluctuation synergy correlation ,in For the degree of correlation of fluctuation synergy, It is a time lag and satisfies To capture the correlation between abnormal fluctuations; Step 220: Process constraint correction, introducing process causal constraint factors. Regarding the aforementioned fluctuation synergy correlation degree Make corrections to obtain the process compatibility correlation. Distinguish between statistical correlation and process causality; Step 230: Calculate the industrial time-series causal entropy based on the process adaptation correlation. Calculate causal entropy Quantify the certainty of causal relationships; Step 240: Causal weight mapping, which maps the causal entropy. Transformed into the causal weights The sum of the causal weights is 1, which provides a priority basis for index hierarchical management; In step 210, the fluctuation synergy correlation degree The calculation formula is: ; In the formula, The start time of the timing window. For the dependent variable data sequence The mean within the time window For result variable data sequence The mean within the time window For the threshold of fluctuation in dependent variable data, For the threshold of fluctuation in the result variable data; when or At that time, in the numerator of the formula for calculating the degree of correlation of fluctuation synergy, the corresponding time point is... product term Set to 0; In step 220, the process adaptation correlation degree The calculation formula is: ; In the formula, For process adaptability coefficient, It is a symbolic function; only when Then proceed to step 230; In step 230, causal entropy The calculation formula is: ; In the formula, It is a volatility smoothing factor and , It is a volatility amplification factor and , It is an exponential function. , To strengthen causality, For weak causality, It is not causal; In step 240, causal weight The calculation formula is: ; In the formula, To the dependent variable data sequence There are time-series related data sequences of all result variables. The number of data sequences for the result variable.
2. The big data analysis and management method for mechanical manufacturing production data according to claim 1, characterized in that, Step 100 further includes: Step 110: Collect equipment operation data, process parameter data, and quality inspection data to form a raw data set; Step 120: Perform noise filtering and missing value completion preprocessing operations sequentially on the original dataset to remove invalid data; Step 130: Perform standardization on the preprocessed raw data, classify it according to data type and map it to a preset numerical range, so that the units of data of the same type are unified, and obtain the dependent variable data sequence. With at least one result variable data sequence ,in Represents the dependent variable data. Represents result variable data, The timestamps are consistent between the dependent variable data sequence and the result variable data sequence, and their time span corresponds to the process time window. Matching The effective time range for characterizing the causal relationship of processes.
3. The big data analysis and management method for mechanical manufacturing production data according to claim 2, characterized in that, In step 130, the standardization process specifically involves: classifying the preprocessed raw data according to preset data type classification rules, with each data type corresponding to a unique preset numerical range; and mapping each type of data to its corresponding range through linear transformation to ensure the stability of the dependent variable data sequence. With result variable data sequence The timing alignment meets the input requirements for subsequent process timing causal analysis.
4. The big data analysis and management method for mechanical manufacturing production data according to claim 1, characterized in that, Step 300 further includes: Step 310: Based on process attributes, equipment operating status, data time series range, and causal weights The four core information categories are used to construct the four-dimensional dynamic index key, clarifying the core identifier dimension of the index, and the corresponding data time series range. ; Step 320: Using the dependent variable data with strong causality as the root node, the effect variable data with strong causality as the first-level child nodes, and the effect variable data with weak causality as the second-level child nodes, excluding non-causal data, construct the hierarchical time-series causal tree index to achieve structured management of related data, where strong causality corresponds to... Weak causal correspondence Non-causal correspondence ; Step 400 further includes: Step 410: Storage of the process timing semantic library Process causal constraint factors Process adaptability coefficient Fluctuation smoothing factor Fluctuation amplification factor Threshold for fluctuation of dependent variable data Threshold for fluctuation of result variable data The initial values and update records are used to automatically calibrate the above parameters when process parameters are adjusted or new processes are added. Step 420: Based on the historical accumulated data of the standardized data sequence and the real-time feedback retrieval accuracy data, update the process time sequence semantic library. The self-learning adaptation time is ≤24 hours to ensure the adaptability of index construction and predictive retrieval.
5. The big data analysis and management method for mechanical manufacturing production data according to claim 1, characterized in that, Step 500 further includes: Step 510: Monitor the standardized data sequence in real time, wherein the fluctuation threshold of the dependent variable data is... The fluctuation threshold of the result variable data is When any data in the standardized data sequence exceeds its corresponding fluctuation threshold, an abnormal signal is triggered, and the abnormal signal is associated with the corresponding data that exceeds the threshold. Step 520: Based on the dependent variable data corresponding to the abnormal signal, call the updated process timing semantic library to predict the associated first-level child node data and second-level child node data from the timing causal tree index; Step 530: Preload the index corresponding to the predicted associated data in advance. After receiving the retrieval request, synchronously output the root node data and all associated child node data corresponding to the abnormal signal. The retrieval response time is ≤1 second, realizing efficient management of abnormal data.
6. A big data analysis and management system for mechanical manufacturing production data, characterized in that, include: The data acquisition and management module is used to collect multi-source production data, and then perform standardized preprocessing on the multi-source production data to obtain a standardized data sequence, thereby realizing full-process data control. The causal entropy analysis module is used to perform quantitative analysis of the process time-series causal intensity of dependent variable data and effect variable data in the standardized data sequence based on the industrial time-series causal entropy algorithm, and outputs the causal weight as the analysis result to provide a quantitative basis for subsequent management. The index building and management module is used to build a four-dimensional dynamic index key that integrates process attributes, equipment operating status, data time series range and causal weight based on the causal weight, and to build a time-series causal tree index in layers according to the strength of causal association, so as to realize the systematic management of the index structure. The self-learning adaptation module is used to build and self-learn to update the process timing semantic library based on the historical data, analysis results and real-time feedback data of the standardized data sequence, and synchronously calibrate the core parameters of the algorithm to ensure dynamic adaptation of analysis and management. The predictive retrieval management module is used to monitor the standardized data sequence in real time, obtain real-time abnormal data that triggers the fluctuation threshold, trigger predictive retrieval based on the updated process time sequence semantic library and the time sequence causal tree index, load the corresponding related data and output it, so as to realize rapid response management of abnormal data. The monitoring module is used to monitor the index building status, algorithm parameter accuracy, prediction hit rate, and data flow integrity to ensure stable system operation. The causal entropy analysis module is also used for: Time-series correlation quantification, calculating the dependent variable data sequence With result variable data sequence Fluctuation synergy correlation ,in For the degree of correlation of fluctuation synergy, It is a time lag and satisfies To capture the correlation between abnormal fluctuations; Process constraint correction, introducing process causal constraint factors. Regarding the aforementioned fluctuation synergy correlation degree Make corrections to obtain the process compatibility correlation. Distinguish between statistical correlation and process causality; Industrial time-series causal entropy calculation, based on the process adaptation correlation degree. Calculate causal entropy Quantify the certainty of causal relationships; Causal weight mapping, which maps the causal entropy Transformed into the causal weights The sum of the causal weights is 1, which provides a priority basis for index hierarchical management; The degree of synergistic correlation of fluctuations The calculation formula is: ; In the formula, The start time of the timing window. For dependent variable data sequence The mean within the time window For result variable data sequence The mean within the time window The threshold for fluctuation of the dependent variable data. For the threshold of fluctuation in the result variable data; when or At that time, in the numerator of the formula for calculating the degree of correlation of fluctuation synergy, the corresponding time point is... product term Set to 0; The process adaptability correlation The calculation formula is: ; In the formula, For process adaptability coefficient, It is a symbolic function; only when At that time, the industrial time-series causal entropy calculation is initiated; The causal entropy The calculation formula is: ; In the formula, It is a volatility smoothing factor and , It is a volatility amplification factor and , It is an exponential function. To strengthen causality, For weak causality, Non-causal The causal weight The calculation formula is: ; In the formula, To the dependent variable data sequence There are time-series related data sequences of all result variables. The number of data sequences for the result variable.
7. An electronic device, characterized in that, The electronic device includes: The processor and the memory are communicatively connected. The memory is used to store at least one executable instruction executed by the processor, the processor being used to execute the executable instruction to implement the big data analysis and management method for mechanical manufacturing production data as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the big data analysis and management method for mechanical manufacturing production data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Hot continuous rolling process fault diagnosis method and device based on multilayer dynamic causal diagram
CN120763513A