Industrial internet big data platform and data processing method

By constructing a tensor that associates field response paths with role labels, the problem of disordered data structure in industrial big data is solved, enabling efficient feature filtering and intelligent data access, and improving data retrieval efficiency and model accuracy.

CN120893020AActive Publication Date: 2025-11-04XIAN AERONAUTICAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511418115.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-04
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

In complex industrial scenarios, data sources are diverse and their structures vary significantly. Existing technologies struggle to achieve unified modeling and analysis of the data, leading to mismatched labels, incorrect field links, and inconsistent behavioral logic. Furthermore, data retrieval efficiency is low and access control is inadequate.

Method used

By constructing a tensor that associates field response paths with role labels, we identify label backtracking and structural misgrouping behaviors. We use a minimum common structure behavior cycle sliding window mechanism to repair the path, and combine K-means clustering, linear correlation analysis and random forest scoring mechanism to screen high-value variables. We then introduce a behavior feature-driven caching recommendation strategy.

Benefits of technology

It improves data structure consistency, achieves efficient feature dimension compression and system response efficiency, and ensures the accuracy of model input and intelligent data access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893020A_ABST
    Figure CN120893020A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial internet big data platform and a data processing method, relates to the technical field of industrial big data processing, and is used for solving the problems of unified structure attribution, field tag mismatching, permission state drift and high-quality modeling data screening of multi-source heterogeneous data. Comprising the links of multi-source data acquisition, field structure attribution judgment, field label behavior correction, modeling data generation, prediction analysis, cache recommendation and the like. According to the method, through unified collection and structured processing of process technological parameters, three-dimensional point cloud, user logs and smelting process data, a behavior tensor containing time, paths and role tags is constructed, field wrong group and tag turn-back behaviors are identified, and structure repair and redundancy cleaning are executed. Discretization, feature screening and rule mining are carried out on variables, data access response and model input quality are optimized, and the method is suitable for high-dimensional data management and intelligent modeling processes in complex industrial scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial big data processing, and more particularly, to an industrial internet big data platform and a data processing method. BACKGROUND

[0002] In a complex industrial scene, data sources are extensive and structural differences are obvious, including process parameters, image point clouds, user logs and message middleware data. These data have significant heterogeneity in acquisition methods, timing characteristics, field structures and permission tag ownership, making it difficult to be directly used for unified modeling analysis. In the prior art, simple integration is usually performed through field name mapping and data type standardization, without dynamic identification and classification processing of label mismatch caused by role switching, log structure chain drift or permission disconnection, resulting in problems such as label pollution, field chain error and inconsistent behavior logic in model input. In addition, in traditional prediction modeling, full-field input is usually used, and there is a lack of joint filtering mechanism based on linear and nonlinear influence degree, so that redundant fields affect the efficiency of the model. At the same time, in the big data access link, a cache recommendation strategy based on behavior characteristics is not introduced, resulting in low data calling efficiency and extensive permission control.

[0003] In view of the above problems, the present application provides a solution. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide an industrial internet big data platform and a data processing method to solve the problems raised in the background art.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical scheme: In a preferred embodiment, it comprises: Collecting process industrial process parameters, three-dimensional visual point clouds, distributed message smelting data and user operation logs, and binding collection time stamps, device identifiers and channel numbers; Introducing field first response channel number, field ownership role label and field write time offset, constructing role state transition track and field response chain, and constructing Γ(t,x,r) tensor in a fixed sliding time window; Taking Γ(t,x,r) tensor as the core state index to calculate the minimum common structure behavior period λn, performing structure ownership consistency analysis to mark the permission state disconnection field, and dividing each field belonging to the permission state disconnection field into Drift, Overlap, Lag or Pseudo type; Performing differential path repair when the field behavior window covers λn and the ownership of the segment is completed, and reconstructing and outputting the complete field ownership path set.

[0006] In a preferred embodiment, the acquisition process industrial process parameters, three-dimensional vision point cloud, distributed message smelting data and user operation log, and bind the acquisition time stamp, device identifier and channel number.

[0007] In a preferred embodiment, the first response channel number, field ownership role label and field write time offset are introduced, then the permission state record table is queried for switching events by taking the field ownership role label as the track node and combining the field write time, and the field node sequence is established by taking the field structure path number and the field response sequence, to construct the role state transition track and the field response chain; In a fixed sliding time window, the above-mentioned role state transition track and field response chain are synchronized and aligned with the user operation session as the boundary, and for each field set in each time period, Γ(t, x, r) tensor is constructed by three-dimensional sparse matrix construction algorithm and combined with hash mapping index according to time t, structure path number x and role label r.

[0008] In a preferred embodiment, by taking Γ(t, x, r) tensor as the core state index, combining role label foldback set, field disintegration chain foldback re-aggregation path graph and window drift recovery table, role label consistency judgment, path ownership continuity verification and window focus state verification are performed to construct field disintegration chain foldback re-aggregation path graph and window drift state segment; According to the minimum time Tbreak from the first label foldback chain interruption to the closed loop reconstruction, the required time Treturn(x) for the field x to be misclassified and foldback to the main path, and the shortest regression cycle Trebind(w) for the window w to re-bind the role context and trigger the correct binding of field ownership, three types of time quantities, self-correlation analysis of field structure path sequence is adopted, and the main period component is calculated and generated by combining fast Fourier transform to extract the minimum common structure behavior period λn.

[0009] In a preferred embodiment, taking λn as the sliding time window length, slice in Γ(t, x, r) tensor according to time dimension, and perform structure ownership consistency analysis on the tensor index set in each sliding time window to mark the permission state disjoint field.

[0010] In a preferred embodiment, taking λn as the sliding time window length, traversing Γ(t, x, r) tensor, calling role state backtracking function to generate main role path sequence according to r-dimensional track, and matching each field structure path belonging to permission state disjoint field with the main path structure set to identify behavior patterns such as path jump, label inconsistency, response out of order, role not switched and causal path missing; According to the matching result, each field belonging to the permission state disconnection field is divided into a main path drift segment Drift, a role overlap segment Overlap, a permission lag segment Lag, or a pseudo-closed window segment Pseudo.

[0011] In a preferred embodiment, after the field perturbation segment division is completed, when the field's behavior window completely covers the minimum common structure behavior period λn, and the segment to which the field belongs completes the structure belonging preliminary judgment and path trajectory convergence within the period, a field segment type identifier Ltype is inserted, and differential path repair is performed according to the field Ltype field; Then, the structure state sequence of all classified fields is extracted from the Γ(t, x, r) tensor with the minimum common structure behavior period λn as the time index window. Based on the field return path set, the field role label re-aggregation sequence and the log channel writing trajectory are combined to perform the final reconstruction of the field belonging label path, forming a complete field belonging path set.

[0012] In a preferred embodiment, the complete field belonging path set is discretized and jointly evaluated to generate a modeling field and extract point cloud features to identify the pose of the target object, and to mine discrete fields to generate a field priority list and a structured knowledge rule table; Then, the environmental load output carbon emission value is calculated by marking resources and emission factors according to the process stage; a time series prediction model is constructed to predict the target variable; and a cache scheduling strategy is generated based on access frequency, security level, and permission level to complete data access and update.

[0013] In a preferred embodiment, it includes: a multi-source data access and structure binding module, a data preprocessing and field belonging analysis module, a field belonging correction and modeling field generation module, a model construction and cache scheduling management module, and signal connections between the modules; The multi-source data access and structure binding module accesses four types of data: process parameters, point clouds, smelting messages, and user logs, and completes protocol configuration, physical channel binding, and structure context ternary information writing; The data preprocessing and field belonging analysis module performs missing completion, topology reconstruction, behavior chain analysis, and Γ tensor construction on time series, point cloud, and log data, respectively, to establish the association between field structure paths and role labels; The field belonging correction and modeling field generation module analyzes field return, label drift, and path break based on the minimum belonging period, completes Ltype classification, field belonging chain repair, and pseudo-merged field elimination, and generates a modeling field set; The model construction and cache scheduling management module discretizes continuous variables, models point cloud objects, and performs time series prediction modeling, and simultaneously performs cache scheduling and permission verification according to access frequency, security level, and permission level.

[0014] The technical effects and advantages of the industrial internet big data platform and data processing method of this invention are as follows: This invention constructs a field response path and role label association tensor to identify label reversal and structural misgrouping behaviors, ensuring closed-loop repair of field attribution chains before modeling and improving data structure consistency. Through a λn minimum period sliding window judgment mechanism, it classifies and corrects field drift, overlap, lag, and pseudo-merging behaviors, resolving structural attribution confusion caused by permission status changes. In terms of modeling data selection, it integrates K-means clustering, linear correlation analysis, and random forest scoring mechanisms to accurately extract high-value variables and effectively compress feature dimensions. For point cloud data, it extracts the principal direction and PPF descriptors to construct an index model, improving 3D recognition efficiency. At the data access end, it introduces a ternary influence factor model to dynamically generate caching strategies based on user behavior, data security, and permission levels, improving system response efficiency. The overall solution possesses advantages such as closed-loop attribution repair, accurate feature selection, and intelligent access response, making it suitable for multi-source industrial big data environments. Attached Figure Description

[0015] Figure 1 This is a timing diagram of an industrial internet big data platform and data processing method according to the present invention.

[0016] Figure 2 This is a schematic diagram of an industrial internet big data platform and data processing method module according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example: This invention discloses an industrial internet data processing method, such as... Figure 1 As shown, it includes: First, a multi-source data channel scheduling mechanism is used to sequentially complete the acquisition and activation process for four types of core data: process parameters of the process industry, 3D visual point clouds, distributed message smelting data, and user operation logs. For each data source, specific communication protocols and physical access paths are configured according to the data channel registration table. Before data import, a three-factor binding of the acquisition timestamp, device identifier, and channel number is completed to ensure the data has a complete acquisition context.

[0019] Among them, for the core time sequence variables such as temperature, pressure and ratio in process industry, the activation operation of the collection node is completed by the automatic collection sub-deployed in the control, the collected data is transmitted into the real-time buffer through the wired bus, and the industrial process data main index pool is written with millisecond level timestamp.

[0020] For the shape recognition task involved in the industrial sorting scene, the on-site deployed deep vision device is used as the data source, the physical connection path with the device is actively established after power-on, and the three-dimensional spatial structure data at each moment is received in the form of image frame pairs. Such point cloud data directly binds spatial coordinates, color information and depth values during the collection process, ensuring that a complete three-dimensional point set can be constructed later.

[0021] For the electrolyte temperature and current density data in the aluminum electrolysis smelting process, the pre-configured industrial message receiver is used to establish a distributed middleware connection, and after completing the data topic subscription, the process message reported by the edge node is continuously received and automatically written into the time series record table.

[0022] For user operation behavior data, automatically start the log data monitoring channel during runtime, monitor user-side browsing behavior logs and server-side database call logs respectively, and preliminarily aggregate them according to time, IP address and user number when the record is generated, forming a multi-source log record stream.

[0023] After completing the data import, immediately start the multi-type data preprocessing process, and execute the differentiated processing path according to the data type label.

[0024] Specifically, for time series continuous variables from process control, a fixed-length time sliding window is used to scan the data curve, and after detecting the missing data points, numerical fitting completion operation is performed combined with the valid data points before and after the window; After completion, based on the median and deviation threshold of historical sampling, perform outlier detection and rejection operation to ensure that extreme disturbance data does not enter the modeling channel; After data cleaning, based on the global maximum and minimum value of the current variable, perform normalization processing to compress the data standard to a unified numerical interval, ensuring that all parameters have the same scale in subsequent modeling operations.

[0025] For the collected three-dimensional point cloud data, first construct a topological structure index graph based on the spatial coordinates in the point set to support neighbor search and structure preservation analysis; On this basis, perform boundary region rejection operation to exclude far-sight background points without semantic content; Then count the neighborhood density of each point in the point set, identify isolated points and high deviation points and perform cleaning operation; Finally, apply spatial voxel division rules to the entire point set, project the point cloud into a cubic grid, reconstruct the point set according to the representative point of each voxel, and realize downsampling processing, which not only preserves the target geometric shape but also reduces the computational load.

[0026] For user operation log data, first, the log records are merged into user session chains in chronological order and IP address, and then a preset field parsing rule set is loaded to perform semantic segmentation and structural field extraction on the log text; Based on the extraction results, field name mapping and data type reconstruction are performed to ensure that all log fields have a unified structure, standard field name and data type before entering the behavior modeling module, and the preliminary binding of field labels and semantic paths is completed.

[0027] It should be noted that in the behavior monitoring scenario in the multi-department joint debugging stage, if deployed in an environment with a complex operation permission model and concurrent virtual desktop access structure, especially when performing cross-post permission simulation testing or multi-role concurrent switching tasks, the phenomenon of structural path misalignment caused by role permission switching lag will occur. This phenomenon only occurs when the following three conditions are met simultaneously in the user operation log: the user account is bound to multiple role permissions, the front-end behavior is loaded by different permission control modules; the log upload channel is a non-blocking asynchronous stream structure; the log data stream is asynchronously merged at multiple permission buffer nodes, and the permission switching and field response events are not synchronized. At the same time, in the log record, multiple role identity switching residual field segments will appear in the same behavior segment in the log record, which will be misidentified as complete behavior chain writing to the structured channel.

[0028] In existing technical means, the log field structuring process mainly aggregates all log records by timestamp and IP to extract fields by regular parsing engine, and directly performs unified naming and standardization processing according to field name and data type. This process does not introduce field attribution path verification, role state locking mechanism or asynchronous field channel identification strategy. The default structure field belongs to the same behavior semantic path, and the field set construction is performed with the behavior window as the granularity. This results in the misalignment of field collection order and real behavior path order, causing the field attribution chain to break; the residual fields of multi-role switching are misidentified as valid fields, causing field label redundancy expansion; and the permission state drift does not trigger field attribution switching, causing multiple role fields to mix into the same behavior window. Therefore, in this embodiment: To solve the problem of permission label switching lag and field structure path misalignment cross interference, three types of structure meta-information are introduced in the log field structure body, namely field first response channel number, field attribution role label and field write time offset.

[0029] Wherein, the field first response channel number is generated and written into the field index domain by the log access controller according to the channel registration number when the field first enters the structure parser buffer; the field belonging role label is derived from the permission binding module loaded by the front-end virtual desktop, and is injected by the UI controller when the field is generated and uploaded with the field; and the field write time offset is calculated from the difference between the structure field write timestamp and the last switching effective time in the current role switching record, and the difference is generated and hung into the structure meta information field group before the field enters the structure belonging judgment module.

[0030] Further, in the structure parsing stage, the user operation session is taken as the boundary, the field belonging role label sequence and the structure path number sequence of all fields are collected within the user operation session, and the role state transition trajectory is generated in combination with the field write time.

[0031] The role state transition trajectory is generated in the following manner: taking the field belonging role label as the trajectory node, if the adjacent field role labels are inconsistent, the platform queries whether there is a switching event between the two from the permission state record table; if the switching event exists and the switching effective time is not greater than the current field write time, a transition edge is added in the trajectory to mark the starting time of the switching event and the target role; otherwise, it is marked as a label drift point and written into the trajectory drift node set.

[0032] At the same time, the field response chain is constructed according to the field structure path number and the response sequence, and each node on the field response chain represents the structure path position of the field response, and the arrangement order is consistent with the actual time of the field response, so as to ensure that the response order of the path sequence is consistent with the physical structure loading order.

[0033] After the role transition trajectory and the field response chain are constructed, the two sequences are synchronized and aligned in a time sliding intersection manner, and a three-dimensional disturbance structure tensor Γ(t, x, r) is constructed in each sliding window time period, where t is the field response time index, x is the structure path number of the field response, and r is the role label bound when the field is submitted.

[0034] It should be noted that the Γ(t, x, r) tensor is constructed by a three-dimensional sparse matrix construction algorithm combined with a hash mapping index. In an optional example, the key value design of the hash mapping adopts a three-level coding rule: Timestamp t: accurate to millisecond level, for example, in the format of: YYYYMMDDHHMMSSmmm; Path x: adopting three-level coding of "workshop-line-equipment", such as C02-L03-M15, corresponding to No. 2 workshop No. 3 line No. 15 equipment; Role r: coded according to the permission level, such as 01-administrator, 02-operator, 03-monitor; The sparsity threshold of the three-dimensional sparse matrix is set to 90% (high-frequency fields are retained, and the rest are stored in compressed form).

[0035] Further, with the Γ tensor as the core state index, the role label consistency judgment and path attribution continuity verification are performed on the fields within each operation session. If a field is marked with different role labels multiple times in the behavior window, and its response structure path is located in the merging area of multiple permission buffer nodes, the platform records the role label turn-back behavior of the field and writes it into the field label turn-back set; If the field reappears with the main path role label in the subsequent fields in the asynchronous channel after its first appearance in the structure path attribution chain, it is taken as a valid return node to construct the field disintegration chain turn-back reunion path graph, which constructs the time span from the initial attribution misclassification time of the field to the reunion time, and marks the structure path segment in the return path and the role label change process.

[0036] Subsequently, while identifying the field label turn-back behavior, the window focus state monitoring is started, and the focus state of each field-attached window number wi is verified. If the window has experienced role state switching before the field submission time, but the window focus state record has not been refreshed, the platform marks the window to enter the context interruption state; then, the structure attribution recombination action is recorded when the window regains focus for the first time, the window drift state segment is generated, and the start and end times of the segment and the role recovery path are written into the window drift recovery table.

[0037] After the above three types of behavior recognition are completed, the turn-back behavior time sequence construction is uniformly called to extract the complete time period of the field attribution state from mismatch to recovery in each type of behavior, and the minimum common structure behavior period λn is constructed with the following three time parameters, according to the formula: λn=sup{min{t∈Tbreak}Δt,min{x∈Φ}Treturn(x),min{w∈Wdetach}Trebind(w)}; Where mint∈TbreakΔt is the minimum time consumption from the first label turn-back chain interruption to the closed-loop reconstruction; Treturn(x) is the time required for field x to be misclassified and turn back to the main path; Trebind(w) is the shortest return period for window w to rebind the role context and trigger the correct binding of the field attribution.

[0038] In an optional specific example, assuming that the steel plant continuous casting equipment historical session data λn calculation example based on a sample size of 1200 production shifts is as follows: Extract Tbreak from Γ(t,x,r) tensor: 327 label turn-back interruption events are screened, the time difference Δt from interruption to reconstruction is calculated, and the minimum value is 480 ms; Calculate Treturn(x): For 15 main paths Φ, count the shortest time of device regression after offline, the minimum value is 320ms; Calculate Trebind(w): For 89 role labels in the detach state, calculate the shortest regression period using the 75% quantile of the ISO42000:2023 standard, and the result is 510ms; Substitute the formula to get λn=sup{480ms,320ms,510ms}=510ms, and the sliding window length is set to 2 times of λn(1020ms).

[0039] At the same time, the platform calculates Tbreak, Φ, Wdetach and their corresponding quantities according to the unified process: using Γ(t, x, r) as an index, the field response time series within the same user operation session is traversed in a sliding manner; the candidate set of Tbreak is calculated by Δt from the first appearance of mismatch to the start and end time of the closed loop of the role label turn-around set; Φ is the set of effective regression nodes in the field collapse chain turn-around recombination path graph, Treturn(x) takes the difference between the misclassification time of each field x and the regression main path time; Wdetach comes from the window drift recovery table, Trebind(w) takes the time difference between the window w entering the context interrupt state and the first recombination action completion. Take the minimum value of the three types of time quantities in their respective sets, and then take the maximum of the three in each session as the current session λn, which is used as the sliding time window length of Γ(t, x, r) slice and the time sequence boundary of consistency evaluation.

[0040] After that, the maximum value of the above three times in each operation session is taken, and the λn period corresponding to the current field sequence is generated as the time sequence boundary of the structure attribution consistency evaluation.

[0041] It should be noted that the existence of λn itself depends on the mutual triggering chain relationship of the triple residual mechanism, and any structure such as channel synchronization, single role, and front-end window drift absence cannot derive this period. Therefore, λn is only applicable to the scenario where role label turn-around, field attribution collapse, and structure path delay recombination occur in the current permission joint debugging state.

[0042] Subsequently, with λn as the sliding time window length, slice in the Γ(t, x, r) tensor according to the time dimension, and perform structure attribution consistency analysis on the field set in each sliding time period. If multiple fields appear continuously in a certain λn time window, and their attribution path numbers and role label combinations are not in the effective combination set of the main role path and the structure response chain, the platform marks the field as a permission state disengaged field, and marks the entire time window as a disengaged segment, and writes it into the structure behavior classification queue.

[0043] Specifically, the platform generates non-overlapping slices on Γ(t,x,r) with the minimum timestamp in the session as the starting point and λn steps, and performs consistency check on the set {(t,xi,ri)} in each slice: first generate the main role path sequence from the r-dimensional trajectory, then calculate the path consistency score based on the mapping relationship between xi and the main path structure set; when the score is lower than the configured threshold and there is no regression record in the turnaround chain, mark it as the authority state dislocation field, and the slice where it is located is recorded as the dislocation segment.

[0044] After determining the authority state dislocation field, immediately start the structure disturbance segment classification process, and perform window-in traversal operation on Γ(t,x,r) tensor with λn period as the sliding window length, to identify all field attribution chain breakage behaviors and path abnormal segments. This operation takes the three-dimensional index of the tensor as the entrance, extracts all tensor slices within the λn time window according to the field response time t, and synchronously reads the structure path number x and role state r corresponding to each field.

[0045] Then, in each λn time window, first call the role state backtracking function to generate the main role path sequence according to the r-dimensional trajectory in the tensor, as the main role state baseline in this window segment. Then, perform path matching calculation on each field structure path xi and the main path structure set in turn, to identify path jump, label inconsistency, response disorder, role switching, and causal path missing behaviors. The specific structure segment division logic is as follows: If the first response bound role label ri of the field is inconsistent with the main role state in this λn period, and the field corresponding path xi in the structure response chain cannot establish an effective mapping relationship with the main path structure set, the platform marks the segment where the field is located as the main path drift segment, defined as type Drift; If the field is repeatedly written by two different role labels in this λn period, and the response order of these fields in the structure response chain is inconsistent or multiple attribution paths appear, the platform marks this segment as the role overlap segment, defined as type Overlap; If the field response time ti is located in the window interval before and after the role switching event, and the field is still bound to the previous role label r(i-1), causing the label state to be not updated synchronously, the platform determines that this segment is the authority lag segment, defined as type Lag; If the field is attributed to the same structure chain in this λn period, but no behavior causal trigger relationship is established between its structure paths, and it is only aggregated because the response time is close, the platform marks this segment as the pseudo-closed window segment, defined as type Pseudo.

[0046] Then, the platform performs uniform parameterization judgment on the field segment within the time window of λn: for the Overlap type, the proportion of field intersection write times to the total write times in the window is taken as the intersection frequency indicator, and the average of the path similarity matrix generated by the structure path function label mapping table is taken as the path coincidence degree indicator; when the intersection frequency and the path coincidence degree are not lower than the empirical quantile threshold obtained by the system based on historical session statistics during the training period, it is confirmed as Overlap.

[0047] In an optional example, the path coincidence degree threshold setting method is as follows: Based on the historical 1000 normal session data, the cosine similarity of the main path and the branch path is calculated; The MAD algorithm is used to calculate the robust threshold, and the result is 0.82, so the coincidence degree ≥ 0.8 is set to trigger the Overlap judgment; The intersection frequency statistics use sliding window counting, the window length = λn, the step = λn / 2, and the intersection write times ≥ 3 times in each window are judged as valid Overlap.

[0048] It should be noted that for the Drift and Lag types, the proportion of unmappable and the duration of label unsynchronization after role switching is taken as the trigger, and the threshold is also robust quantile (such as median upper quantile) of historical sessions, which is fixed as a configuration item during deployment. The Pseudo type needs to meet the conditions of no causal dependence between paths and only due to close response time. The platform takes the failure of causal test and the time proximity exceeding the window neighbor statistics threshold as the judgment basis. All thresholds come from the statistical learning process of historical undisturbed / disturbed samples, and the write configuration table is written during deployment to ensure consistency.

[0049] After the field disturbance segment division is completed, the platform inserts the field segment type identifier Ltype in the field structure attribution field set, and the value range is {Drift, Overlap, Lag, Pseudo}. At the same time, the platform sets the write constraint condition, and only when the field behavior window completely covers the λn period and its belonging segment completes the structure attribution preliminary judgment and path trajectory convergence in the period, the Ltype label is written into the structure body, avoiding early classification of the fields still in the link return process.

[0050] After entering the label-driven scheduling stage, the platform performs differential path repair actions according to the field Ltype field. The specific operation logic is as follows: For the field segment marked as Drift and Lag, the platform first verifies whether the field has formed a complete turnaround chain in the λn period. If it is confirmed that the field has been identified by label drift and completed the main path regression, the platform calls the r-dimensional role trajectory of the Γ(t, x, r) tensor, performs field attribution chain reconstruction operation, backfills the role label and path number of the field according to the main path state, and repairs the attribution path breakpoint; For the Overlap type field segment, the platform counts the number of role label types and switching frequency bound by the field in the segment. If the field cross-write times exceed the threshold in the λn period, and the path overlap degree is higher than the configuration comparison threshold, the platform calls the structural path function label mapping table, performs field path function label comparison operation, further identifies and strips the redundant field structure path through the similarity matrix, and clears the semantic overlap segment; For the Pseudo type segment, the platform checks whether there is a logical causal dependence relationship in the structural path response chain. If no valid causal event chain can be detected in the field sequence, the platform determines that the segment is a false splicing aggregation result, calls the structural behavior elimination module, marks the segment field structure behavior as a pseudo-merged field, and logs out its structural path record from the attribution chain structure mapping table to avoid entering the modeling process.

[0051] In the field structure merging phase, With λn period as the time index window, the structural state sequence of all classified fields in the period is extracted from the Γ(t, x, r) tensor. Based on the field turnaround path set, the system jointly executes the final reconstruction of the field attribution label path with the field role label re-aggregation sequence and the log channel write trajectory, and forms a complete field attribution path set.

[0052] After the field attribution path set is written, the platform outputs the structural field set to the modeling module with the set as the only input source. In the set, the field role state has been closed loop, the label state has been confirmed consistent, and the structural attribution chain has been completed backtracking repair, ensuring that the structural path misassembly, label drift pollution, and permission error chain delay problems have been completely eliminated before attribution judgment, and the unique legal entrance of the field semantic structure in the model training stage is constructed.

[0053] Subsequently, supervised discretization is performed for continuous variables in process industry. The process takes the numerical distribution of variable historical sampling data as input, calls the dynamic clustering engine to perform multiple rounds of clustering boundary calculation operations in the variable dimension. Specifically, the platform loads the numerical distribution curve of the current variable in the collection period, takes the unique ID of each variable field as the index, inputs the K-means cluster, performs clustering division to obtain the preliminary interval boundary. The platform synchronously calls the expert process weight configuration table, marks the main control variable affecting the core production index as a fine-grained parameter, and performs an upward processing on the number of cluster centers to allow denser interval division; for auxiliary parameters in the edge correlation, the clustering dimension controller proportionally reduces the number of intervals to perform coarse discretization. Finally, the discrete label of each variable is written into the feature encoding dictionary table with variable ID as the key and discrete level as the value.

[0054] For time series type process parameters, the platform first performs joint evaluation of linear and nonlinear influence strength on each field to screen a set of high-value modeling fields. In the linear evaluation stage, the platform takes the field value matrix as input to calculate the Pearson correlation coefficient matrix between variables to identify the linear influence degree of the main control variable on the target index; in the nonlinear evaluation stage, the platform constructs a random forest regressor model with the current field set as input, and takes the feature importance score as the nonlinear influence degree indicator of the variable.

[0055] Meanwhile, the platform fuses the linear and nonlinear score results through a weighted rule, loads the field priority confirmation table provided by the process expert to cross-verify the model results, and generates a field priority list. All screening results are registered in a unified field ID mapping table.

[0056] Specifically, the linear evaluation takes the absolute value of the correlation coefficient matrix as the score; the nonlinear evaluation takes the feature importance of the random forest regressor as the score; the two are fused by weighting, and the weight is accurately determined and fixed during deployment according to the optimal performance of the validation set. To avoid information leakage, the correlation coefficient and the feature importance are calculated on the training set, and the validation / test set is only used for threshold selection and effect evaluation; for strongly collinear features, the platform removes high redundant items based on the upper triangular traversal of the correlation matrix, and writes the remaining results to the unified field ID mapping table.

[0057] For point cloud data collected from Kinect or other three-dimensional sensors, the platform performs a local surface morphology feature extraction process. After each frame of point cloud input, the platform constructs a K-neighbor point set, and calls a principal direction estimation algorithm to fit a normal vector direction for each point neighborhood using the PCA method to generate a normal vector field.

[0058] To improve the consistency of the description, the platform performs a global direction reorientation process on all normal vectors to ensure that the normal vectors of the isomorphic surface area are oriented consistently. Subsequently, the platform constructs PPF descriptors based on the sampling points, calculates the Euclidean distance, normal angle, and pose difference between the point pairs, and constructs a vector set containing four-dimensional description features. Then the platform indexes all PPF descriptors through Hash mapping to generate a point pair feature voting model for subsequent identification tasks, ensuring that all point cloud data have vectorizable, indexable, and comparable structural expression capabilities.

[0059] For the multi-parameter redundancy analysis task in the process industry, the platform is based on the discretized field set described above to perform frequent item mining and association rule extraction. The platform loads the process variable set represented by the field discretization level, constructs the item set space, and performs frequent item expansion and support calculation on the space based on the Apriori algorithm. In each iteration, the platform eliminates item sets below the support threshold and confidence threshold, and records the appearance combinations of the remaining items. Finally, a set of association rules such as {CaO: high, SO3: medium}→{clinker activity index: excellent} is output, and the lift (Lift) index is calculated. The platform synchronously loads the expert confirmation rule list, cross-selects all generated rules and experience knowledge rules, and writes all remaining rules in the form of triples into the structured knowledge rule table.

[0060] wherein the support threshold and confidence threshold of the frequent item are selected and fixed during the training period according to the experience quantile of the sample transaction number and the item set distribution; in addition to meeting the support and confidence, the rule retention also requires the lift to be no less than the configured threshold, and to be consistent or complementary with the expert confirmation rule table; when multiple rules point to the same target field, the platform retains them in the order of Lift priority and then confidence, to prevent rule redundancy and conflict.

[0061] To measure the degree of environmental impact of the industrial process, the platform introduces the Life Cycle Assessment (LCA) method. When processing environmental load data, the platform divides the entire production process into three stages: raw material transportation, process execution, and waste discharge, according to the ISO14040 standard. In the data preparation stage, the platform calls the process equipment resource consumption record table and the pollutant emission flow table, and marks each type of resource and emission item as a mappable factor item. Then an environmental impact factor matrix is constructed, and all emission species are mapped to equivalents, equivalents, etc. standard environmental load units. The platform sums the environmental pressure indicators of each production stage by combining the process path and stage factor weights, and finally outputs the comprehensive carbon emission value and resource load indicator corresponding to the unit output. This indicator is written into the environmental performance evaluation module structure table to support green process selection.

[0062] In the point cloud object recognition process, the platform takes the pre-constructed PPF feature index table as the template model, loads the real-time point cloud data of the scene to be recognized, and performs point pair feature sliding matching process. Then, the window step is set to sample all point pairs in the scene point set, extract their PPF features, and call the Hough space voting mechanism to vote each pair of features into the six-dimensional pose parameter space.

[0063] After the voting is completed, the platform selects the pose parameter with the highest number of votes as the initial matching pose, and then calls the ICP iterative registration algorithm to optimize the initial pose with the minimum Euclidean distance error as the objective function, and finally determines the 6D pose of the target workpiece. The recognition result is stably output within a millimeter error range.

[0064] Among them, the construction of PPF descriptor adopts fixed step quantization in three dimensions of point pair distance, normal angle, and pose difference, and writes into Hash index. The quantization step and Hash bucket number are automatically selected and solidified as the configuration during the deployment period by statistical analysis of the discrete degree of the template point cloud and the scene point cloud. The Hough space voting takes the six-dimensional grid of the pose parameter as the carrier, and the grid density is also adaptively generated according to the sample discrete degree. The termination criteria of ICP include one of the two conditions triggered: the improvement amplitude of the mean square distance of the adjacent two iterations is lower than the threshold, and the maximum iteration round reaches the upper limit of the configuration; to improve the robustness, a robust rejection strategy based on the median absolute deviation is used for the outlier corresponding points during the registration process, and the final pose is confirmed by the voting peak neighborhood and the ICP optimal solution.

[0065] In the time series prediction task, the platform loads the set of process parameters determined through feature screening to construct a deep prediction model for predicting the target variable. The deep prediction model structure adopts a BiLSTM bidirectional long short-term memory network. In the training phase, the platform uses an equally spaced sliding window to divide the time series data, and each window segment contains information in both past and future directions, which are input into the LSTM units in both directions. During training, the MSE mean square error is used as the loss function to perform back propagation until convergence. After training, the model takes the current 6-dimensional process parameter sequence as input and continuously outputs rolling prediction values during the test period.

[0066] To avoid non-reproducible training process, the platform fixes data division and preprocessing before entering the training period: normalizes the target process parameter sequence and divides it into training / verification / testing parts in chronological order; the sliding window length is set to the empirical value of the first decay inflection point of the autocorrelation of the target variable and is solidified as the configuration, and the step size and batch size are recorded during the training period; hyperparameters are obtained through grid search on the validation set and written into the training log; training adopts early stopping strategy and gradient clipping to prevent overfitting and gradient explosion; finally, the MSE and the robust statistics of the rolling prediction error are evaluated on the test set, and the optimal weight and inference configuration are persisted and registered with the model.

[0067] In the data access phase, a user behavior driven cache recommendation mechanism is introduced, and a ternary influence factor model is constructed to perform data type priority sorting and cache space proportion dynamic allocation. Specifically, first, based on the user access log recorded in the log structured index pool, the behavior trajectory of all user accounts is parsed one by one. The platform performs behavior aggregation operation on the operation behavior of each user in the historical access period, extracts the access frequency count of each user in all data type dimensions, and calculates the access frequency proportion index by the proportion of access times in the total access behavior number of the user, to form the first influence factor W1(u, i).

[0068] Next, the data type identifier i recorded in the access behavior is extracted to extract its system defined second influence factor security level and third influence factor user identity permission level. The two indicators are called from the data type security level table and the user permission registration table respectively, and written into the temporary factor buffer table according to the data type and user number.

[0069] Subsequently, the platform constructs a ternary influence factor model with access frequency proportion W1(u, i), data type security level and user permission level as factor sources, and performs weighted calculation on the active user subset in the current user set.

[0070] Among them, the basis for determining the active user subset is that the number of accesses of the user in the current access period exceeds the activity threshold Lact. The platform calculates the recommended value R(i, k) of each data type i under the permission level k according to the following formula: R(i, k)=Σ{u}[Iu×W1(u, i)×δ(Pu≥Pk)]; Where Iu is the activity score of user u, which is generated by weighting the user access frequency and continuous access days; W1(u, i) is the access frequency proportion of user u to data type i; δ(Pu≥Pk) is a Boolean judgment function, which takes value 1 when the user permission level Pu is greater than or equal to the current evaluation level Pk, otherwise 0.

[0071] It should be noted that, in order to ensure the reproducibility of R(i, k), Iu is linearly weighted by two parts of normalized quantity: one is the normalized value of the user's access frequency in the current access period, and the other is the normalized value of the number of consecutive access days; the weight of the two is optimized to a fixed configuration through grid search of historical access logs in the training period. The determination of the activity threshold Lact is based on the empirical quantile calculation of the historical access distribution, and is solidified in the deployment period; δ(Pu >= Pk) is obtained by matching the user permission registration table and the data type security level table. The cache space proportion allocation of Stotal is normalized by the recommendation value matrix as the weight, and the backfill of main library loading and cache copy writing is triggered in the failure path to support the subsequent hit strategy in a closed loop.

[0072] Further, after calculating the complete recommendation value matrix for all data type and permission level combinations, sorting operation is performed on each data type under each permission level according to the recommendation value, and a cache scheduling reference list is generated. The platform allocates the current allocatable cache space Stotal according to the recommendation value as the weight, and generates a cache space allocation strategy; When the user actually initiates a data download request, the user identity recognition module is first called to obtain the user number and role permission level Pu of the current user, and then the data type i in the request message is parsed, and the data access permission controller is called to perform permission verification operation. The permission verification logic is based on the matching result of data type access level and user permission level to make a Boolean judgment, if the verification fails, the data response process is directly blocked and a permission insufficient prompt is returned; if the verification passes, the platform enters the cache access judgment process.

[0073] In the cache access phase, the platform calls the data type and cache space index table to find out whether the data type i has been preloaded into the cache under the current permission level. If the cache hits, the data is directly read from the cache area and returned to the user terminal; if the cache misses, the platform extracts the target data from the main database, and after completing the data transmission, the cache copy writing operation is synchronized to ensure that this type of data can be pre-fetched in subsequent access.

[0074] The application also proposes an industrial internet big data platform, as shown in Figure 2 The figure shows that it includes: multi-source data access and structure binding module, data preprocessing and field attribution analysis module, field attribution correction and modeling field generation module, model construction and cache scheduling management module, signal connection between modules.

[0075] The multi-source data access and structure binding module accesses four types of data: process parameter, point cloud, smelting message and user log, completes protocol configuration, physical channel binding and structure context ternary information writing; The data preprocessing and field attribution analysis module performs missing value completion, topology reconstruction, behavior chain analysis and Γ tensor construction on the time series, point cloud and log data respectively, and establishes the association between the field structure path and the role label; The field attribution correction and modeling field generation module analyzes the field turn-back, label drift and path break based on the minimum attribution period, completes the Ltype classification, field attribution chain repair and pseudo-merged field elimination, and generates the modeling field set; The model construction and cache scheduling management module discretizes continuous variables, models point cloud objects and predicts time series, and simultaneously performs cache scheduling and permission verification according to access frequency, security level and permission level.

[0076] The above formulas are dimensionless numerical calculations, and the formulas are obtained by software simulation of a large amount of collected data to obtain a formula of the most recent real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0077] The above embodiments can be realized wholly or partially by software, hardware, firmware or any combination thereof. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product.

[0078] Those skilled in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application of the technical solution and the constraints of the invention. A person skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0079] In addition, each functional module in each embodiment of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0080] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0081] Finally, the above is only a preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. An industrial internet data processing method, Its characteristics include: Collect process parameters of the industrial process, 3D visual point cloud, distributed message smelting data and user operation logs, and bind the collection timestamp, device identifier and channel number; Introducing the field's first response channel number, the field's assigned role label, and the field's write time offset, we construct the role state transition trajectory and the field response chain, and construct the Γ(t,x,r) tensor within a fixed sliding time window; The minimum common structural behavior period λn is calculated using the Γ(t,x,r) tensor as the core state index. The structural affiliation consistency analysis is performed to mark the permission state disconnect field. Each field belonging to the permission state disconnect field is classified into Drift, Overlap, Lag or Pseudo type. When the field behavior window covers λn and the segment has completed the initial classification and trajectory convergence, perform differentiated path repair, reconstruct and output the complete set of field classification paths.

2. The industrial internet data processing method according to claim 1, characterized in that: Collect process parameters, 3D visual point clouds, distributed message smelting data, and user operation logs from the industrial process, and bind the collection timestamp, device identifier, and channel number.

3. The industrial internet data processing method according to claim 1, characterized in that: The system introduces the field's first response channel number, the field's assigned role label, and the field's write time offset. Then, it uses the field's assigned role label as the trajectory node and combines the field's write time to perform a switch event query on the permission status record table. The system also establishes the field node order based on the field structure path number and the field response order, thus constructing the role status transition trajectory and field response chain. Within a fixed sliding time window, the above role state transition trajectory and field response chain are synchronously aligned with the user operation session as the boundary. For the field set in each time period, the Γ(t,x,r) tensor is constructed by time t, structural path number x and role label r through a three-dimensional sparse matrix construction algorithm combined with hash mapping index.

4. The industrial internet data processing method according to claim 3, characterized in that; By using the Γ(t,x,r) tensor as the core state index, combined with the role label backtracking set, the field disintegration chain backtracking and regrouping path graph, and the window drift recovery table, we perform role label consistency judgment, path attribution continuity verification, and window focus state verification to construct the field disintegration chain backtracking and regrouping path graph and the window drift state segment. Based on three time parameters—the minimum time Tbreak for the initial label backlink chain interruption to closed-loop reconstruction, the time Treturn(x) required for field x to be misclassified and backlinked to the main path, and the shortest regression period Trebind(w) for window w to rebind the role context and trigger the correct field attribution binding—autocorrelation analysis of the field structure path sequence is used, and the minimum common structure behavior period λn is calculated by extracting the main periodic components using fast Fourier transform.

5. The industrial internet data processing method according to claim 4, characterized in that: Using λn as the sliding time window length, slice the Γ(t,x,r) tensor by the time dimension, and perform structural ownership consistency analysis on the tensor index set within each sliding time window, marking the disconnected fields of permission status.

6. The industrial internet data processing method according to claim 5, characterized in that: Using λn as the sliding time window length, traverse the Γ(t,x,r) tensor, call the role state backtracking function to generate the main role path sequence according to the r-dimensional trajectory, and match the field structure path belonging to the permission state disconnected field with the main path structure set to identify path jump, inconsistent labels, out-of-order response, role not switched and causal path missing behavior patterns. Based on the matching results, each field belonging to the permission status disconnect field is divided into the main path drift segment (Drift), the role overlap segment (Overlap), the permission uncut segment (Lag), or the pseudo-closed window segment (Pseudo).

7. The industrial internet data processing method according to claim 6, characterized in that: After the field disturbance segment is divided, when the behavior window of the field completely covers the minimum common structure behavior period λn, and the segment to which the field belongs completes the initial judgment of structure ownership and path trajectory convergence within the period, the field segment type identifier Ltype is inserted, and differentiated path repair is performed according to the Ltype field. Next, using the minimum common structural behavior period λn as the time index window, the structural state sequence of all classified fields is extracted from the tensor Γ(t,x,r). Based on the set of field return paths, the field role label re-aggregation sequence and the log channel write trajectory are combined to perform the final reconstruction of the field affiliation label path, forming a complete set of field affiliation paths.

8. The industrial internet data processing method according to claim 1, characterized in that; The complete set of field attribution paths is discretized and jointly evaluated to generate modeling fields and extract point cloud features to identify the pose of target objects. Discrete fields are mined to generate a field priority list and a structured knowledge rule table. Then, according to the process stage, resources and emission factors are marked to calculate the environmental load output carbon emission value; a time series prediction model is constructed to predict the target variable; and a cache scheduling strategy is generated based on access frequency, security level and permission level to complete data access and update.

9. An industrial internet big data platform, used to implement the industrial internet data processing method according to any one of claims 1-8, characterized in that, include: The module includes a multi-source data access and structure binding module, a data preprocessing and field attribution resolution module, a field attribution correction and modeling field generation module, a model construction and cache scheduling management module, and signal connections between the modules. The multi-source data access and structure binding module accesses four types of data: process parameters, point cloud, smelting messages and user logs, and completes protocol configuration, physical channel binding and writing of structural context ternary information. The data preprocessing and field attribution parsing module performs missing completion, topology reconstruction, behavior chain parsing and Γ tensor construction on time series, point cloud and log data respectively, and establishes the association between field structure path and role label; The field attribution correction and modeling field generation module analyzes field reversal, label drift and path breakage based on minimum attribution cycle, completes Ltype classification, field attribution chain repair and pseudo-merged field removal, and generates a modeling field set. The model building and cache scheduling management module performs continuous variable discretization modeling, point cloud object recognition and time series prediction modeling, and performs cache scheduling and permission verification based on access frequency, security level and permission level.

Citation Information

Patent Citations

  • Intelligent visual management method and system for enterprise big data

    CN120144416A

  • Method and platform for integrating and sharing health data of old people

    CN120321057A

  • Human resource data security management method and system

    CN120337287A

  • Intelligent scheduling and collaboration system and method for city reconstruction full life cycle

    CN120672163A

  • Data acquisition visualization processing system based on big data analysis

    CN120687807A