Industrial data storage method and system based on knowledge graph
By using a knowledge graph-based approach, cooling tower operation data is acquired and identified and analyzed for connectivity. This solves the problem of instability in establishing node relationships in multi-source data association, achieves accurate data classification and storage continuity, and improves the stability and logical integrity of data organization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies lack a joint analysis mechanism for record time sequence and behavioral labels in multi-source data association. The establishment of node relationships relies on static graph structures, which are difficult to reflect the real transmission paths between data, resulting in fuzzy direction identification and disordered node links. In particular, when processing data with frequent field changes or irregular recording cycles, it is impossible to accurately classify path channels, affecting the continuity of data expression and the integrity of structural logic.
By using a knowledge graph-based approach, cooling tower operation data is acquired, data is aggregated by time, identifiable content is added and input into the graph node recording area, the direction of information transmission is analyzed, adjacent nodes are connected, node attribute tags are extracted, the direction of data transmission is identified, the direction of connection relationships is updated, path channels with consistent order are divided, and node content and connecting lines are filled into the processing area and labeled for structured storage.
It enhances the organizational connections, process continuity, and structural logic among data, improves the path stability, node orderliness, and data coordination in graph construction, and ensures accurate data classification and continuous storage.
Smart Images

Figure CN121636501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial data storage technology, and in particular to an industrial data storage method and system based on knowledge graphs. Background Technology
[0002] The field of industrial data storage technology involves methods for the unified collection, encoding, transformation, organization, indexing, archiving, and management of structured and unstructured data generated by production equipment, operating systems, and monitoring terminals in industrial scenarios. Core aspects of this field include standardized modeling of multi-source heterogeneous industrial data, construction of semantic relationships between time-series data and equipment entities, optimization of data storage structures for querying and analysis, and traceable management of historical data. Methodologically, it typically relies on predefined data collection specifications to establish field mapping rules, achieves data format unification through preprocessing, and utilizes database systems to construct data storage structures to support high-frequency writing and conditional retrieval. Traditional knowledge graph-based industrial data storage methods address the complex relationships between equipment entities and their operational data in industrial systems by constructing graph structures for semantic representation and storage. This patent primarily addresses the semantic relationship expression problem between operating condition data, operating parameters, and maintenance events generated during the operation of cooling tower equipment.
[0003] Existing technologies lack a joint analysis mechanism for record time sequence and behavioral labels in multi-source data association. The establishment of node relationships relies on static graph structures, which are difficult to reflect the real transmission paths between data. In actual working conditions, problems such as fuzzy direction identification and disordered node links are prone to occur. In particular, when processing data with frequent field changes or irregular recording cycles, it is impossible to accurately classify path channels, resulting in structural organization imbalance and semantic matching interruption. For example, during the sudden change of the operating state of a cooling tower system, traditional methods are difficult to quickly identify abnormal nodes and reorganize the structure, affecting the continuity of data expression and the integrity of structural logic. Summary of the Invention
[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide an industrial data storage method based on knowledge graphs; To achieve the above objectives, the present invention adopts the following technical solution: an industrial data storage method based on knowledge graphs, comprising the following steps: S1: Obtain cooling tower operation data, collect data by time, locate the data source according to the number field, add corresponding identification content to each record, connect the identification with the value in the corresponding field, input it into the graph node record area, add the original data field, and obtain the semantic mapping content. S2: Based on the semantic mapping content, select consecutive record nodes, call the time tag of each group of data, analyze the information transmission direction and connect adjacent nodes, input the connection result into the graph path, and obtain the process connection relationship content; S3: Based on the process connection relationship content, extract node attribute tags, read the behavior type of the current node, identify the data transmission direction, update the connection relationship direction and write it into the graph to obtain the content corresponding to the structural direction; S4: Based on the content corresponding to the structural direction, retrieve the path links to adjacent nodes, extract the period and field change order, combine paths with the same order into the same channel, divide paths with different orders, and obtain path ownership information. S5: Based on the path attribution information, extract the nodes and connection relationships, fill the node content and connection lines into the processing area, attach labels and write the node relationships to obtain the structured storage result.
[0005] As a further aspect of the present invention, the semantic mapping content includes cooling water temperature identifier, fan speed identifier, inlet and outlet water pressure identifier, acquisition time label, number field mapping, and raw data field; the process connection relationship content includes time label sequence, node connection order, information transmission path, and process node pair; the structural direction corresponding content includes behavior type label, sequence information, action direction label, and connection direction data; the path attribution information includes node record period, field change order, path grouping label, and processing channel identifier; and the structured storage result includes path area division, node connection lines, label additional information, and node association data.
[0006] As a further aspect of the present invention, the corresponding identification content refers to the unique number information added to each data record during the process of collecting cooling tower operation data, which identifies the data source and the device to which it belongs. The raw data fields refer to the raw monitoring data fields of cooling water temperature, fan speed, and inlet and outlet water pressure obtained during the collection process, along with the units and sampling information.
[0007] As a further embodiment of the present invention, the graph path refers to the directed path generated between data nodes according to the connection logic of time and behavior sequence, reflecting the data transmission, behavior causality and process relationship, and is recorded in the graph structure; The processing area refers to the target data structure area used to receive and store the node and connection relationships. The data is sorted, tagged, and written into the structured storage result.
[0008] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Acquire cooling water temperature, fan speed, inlet water pressure and outlet water pressure data during the operation of the cooling tower, compare the data frames according to the acquisition time, arrange data with the same time in the same sequence, extract the field set in the sequence and locate the field position, associate the data within the time series, and obtain the data mapping table within the acquisition period. S102: Based on the data mapping table within the collection period, call the number field information, identify the data source, arrange the field values in the order of the numbers, append the number values to the corresponding field positions, and output the data content with the source identifier after splicing to obtain the data source identifier mapping result; S103: Based on the data source identifier mapping result, import the data content into the graph node area, write the correspondence between the number value and the field value into the node structure, attach the original field information and timestamp content, expand the graph node representation range, and obtain semantic mapping content.
[0009] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Based on the semantic mapping content, select the data nodes with continuous time in cooling water temperature, fan speed and inlet and outlet water pressure, call the timestamp field of each data, arrange the nodes in the data sequence according to time, locate the node pair number under continuous time, and obtain the node time sequence index pair. S202: Based on the node time series index pair, extract the timestamp values of adjacent nodes, compare the order of time fields, identify the order of time arrangement, connect the nodes according to the time arrangement order to form a data transmission path, and output the time sequence information corresponding to the path as a sequence to obtain the node transmission direction sequence. S203: Based on the node transmission direction sequence, link adjacent nodes sequentially according to the transmission direction, write each pair of nodes into the path order field, and attach the order value and node number together to the graph path area. According to the time progression relationship number path content position, obtain the process connection relationship content.
[0010] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Based on the process connection relationship content, extract the field labels in each pair of connected nodes, identify the node behavior type, extract the behavior type information according to the order of the fields, and classify the labels and behavior types into the corresponding relationship to obtain the node behavior attribute set. S302: Based on the set of node behavior attributes, select the dual-node behavior type in the connection path, call the node time field, determine the order of behavior occurrence, arrange the recognition results in time order as the direction of action, and obtain the behavior transmission sequence. S303: Based on the behavior transmission sequence, update the direction field of the corresponding connection line in the graph structure, replace the original direction data, fill the new direction sequence into the connection field position according to the time progression order, output the direction update result, and obtain the content corresponding to the structure direction.
[0011] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Based on the content corresponding to the structural direction, retrieve adjacent data nodes in the path, extract the collection period and field change sequence of each group of nodes, use the field sequence as the identification basis, extract the sequence features between node pairs, and obtain the node field sequence information set. S402: Based on the node field order information set, group together data paths with consistent field order, split data paths with changing field order, and separate data paths of differentiated field sequences into independent channels based on the field position order to obtain a field order classification information set. S403: Based on the field order classification information set, the paths corresponding to the numbers are placed in independent sequence regions, and the path information is collected within the regions. The data sequence belonging relationship is output according to the number partitioning result to obtain the path belonging information.
[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Based on the path attribution information, extract the data nodes and corresponding connection data fragments within the path, call the path identifier field, extract the node content and connection information within the same path respectively, and arrange them into the path channel area according to the order of data appearance to obtain the path channel permutation set; S502: Based on the path channel arrangement set, compare the node content tags, retrieve the existing attribute tag content of each node, and concatenate the tag information with the current node's fields. At the same time, juxtapose the associated connection data content to obtain the node tag concatenation sequence. S503: Based on the spliced data content in the node label splicing sequence, write the node field and the connection field together into the target data storage area, and use the path channel as the dividing basis to divide the area to obtain the structured storage result.
[0013] A knowledge graph-based industrial data storage system includes: The data acquisition module collects cooling tower operation data by time, locates the data source according to the number field, adds corresponding identification content to each record, links the identification with the value in the corresponding field, and inputs it into the graph node record area, adds the original data field, and obtains semantic mapping content. Based on the semantic mapping content, the node mapping module selects continuously recorded data nodes, calls the time tag of each group of data, analyzes the information transmission direction and connects adjacent nodes, inputs the connection results into the graph path, and obtains the process connection relationship content. The relationship identification module extracts node attribute tags based on the process connection relationship content, reads the behavior type of the current node, identifies the data transmission direction, updates the connection relationship direction and writes it into the graph to obtain the content corresponding to the structural direction. The path attribution module retrieves adjacent nodes in the path based on the content corresponding to the structural direction, extracts the period and field change order, combines paths with the same order into the same channel, and divides paths with different orders to obtain path attribution information. Based on the path attribution information, the structure storage module extracts the nodes and connection relationships, fills the node content and connection lines into the processing area, adds labels and writes the node relationships to obtain the structured storage result.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, multi-field synchronous collection is achieved by binding running data with identified content. Information flow paths are constructed according to time sequence. The transmission direction between nodes is marked by attribute tags and behavior types. Channel division and classification are carried out by selecting paths with consistent order based on field change patterns. The path and node content are uniformly arranged and filled into the structural area with added tag information to enhance the organizational association, process continuity and structural logic between data, and improve the path stability, node orderliness and data coordination in the graph construction. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7This is a system module diagram of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0022] Please see Figure 1 This invention provides an industrial data storage method based on knowledge graphs, comprising the following steps: S1: Acquire cooling water temperature, fan speed, inlet water pressure and outlet water pressure data generated during the operation of the cooling tower. Using the acquisition time as a reference, arrange data with the same time in the same processing sequence, call the number field to locate the data source, add corresponding identification content to each record, connect the identification and value in the corresponding field, and input it into the graph node record area. At the same time, add the original data field to obtain the semantic mapping content. S2: Based on semantic mapping content, select data nodes that record consecutively, determine the order of records of adjacent nodes according to time sequence, call the time tag of each group of data, analyze the direction of information transmission, link the nodes before and after time in first-out order, put the linking results into the graph path record, and obtain the process connection relationship content. S3: Based on the process connection relationship content, extract the attribute tags of the nodes at both ends of the connection, read the behavior type of the current node from the corresponding node, identify the direction of action according to the order of appearance of the behavior type in the data, replace the direction relationship, write the current direction information into the connection data position, and obtain the content corresponding to the structural direction. S4: Based on the content corresponding to the structural direction, retrieve adjacent nodes in the data nodes of the path link, extract the node record period and field change order, the data arrangement method between segmented paths, put paths with consistent field change order into the same processing channel, separate paths with inconsistent order, and obtain path ownership information. S5: Based on path attribution information, extract the nodes and corresponding connection content in the path, set up an independent area for each group of paths in the data processing structure, fill the node content and connection lines into the corresponding positions, attach existing label information and write the node association data during the data entry process, and obtain the structured storage result.
[0023] The semantic mapping content includes cooling water temperature identifier, fan speed identifier, inlet and outlet water pressure identifier, acquisition time label, number field mapping, and raw data field. The process connection relationship content includes time label sequence, node connection order, information transmission path, and process node pair. The structure direction correspondence content includes behavior type label, sequence information, action direction label, and connection direction data. The path attribution information includes node record period, field change order, path group label, and processing channel identifier. The structured storage result includes path area division, node connection lines, label additional information, and node association data.
[0024] Please see Figure 2 The specific steps of S1 are as follows: S101: Acquire cooling water temperature, fan speed, inlet water pressure and outlet water pressure data during the operation of the cooling tower, compare the data frames according to the acquisition time, arrange data with the same time in the same sequence, extract the field set in the sequence and locate the field position, associate the data within the time series, and obtain the data mapping table within the acquisition period. When acquiring data on cooling water temperature, fan speed, inlet water pressure, and outlet water pressure during the operation of a cooling tower, periodic data collection is required for each of these four parameters. Cooling water temperature is converted into a digital signal via a transmitter connected to a resistance temperature detector (RTD) sensor. Fan speed is detected by a speed sensor on the impeller shaft, which detects speed pulses and converts them into numerical values. Inlet and outlet water pressures are collected by pressure monitors installed on the inlet and outlet pipes. The data collection interval for each parameter is set to a fixed interval. After completion, a data frame containing all four parameters is generated. The time fields in the data frames are then compared to determine if the time labels of each data frame are consistent. If they are consistent, they are placed in the same sequence. The parameter fields in the current frame are extracted and their positions are marked (e.g., temperature is ranked 1st, fan speed 2nd, water pressure 3rd and 4th). Then, in the same processing sequence, the fields are processed sequentially according to their order of arrangement. The process involves confirming the location, extracting the original data values corresponding to the fields, and binding them to the same acquisition period. All field contents within the same period are sequentially aggregated into a field set. If there are consecutive acquisition contents, the field sets of each period are connected sequentially in time order to form a sequence structure. If the time tag interval of adjacent frames is inconsistent with the acquisition period, the abnormal frames are skipped without processing. Subsequently, the field data within each period is bound according to the time series index. For example, wind speed value of 1532 rpm and inlet water pressure of 0.32 MPa are bound to specific time nodes. Then, all bound contents are processed and the field names and acquisition values are filled into the structure rows and columns. The rows represent the acquisition period number, and the columns represent the field item names. The total number of sequence rows accumulated during the acquisition process corresponds to the number of acquisitions. The number of field items is equal to the number of channels participating in the acquisition, resulting in a data mapping table within the acquisition period.
[0025] S102: Based on the data mapping table within the collection period, call the number field information, identify the data source, arrange the field values in order of number, append the number value to the corresponding field position, and output the data content with source identifier after concatenation to obtain the data source identifier mapping result; First, locate the corresponding ID field for each row of data. This field serves as a unique identifier for the data source. The ID field is typically preset to the ID parameter of the data acquisition device. The ID value is in integer or unique character format. Then, retrieve the content of the ID field and perform corresponding operations against the field values in the data row containing that ID. Sort the fields under the current ID according to the order of the ID values. Extract the four values (cooling water temperature, fan speed, inlet pressure, and outlet pressure) from each record according to the field definition order, and append this ID value to the end of each field. For example, if the cooling water temperature field value for ID 0003 is 32.1, then... It should be "32.1-0003". Append operations are performed on all four fields sequentially to create a record with a number. Further, based on the number, field values from devices with the same source are grouped into the same content set. In the data output structure, each content set is expanded sequentially according to its number. Field values corresponding to different numbers form different data blocks. The field order in the structure remains consistent with the original field order in the mapping table. This operation establishes the field value attribution relationship for each data acquisition device in the cooling tower and clarifies the source of each record value in the same data channel. After all numbered fields are concatenated, they are merged into the processing channel, ultimately yielding the data source identifier mapping result.
[0026] S103: Based on the data source identifier mapping result, import the data content into the graph node area, write the correspondence between the number value and the field value into the node structure, attach the original field information and timestamp content, expand the graph node representation range, and obtain the semantic mapping content; First, the correspondence between the ID value and the field value in each data entry is extracted. A one-to-one mapping operation is performed between the field content and its ID. For the four fields of cooling water temperature, fan speed, inlet water pressure, and outlet water pressure, the values are bound according to the ID field. Each ID corresponds to a complete set of field data. Then, the data source represented by each ID is identified in the mapping results. The ID-field mapping relationship is input into the graph node area. The node area is arranged with field type as the basic classification item, ID as the row index, and field item as the column index. For example, when the ID is 001, its corresponding temperature is 29.4, fan speed is 1510, and inlet water pressure is... If the water pressure is 0.28 and the output pressure is 0.21, the data is entered into the node area in the format "001, 29.4, 1510, 0.28, 0.21". Then, original field information is appended to the field values of each record in the node area. This original field information includes the sampling unit and sensor parameter description; for example, the unit for the temperature field is "℃", the unit for the wind speed field is "rpm", and the unit for the pressure field is "MPa". A field identifier string is formed by concatenating the field name and unit. If the temperature field is 29.4, its identifier string is "Temp-℃-29.4". All field identifier strings are arranged in field position order under the same number. Next, the timestamp information corresponding to the data record is read. The timestamp is represented in year-month-day-hour-minute-second format to identify the time of data collection. The index value within the time period is used to replace the specific time; for example, the time index for number 001 is T1, and for number 002 it is T2. The time index is then appended as an extended attribute of the node to the field mapping entry, realizing the association of time series at the node level. Within the node area, to avoid duplicate inputs under the same ID, a field overwrite judgment is performed. This judgment is based on the uniqueness of the ID and time index combination key. When the same combination key exists, only the latest input value is retained, overwriting the previous data with the same combination key, thus ensuring that there is no duplicate stacking of field content within the node structure. Subsequently, the input node content undergoes semantic expansion. Each node data is mapped to a semantic label table according to its field type. Each field in the label table has a corresponding semantic definition; for example, the temperature field is defined as "thermal feature," wind speed as "dynamic feature," and pressure as "flow feature." The system determines the semantic category through field matching and appends the semantic definition to the node field description based on the matching result. In this process, each data entry forms a complete semantic entry through ID, field name, unit, time index, and semantic label. These semantic entries are stored in the graph node area and bound to the corresponding ID, completing the expansion of the node's semantic information. In the final output, each numbered node contains four types of fields, four types of semantic identifiers, and a set of time indexes, collectively forming a multi-dimensional semantic expression structure for the node, resulting in semantically mapped content.
[0027] Please see Figure 3 The specific steps of S2 are as follows: S201: Based on semantic mapping content, select data nodes with continuous time in cooling water temperature, fan speed, and inlet and outlet water pressure, call the timestamp field of each data, arrange the nodes in the data sequence according to time, locate the node pair number under continuous time, and obtain the node time sequence index pair. First, the field names, values, and corresponding timestamps in each node's data are analyzed. The node number and timestamp pair are extracted, and the validity of the field value is determined. It is also verified that the value corresponds to the normal operating range of the cooling tower, such as a temperature range of 15 to 40 degrees Celsius, a fan speed range of 1000 to 2000 RPM, and a pressure range of 0.1 to 0.5 rpm. Data exceeding these ranges is discarded to avoid misleading the identification of consecutive nodes. Next, the timestamp field of each valid record is retrieved, and the time sequence number in the time label is read and sorted numerically. The time series structure is constructed using the ordered data index. For example, if the time labels are T102, T103, T104, T106, and T108, a sorted vector of [102, 103, 104, 106, 108] can be constructed, with the time breakpoint between T104 and T106. For the sorted node vector, the difference between adjacent time tags is compared item by item. A node interval baseline of 1 is set. When determining whether adjacent time tags form a continuous sequence, the difference is used to determine continuity. If the difference is 1, it is considered a continuous node pair. In the example above, [T102, T103] and [T103, T104] are continuous node pairs, and their corresponding node numbers, matched by time index, are N1, N2, N3, etc., thus obtaining continuous node number pairs [N1, N2] and [N2, N3]. For nodes with interrupted positions (such as T105 missing between T104 and T106), they will not be included in the continuous sequence during the identification process to ensure the accuracy of continuous node pair identification. The node numbering process relies on the combination relationship between the number field and field content in the semantic mapping. The original number is extracted from the field content, and a mapping structure between the number and the time index is established. Through this mapping structure, the number of each record after time sorting is extracted, thereby mapping the time order to the node order. Based on this, node sequence pairs are generated. Each node sequence pair represents a set of temporally consecutive node combinations, forming a basic node chain structure, such as [N1, N2], [N2, N3], [N3, N4], and so on. To avoid misjudgments due to duplicate or incorrect numbers, a uniqueness check is performed on the number values before generating the node number pairs. The combination of node number and timestamp is used as the key for verification. If the key is duplicated, the later occurrence in the time series is retained to avoid conflicts in the data structure and ensure the uniqueness and validity of consecutive node pairs. Finally, all temporally consecutive number combinations are extracted in the form of consecutive node pairs to obtain the node time series index pairs.
[0028] S202: Based on the node time series index pairs, extract the timestamp values of adjacent nodes, compare the order of time fields, identify the order of time arrangement, connect the nodes according to the time arrangement order to form a data transmission path, and output the time sequence information corresponding to the path as a sequence to obtain the node transmission direction sequence. Extract the timestamp values contained in each pair of nodes, call the "timestamp" field of the first and second nodes, compare their time values, and determine the temporal relationship of the data through numerical difference calculation. The timestamp field uses integer format to represent the time progression order. For example, if the timestamp of node A is 11521 and the timestamp of node B is 11522, then the difference Δt is 1, indicating that node A appears before node B. Based on the difference result, it is confirmed that node A is the forward node of the current path and node B is the backward node. During the monitoring of the cooling tower operation status, if node A is the fan speed and node B is the outlet water pressure, it indicates that the speed change precedes the pressure change, and the transmission relationship from speed to pressure is established. Subsequently, the timestamp comparison results for each group are aggregated. Using the node number as the key, the identified order is associated and labeled, forming a linked list structure. For example, node number pairs [N3, N4], [N4, N5], and [N5, N7] form a data path chain [N3→N4→N5→N7] according to the comparison order. Node pairs on each path are connected sequentially according to the time progression. In the above linked list structure, each node originates from the constructed semantic mapping node pool, and there is a mapping relationship between the number value and the semantic field. The semantic content of the field can be obtained by tracing the source through the number. Combining the semantic fields, for example, if node N3 represents "cooling water temperature" and N4 represents "inlet water pressure", then the path segment [N3→N4] can represent "temperature change leads to inlet water pressure change". After completing the node connections, a corresponding time sequence is constructed based on the arrangement order of the node numbers in the linked list path. For example, the timestamp sequence corresponding to the numbered path [N1→N2→N5→N8] is [11520, 11521, 11524, 11529]. This timestamp sequence serves as the time basis reference for the transmission path, and a node time transmission vector is derived. This vector is constructed using a numbered index method for subsequent use in the graph. For example, the vector [1, 2, 5, 8] corresponds to the index number order and is represented as path flow line information according to the standard graph structure definition. During the construction of the node flow structure, to eliminate path anomalies caused by timestamp reading errors or duplicate writing, a uniqueness check is performed on the extracted timestamp field. If a node pair is found to have overlapping timestamps, the path is directly cut off on that node pair to prevent cyclic paths from affecting the overall link. This structural correction method ensures that the path only advances backward and does not backtrack. Finally, the time index paths are concatenated according to the node order to form a set of vectorized sequences describing the time relationships between nodes, resulting in the node transmission direction sequence.
[0029] S203: Based on the node transmission direction sequence, link adjacent nodes sequentially according to the transmission direction, write each pair of nodes into the path order field, and attach the order value and node number together to the graph path area. According to the time progression relationship number path content position, the process connection relationship content is obtained. First, extract the node pair information. Perform bidirectional matching on each node pair to confirm its position index in the path graph. Use the node number to find its corresponding structure in the graph node area. Compare the timestamp field of the structure with the data generation time of adjacent nodes, using ascending order to determine the relationship. Construct a connecting path based on chronological order. Merge the start and end numbers of this path with its sequence value and write it into the graph path area. Simultaneously, set the sequence value to a positive integer and increment it to form the path progression sequence. This applies to node 3 (cooling water temperature) and node 3 (fan speed). Node 5, at timestamps t3=18560 and t5=18565, records the path relationship as 3-5, with a path order value of 1. Multiple path order fields are consecutively numbered. For example, if the timestamps of nodes 5 and 9 are 18565 and 18572, the path order is numbered as 2. This recursively forms a path progression chain. The content format is represented by a structured key-value pair. Each path field includes a start number, an end number, and an order value. Multiple fields are stored sequentially in the path structure according to their numbering order, and the entire connection order in the path is output to obtain the process connection relationship content.
[0030] Please see Figure 4 The specific steps of S3 are as follows: S301: Based on the process connection relationship content, extract the field labels in each pair of connected nodes, identify the node behavior type, extract the behavior type information according to the order of the fields, and classify the labels and behavior types into the corresponding relationship to obtain the node behavior attribute set; First, obtain the field structure information between all paired nodes in the graph and extract the field label field. By reading the field identifier field in the node field, identify the data item describing the operation action and determine whether it contains operation behavior keywords, such as verbs like "start," "adjust," and "transport." If such information exists, the field is determined to be an operation type field. Based on this, according to the order of the fields in the original data table, match the two fields of each connected node sequentially and retrieve the expression content of their respective behavior types. Extract the data item in the field corresponding to the node number to form the response mapping of the operation behavior between nodes. Considering the difference in the number of fields in different nodes, call the field order corresponding to the field in the node structure. Information is extracted using a priority order as the reference order for behavior type extraction, ensuring that the behavior information extracted within the same connection has consistent field sources. Further, the label field content and behavior type extraction results are obtained. Metadata items associated with the field label are searched in the field structure of each node, and association pairs between the label and its corresponding behavior type are established. For example, if node A's field is "Wind turbine status: Running", its field label is "Wind turbine status", and its behavior type is "Running", then its corresponding relationship is "Wind turbine status - Running". These structure pairs are stored between two field information tables, and the same operation is performed on the remaining fields in the node structure. Finally, the structural content is extracted from the mapping relationship between field labels and behavior types of all connected nodes to obtain the node behavior attribute set.
[0031] S302: Based on the set of node behavior attributes, select the dual-node behavior type in the connection path, call the node time field, determine the order of behavior occurrence, arrange the recognition results in time order as the direction of action, and obtain the behavior transmission sequence. First, extract the content of the behavior types corresponding to each pair of adjacent nodes in the graph connection path. The behavior types are usually given by the verbal expressions in the node fields, such as structures like "start the fan", "increase the water pressure", "stop the circulation", etc. During the process of extracting the node behavior type content, read and identify the corresponding behavior description fields for the field tags in the double nodes. By analyzing the behavior roots in the fields, ensure that the selected fields are unique and used to describe the node behavior. Further, read the timestamp field included in this behavior field as the timing identifier for the occurrence of the behavior. Compare the time field values of each group of double nodes. The judgment is based on the magnitude relationship of the absolute time values. If the time value corresponding to the behavior field of node A is earlier than that of node B, mark that the behavior of A takes precedence over B. During the process of processing the field values, convert the timestamp values into a unified floating-point representation, and use a format correction method to eliminate the original format differences, ensuring that the time judgment results have continuity and an increasing relationship. For example, if the behavior field of node A "fan start" occurs at time t1, and the field of node B "cooling water pressure increase" occurs at time t2, if t1 < t2, then define that the fan start action precedes the water pressure increase. Based on the judgment results, construct the transfer direction between the behaviors, forming the causal relationship of operations between the nodes. After all the double-node behavior pairs in the path have completed the above judgment operations, arrange their order according to the judgment results in sequence. Finally, establish a complete timing connection path between all the behavior action pairs, and output it in text as a record of the transfer structure of "behavior 1 → behavior 2" to obtain the behavior transfer order sequence.
[0032] S303: Based on the behavior transfer order sequence, update the direction field of the corresponding connection line in the graph structure, replace the original direction data, fill in the new direction order according to the time progression order in the connection field position, output the direction update result, and obtain the corresponding content of the structure direction; First, the behavioral transmission direction information between each group of nodes in the sequence is read, and the corresponding connection line field is retrieved from the graph structure data. The path number field in the connection line field is used to determine its position index in the connection data table. The original direction value in the connection field is compared and replaced. If the current behavioral transmission direction is inconsistent with the original direction, the old value is overwritten with the corresponding direction identifier in the sequence. The overwrite operation uses the unique number of the connection field as the retrieval condition to ensure that the direction field of each path is updated only once. Then, the time progression order in the behavioral transmission sequence is called to adjust the arrangement of the connection lines in the path table in a temporal order, placing the connection records with earlier times at the beginning of the path and the later times at the end. The connections are sequentially pushed to the end of the path according to the order of the behavior occurrence. During this process, the index value of the direction field of each connection record is reassigned. The index value is determined by the time interval Δt within the sequence. Δt is the difference in timestamps between two adjacent direction records. For example, if the behavior occurrence times between node A and node B are t1 and t2 respectively, their direction numbers are calculated and arranged according to the order difference of t2−t1 to ensure that the time progression order of the connection direction is consistent with the behavior occurrence order. After all the path node direction fields are updated, the updated direction values are filled into the corresponding positions of the graph connection fields to form a complete direction mapping table, and the connection status after the direction transformation in the overall structure is output to obtain the content corresponding to the structural direction.
[0033] Please see Figure 5 The specific steps of S4 are as follows: S401: Based on the content corresponding to the structural direction, retrieve adjacent data nodes in the path, extract the collection period and field change sequence of each group of nodes, use the field sequence as the identification basis, extract the sequence features between node pairs, and obtain the node field sequence information set. First, the numbers of each adjacent node in the path are read. The node field values and collection time data recorded in the graph structure are located by number. The content of the field group is extracted from the field recording area. The collection order is distinguished by the timestamp value attached to the field order. Then, the field name and timestamp value are combined to form a field order pair set. By comparing the position of the same field in different nodes in the set, it is determined whether the field maintains its original order in the path. If the position of the field in the later node is later than the position of the earlier node, it is marked as "moved to the back". If the position of the field in the later node is earlier, it is marked as "moved to the front". If the positions are the same, it is marked as "no change". In the field order comparison process, the field name is used as the retrieval basis. The sorting results of the field in each node are compared one by one, and the order status is filled into the field order status mark. Then, the comparison results of each group of fields are collected by path number. All field names and their corresponding order change status are uniformly collected on the field structure corresponding to each group of paths. By performing this kind of identification process on all node fields in the path, a list with field order change characteristics is finally output, resulting in a node field order information set.
[0034] S402: Based on the node field order information set, consolidate data paths with consistent field order, split data paths with changing field order, and separate data paths of differentiated field sequences into independent channels based on the field position order to obtain a field order classification information set; First, the field order list corresponding to each path is retrieved. Paths marked as "unchanged" are merged, treating their field order as the same structural pattern. The same identifier is assigned to these paths in the storage structure to maintain consistency. Next, paths with changed field orders are split. The range of the changed segment is determined by identifying the start and end positions of the field change. The data segment of the changed path is partitioned based on the order of the fields among nodes. Within each region, the changed field names and their corresponding node numbers are extracted, generating a field order difference table. This table describes the differences in field position between different paths. For example, when a field in a certain path... When node f2 is in position 2 at node A and position 4 at node B, the path is identified as a path of changing order and is grouped into a differentiated path group. Subsequently, based on the positional changes of the differentiated fields among the nodes, these paths are divided into independent channels. Each channel contains only a set of paths with consistent field arrangement trends. To ensure the accuracy of path grouping, the field order within each channel is checked again to eliminate overlapping or reversed node relationships in the sequence, ensuring that each channel maintains a single-direction field order pattern. Finally, the field path data corresponding to all channels are compiled into a classification index table, recording the channel number, field order, and node mapping information, and outputting the summarized path partition structure to obtain the field order classification information set.
[0035] S403: Based on the field order classification information set, the paths corresponding to the numbers are placed in independent sequence areas, and the path information is collected in the areas. The data sequence belonging relationship is output according to the number partitioning result to obtain the path belonging information. First, obtain the corresponding number information and path number for each field's sequence structure. Perform a traversal operation on each path number, identifying its corresponding field sequence number during the traversal. Paths with the same field sequence are grouped together in the same logical region for parallel processing, while paths with different numbers are separated into independent logical regions. Each region is configured with an independent dataset container for path content aggregation. During path import, extract the path's node order, field order, direction information, and behavioral characteristics, filling them into the path data record format and binding them with the field sequence number to form a number-path mapping pair. Within the same logical region, organize the data according to the field sequence number to ensure consistent information structure among paths. If a path is found during the aggregation operation... If some fields are missing or inconsistent in order in a path, the path's inclusion in the current region is stopped and the process moves to the next path. This ensures a consistent path structure within the region. A path information table is constructed using the path number as the primary index. In this table, each path is assigned fields such as the region number, the number of field items within the path, the path length, and the identifiers of the first and last nodes. This structure completes the path attribution label definition operation. Combining the field sequence number and the path number as dual references, an attribution mapping record table is generated. By traversing and outputting, all path attribution relationships in the table are exported as structured text information. A path attribution index list is generated by sorting by path number. This list records the one-to-one correspondence between each path number and the region number, ultimately yielding the path attribution information.
[0036] Please see Figure 6 The specific steps of S5 are as follows: S501: Based on path attribution information, extract data nodes and corresponding connection data fragments within the path, call the path identifier field, extract the node content and connection information within the same path respectively, and arrange them into the path channel area according to the order of data appearance to obtain the path channel permutation set; First, the path number field is retrieved from the path attribution table. Each path's node list and connection segment information are then searched sequentially. During this process, the unique identifier of each node and the connection index value between adjacent nodes are extracted, establishing a one-to-one mapping relationship between nodes. This mapping determines the preceding and following connections for each node, completing the identification of the path's internal structure. Then, for each path, the corresponding data content for each node is extracted sequentially, including key operating parameters such as cooling water temperature, fan speed, inlet water pressure, and outlet water pressure. The timestamp, field name, and measured value for each node are extracted simultaneously. The nodes are arranged in the same logical sequence based on the natural order of the time field. Simultaneously, the start and end numbers of the connection segments are matched, creating a continuous chain relationship between node content and connection segments. Finally, the path number is used as the index... The process involves summarizing and organizing all node and connection information, maintaining the correspondence between node sequences and connection segments in the memory structure, and constructing a complete path data path instance. For example, in the path numbered P-12, the connection segment C-1 between nodes N-1 and N-2 will be arranged in the order of timestamps t-1 to t-2, with node N-3 following it to ensure consistent order within the path path. Then, based on the sequential position of the nodes in the path, channel position parameters are assigned to each node, and the channel numbers of adjacent nodes are arranged consecutively. Finally, all node sequences corresponding to the path are stored in the channel area to form a path channel arrangement structure set. This structure uses path number, node sequence number, and connection index as core fields to achieve partitioned arrangement and mapping of the internal structure of the path, resulting in a path channel arrangement set.
[0037] S502: Based on the path channel arrangement set, compare the node content tags, retrieve the existing attribute tag content of each node, and concatenate the tag information with the current node's fields. At the same time, juxtapose the associated connection data content to obtain the node tag concatenation sequence. First, the node field information is read item by item for each node in the path sequence. The node name and node number fields are used as indexes to point to the tag content table. Based on the node number field, the attribute tag value sequence corresponding to the current node is retrieved from the tag set. At the same time, the relative position of the node in the current channel path is marked. Then, combined with the upstream and downstream segments connected to the node, the content text of the corresponding connection segment is extracted through the connection index field. The relational terms and operation instructions in the connection segment content are decomposed and assigned to the current node tag field structure. They are then merged with the node tag items into the same output field group through field shifting. For example, in a certain path, the node numbered N- The node 01 has the attribute "temperature control" and the connection segment is described as "control output → heating". Therefore, "control output" is extracted as the output field fragment of this node and concatenated as its tag extension after the "temperature control" field to form the combined tag "temperature control - control output". Then, the nodes in the path channel are processed in the same way for each subsequent node tag. During the process, there is no need to read the node tag table repeatedly. Instead, the fields are quickly associated with the already called tags in the cached area by referring to the dictionary. This avoids multiple accesses to the tag library affecting the channel processing efficiency. Finally, all node tags are concatenated and the paragraph text they are connected to are imported into the field block to obtain the node tag concatenation sequence.
[0038] S503: Based on the concatenated data content in the node label concatenation sequence, write the node field and connection field together into the target data storage area, and use the path channel as the dividing basis to divide the area to obtain the structured storage result; First, the node field group and the connection field group are read from the concatenated sequence. The node number, name, attribute label, and corresponding time series field are extracted sequentially from the node field group, and a field index table is created as a reference for data writing. In the field index table, the path number is used as the primary key, and the node number as the secondary key. Field-level alignment is performed on the node fields and connection fields to ensure a one-to-one correspondence between node content and its connection fields under the same path channel. Then, the field mapping structure of the target storage area is called to create a corresponding storage block area for each channel. Each storage block has two sub-areas: a node data segment and a connection data segment. Node field content is written to the node data segment, and connection field content is written to the connection data segment, maintaining the consistency of their order within the data block. During the writing process, the path number is adjusted accordingly. The storage block index is automatically switched to achieve partitioned writing of path channels. For example, when the path channel number is C-01, its node fields such as temperature, pressure, and fan speed will be written to the node data segment in sequence. The corresponding connection fields such as flow direction and adjacent node number will be written to the connection data segment at the same time, and are uniformly assigned with the identifier C-01. After all the node and connection data of all path channels have been written, the path partition index table is called to check whether the number of nodes and the number of connections in each storage block are consistent. If there is a discrepancy in the number of fields, the missing items are filled or the redundant items are truncated according to the node number increment rule. Finally, all the node data and connection data of all path channels have been stored accordingly, forming a path partitioned data set, and a structured storage result is obtained.
[0039] Please see Figure 7 A knowledge graph-based industrial data storage system includes: The data acquisition module collects cooling tower operation data by time, locates the data source according to the number field, adds corresponding identification content to each record, links the identification with the value in the corresponding field, and inputs it into the graph node record area, adds the original data field, and obtains semantic mapping content. The node mapping module selects continuously recorded data nodes based on semantic mapping content, calls the time tag of each group of data, analyzes the information transmission direction and connects adjacent nodes, inputs the connection results into the graph path, and obtains the process connection relationship content. The relationship identification module extracts node attribute labels based on the process connection relationship content, reads the behavior type of the current node, identifies the data transmission direction, updates the connection relationship direction and writes it into the graph to obtain the content corresponding to the structural direction. The path attribution module retrieves adjacent nodes in the path based on the content corresponding to the structural direction, extracts the period and field change order, combines paths with the same order into the same channel, and divides paths with different orders to obtain path attribution information. The structure storage module extracts nodes and connection relationships based on path attribution information, fills the node content and connection lines into the processing area, adds labels and writes the node relationships to obtain the structured storage result.
[0040] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for storing industrial data based on a knowledge graph, characterized in that, The method comprises the following steps: S1: obtaining cooling tower operation data, collecting data by time, locating data source according to number field, adding corresponding identification content to each record, connecting identification and value in corresponding field, and inputting into graph node record area to obtain semantic mapping content, adding original data field; S2: based on the semantic mapping content, selecting continuous record nodes, calling each group of data time labels, analyzing information transmission direction and connecting adjacent nodes, inputting the connection result into the graph path to obtain the process connection relationship content; S3: based on the process connection relationship content, extracting node attribute labels, reading the behavior type of the current node, identifying the data transmission direction, updating the connection relationship direction and writing into the graph to obtain the structure direction corresponding content; S4: based on the structure direction corresponding content, searching the path link adjacent nodes, extracting the period and field change order, combining the paths with the same order in the same channel, and dividing the paths with different orders to obtain the path attribution information; S5: based on the path attribution information, extracting nodes and connection relationships, filling node content and connection lines into the processing area, adding labels and writing into node relationship to obtain the structured storage result. 2.The knowledge graph-based industrial data storage method according to claim 1, characterized in that, The semantic mapping content includes cooling water temperature identifier, fan speed identifier, inlet and outlet water pressure identifier, collection time label, number field mapping and original data field. The process connection relationship content includes time label sequence, node connection order, information transmission path and process node pair. The structure direction corresponding content includes behavior type label, sequence information, action direction label and connection direction data. The path attribution information includes node record period, field change order, path grouping label and processing channel identifier. The structured storage result includes path area division, node connection line, label addition information and node association data. 3.The knowledge graph-based industrial data storage method of claim 1, wherein, The corresponding identification content refers to the unique number information added to each data record in the process of collecting cooling tower operation data, which identifies the data source and attribution equipment; The original data field refers to the original monitoring data field of cooling water temperature, fan speed and inlet and outlet water pressure obtained in the collection process, with unit and sampling information. 4.The knowledge graph-based industrial data storage method of claim 1, wherein, The graph path refers to the directed path generated according to time and behavior sequence connection logic between data nodes, reflecting data transmission, behavior causality and process relationship, recorded in the graph structure; The processing area refers to the target data structure area for receiving and storing node and connection relationship, which is arranged, labeled and written into the structured storage result. 5.The knowledge graph-based industrial data storage method according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: obtaining cooling water temperature, fan speed, inlet water pressure and outlet water pressure data in the cooling tower operation process, comparing data frames by collection time, arranging data with consistent time in the same sequence, extracting field set in the sequence and positioning field position, and associating data in the time sequence to obtain data mapping table in the collection period; S102: Based on the data mapping table in the collection period, call the number field information, identify the data source, arrange the field values in order according to the number sequence, append the number value to the corresponding field position, splice and output the data content with source identification, and get the data source identification mapping result; S103: Based on the data source identification mapping result, import the data content into the graph node area, write the number value and field value correspondence into the node structure, append the original field information and timestamp content, expand the expression range of the graph node, and get the semantic mapping content. 6.The knowledge graph-based industrial data storage method according to claim 1, characterized in that, The specific steps of S2 are: S201: Based on the semantic mapping content, select the data nodes with continuous time in the cooling water temperature, fan speed and inlet and outlet water pressure, call the timestamp field of each data, arrange the nodes in the data sequence according to the time sequence, locate the node pair number under continuous time, and get the node time sequence index pair; S202: Based on the node time sequence index pair, extract the timestamp values of adjacent nodes, compare the time field sequence, identify the time arrangement order, connect the nodes according to the time arrangement order as data transmission path, output the time sequence information corresponding to the path as sequence, and get the node transmission direction sequence; S203: Based on the node transmission direction sequence, link adjacent nodes in transmission direction in turn, write each node pair into path sequence field, and append the sequence value and node number to the graph path area. According to the time advancing relationship, number the path content position, and get the process connection relationship content. 7.The knowledge graph-based industrial data storage method according to claim 1, characterized in that, The specific steps of S3 are: S301: Based on the process connection relationship content, extract the field label in each pair of connected nodes, identify the node behavior type, extract the behavior type information according to the field sequence, extract the result field, and put the label and behavior type into the corresponding relationship, get the node behavior attribute set; S302: Based on the node behavior attribute set, select the behavior types of the double nodes in the connection path, call the node time field, judge the order of behavior occurrence, arrange the recognition results in time sequence as action direction expression, get the behavior transmission sequence, and get the structure direction corresponding content. S303: Based on the behavior transmission sequence, update the direction field of the corresponding connection line in the graph structure, replace the original direction data, fill the new direction sequence into the connection field position according to the time advancing sequence, output the direction update result, and get the structure direction corresponding content. 8.The knowledge graph-based industrial data storage method of claim 1, wherein, The specific steps of S4 are: S401: Based on the structure direction corresponding content, retrieve adjacent data nodes in the path, extract the collection period and field change order of each group of nodes, use the field order as the identification basis, extract the order features between node pairs, and get the node field order information set; S402: Based on the node field order information set, group the data paths with consistent field order, split the data paths with field order transformation, distinguish the data paths with different field sequences according to the field position order, and get the field order classification information set; S403: Based on the field order classification information set, the numbered corresponding path is placed in an independent sequence area, and the path information is collected in the area. The data sequence attribution relationship is output according to the numbering partition result, and the path attribution information is obtained. 9.The knowledge graph-based industrial data storage method of claim 1, wherein, The specific steps of S5 are: S501: Based on the path attribution information, the data nodes and corresponding connection data segments in the path are extracted, the path identification field is called, the node content and connection information in the same path are extracted respectively, and the data appearance order is arranged to the path channel area to obtain the path channel arrangement set; S502: Based on the path channel arrangement set, the node content label is compared, the existing attribute label content of each node is searched, the label information is spliced with the current node, and the associated connection data content is juxtaposed to obtain the node label splicing sequence; S503: Based on the spliced data content in the node label splicing sequence, the node field and the connection field are written into the target data storage area, and the path channel is used as the division basis to divide the area to obtain the structured storage result.
10. An industrial data storage system based on a knowledge graph, characterized by, The system is used to realize the knowledge graph-based industrial data storage method of any one of claims 1-9, and the system comprises: The data acquisition module collects cooling tower operation data, collects data by time, positions the data source according to the number field, appends corresponding identification content to each record, connects the identification and the value in the corresponding field, and inputs into the graph node record area to add the original data field to obtain the semantic mapping content; The node mapping module selects the data nodes of continuous records based on the semantic mapping content, calls the data time label of each group, analyzes the information transmission direction and connects the adjacent nodes, inputs the connection result into the graph path, and obtains the process connection relationship content; The relationship identification module extracts the node attribute label based on the process connection relationship content, reads the behavior type of the current node, identifies the data transmission direction, updates the connection relationship direction and writes it into the graph to obtain the structure direction corresponding content; The path attribution module extracts the period and the field change order based on the structure direction corresponding content, combines the paths with the same order in the same channel, divides the paths with different orders, and obtains the path attribution information; The structure storage module extracts the node and the connection relationship based on the path attribution information, fills the node content and the connection line into the processing area, adds the label and writes it into the node relationship to obtain the structured storage result.
Citation Information
Cited By
Intelligent data processing method and system based on knowledge graph
CN121051252A
Big data processing method and system for filtrate detection
CN122196508A