Non-cooperative industrial data scheduling and storage method
Through non-collaborative industrial data scheduling and storage methods, the tree-like structure data and persistent index of hierarchical markers are used to solve the problems of data processing logic black boxes and resource waste in the collaborative method, and more efficient and flexible data processing and storage are achieved.
Patent Information
- Application Number
- CN202510175508.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-27
AI Technical Summary
The existing collaborative industrial data scheduling and storage methods have problems such as data processing logic black boxes, waste of resources, poor adaptability, and poor performance, which are difficult to meet the complex and changeable industrial data processing needs.
Non-collaborative industrial data scheduling and storage methods are adopted to generate persistent data indexes by receiving tree-structured industrial data with hierarchical markers, and a combination of volatile and nonvolatile storage is adopted to achieve efficient preprocessing, persistence and query of data.
It improves the system's compatibility with heterogeneous data, improves data processing and persistence efficiency, provides better query performance and flexibility, and reduces resource usage and storage space.
Smart Images

Figure CN120218464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a non - collaborative industrial data scheduling and storage method, belonging to the fields of industrial data acquisition, processing, and application. Background Art
[0002] Currently, the acquisition, processing, storage, query, and application of industrial data generally adopt a collaborative solution, as Figure 1 shown. For scenarios with simple business requirements, small data volume, and small data calculation volume, a direct cloud acquisition solution is adopted. In scenarios with diverse businesses, multiple data types, high data acquisition frequencies, and complex logical relationships between data, an edge acquisition and edge pre - processing solution is adopted, and the data is uploaded to the cloud through a collaborative data interface and a collaborative data format.
[0003] In the industrial data processing link, for different types and multi - source heterogeneous data types, engineering and technical personnel in related fields often need to be deeply involved, communicate deeply with data owners, and then complete corresponding coding operations according to the model and characteristics of the edge acquisition gateway to transmit the pre - processed data to the cloud. This process actually severs the relationship between data owners and data processing behaviors. Each industry and each type of production equipment has its own data processing logic. Only data owners who have been deeply involved in this industry for many years can give full play to their experience and use the most appropriate calculation logic to complete the processing of raw data. In the current market, data owners only act as messengers, and the data processing logic is a black box to them. They cannot flexibly and timely adjust the data processing logic according to business changes or other reasons.
[0004] In the collaborative data scheduling and storage solution, even if some configurable items that can be modified by data owners and have a certain degree of freedom are added to the cloud, a series of problems will still exist. The most prominent one is that these logics are not hard - coded. Compared with the hard - coding mode, the central processing unit, memory, and disk of the server need to spend more hardware resources and time to complete the same processing logic, and these extra costs will increase exponentially according to the complexity of the processing logic.
[0005] In the data transmission link, some collaborative industrial data scheduling methods will package data itself, data meaning, events to be processed, etc. into a large data packet through data redundancy for transmission, so that the next data processing link can complete the processing logic through redundant data information. However, this approach will cause a large amount of network bandwidth occupation, additional memory and processor resource overhead, and in high - frequency acquisition scenarios, thousands of large data packets have the same information, which is undoubtedly a waste of resources.
[0006] In the industrial data storage process, it is necessary to ensure the integrity, timeliness, and accuracy of data. However, in the current market, there is actually no relatively perfect solution. Among them, the collaborative data scheduling method can ensure better performance during storage due to its relatively fixed data receiving interfaces and data formats, and has good adaptability in fixed production scenarios. However, once there are minor changes in the production scenario and change requirements arise, this collaborative data scheduling method cannot well adapt to the new scenario, and engineering and technical personnel need to intervene to expand the cloud processing logic in order for the system to continue to provide services, and this workload will also vary according to the changes in the new scenario. Generally, non-collaborative approaches are even more immature, usually having very poor performance, and even in simple production scenarios, they cannot provide timely and high-quality services to data owners. Therefore, the collaborative data scheduling method is still relatively popular in the current market.
[0007] In industrial data query, similar to the data storage process, the collaborative data scheduling method needs to write corresponding hard-coded programs according to the established data interfaces, data formats, and unique data attributes to provide services to data owners, greatly limiting flexibility. Non-collaborative data query is even difficult to implement. Due to different data formats, engineering and technical personnel must intervene to complete the hard-coded programs for query methods to be completed.
[0008] In the collaborative data scheduling method, engineering and technical personnel will pre-conceive the transfer method for each piece of data in advance, and the data will proceed through its life cycle accordingly. Therefore, it is difficult to continue the subsequent continuous role of the data. As a data owner, it is also very difficult to integrate one's own understanding of the data and industry experience into the digital system, and it is difficult for the system to continue to integrate. Summary of the Invention
[0009] Object of the Invention: Aiming at the problems existing in the above-mentioned prior art, the object of the present invention is to provide a non-collaborative industrial data scheduling and storage method to improve the compatibility of the system with heterogeneous data from different sources, improve the efficiency of data processing and data persistence, and further provide better performance, compatibility, and flexibility for data query and secondary processing.
[0010] Technical Solution: To achieve the above object of the invention, the present invention adopts the following technical solutions:
[0011] In the first aspect, the present invention provides a non-collaborative industrial data scheduling and storage method, including the following steps:
[0012] Receiving tree-structured industrial data with hierarchical markings; the data source information is marked in the tree structure, and the data format is located through the unique marking corresponding to the data source, and the parsing marking of each data item in the data format is mapped through a defId;
[0013] Add the raw data to the data queue in the first area of the volatile storage; retrieve a piece of raw data from the data queue, locate the srcId of the data source, and after filtering and classifying according to the custom policy of each data node in this piece of raw data, obtain several groups of preprocessed data with the same data source and the same processing policy, and store each group of preprocessed data in a continuous area in the second area of the volatile storage;
[0014] Generate a persistent data index for the data in the second area; the persistent data index is a hash index generated according to the combination of the data source srcId or the agreed preset number of digits therein, the data item marker mapping defId or the agreed preset number of digits therein, and the year-to-minute information of the time stamp corresponding to the data;
[0015] Write the entire group of data with the persistent data index that meets the write condition into the non-volatile storage.
[0016] Preferably, when writing data into the non-volatile storage, only write the hash index, the formatted data, and the data type enumeration; the data type enumeration is used to indicate the type of data, including one or more of text type, number type, list type, range type, set type, and mapping type.
[0017] Further, the steps for querying the real-time data of a specified data item in a specified data source include:
[0018] Obtain the specified data source or the agreed preset number of digits therein, and the specified data item marker or the agreed preset number of digits therein;
[0019] Directly query the latest piece of data in the corresponding category in the second area of the volatile storage. If it cannot be found, generate a hash index by combining the data source, the data item marker, and the current time information, and locate the corresponding data in the non-volatile storage.
[0020] Further, the steps for querying the data of a specified data item in a specified data source within a specified time period include:
[0021] Obtain the specified data source or the agreed preset number of digits therein, and the specified data item marker or the agreed preset number of digits therein;
[0022] Generate a hash index for each query according to the number of time points within the specified time period, and use the hash index to locate the data in the non-volatile storage.
[0023] Further, the steps for querying the duration during which the data of a specified data item in a specified data source within a specified time period meets a specified rule include:
[0024] Obtain the specified data source or the agreed preset number of digits therein, and the specified data item marker or the agreed preset number of digits therein;
[0025] Generate a hash index for each minute within a specified time period, use the generated hash index to locate a block of data, copy the data into the server memory, filter the data that meets the preset conditions of the query, and merge the time information of the qualified data to obtain the final result.
[0026] Furthermore, it also includes basic operations such as built-in addition, subtraction, multiplication, division, logical operations, mapping tables, and string concatenation to support secondary processing of data, and generate corresponding srcIds and defIds for each item of data generated by the secondary processing; the mapping table is a user-defined operation for obtaining an output value based on the input parameters.
[0027] In a second aspect, the present invention provides a cloud platform for non-collaborative industrial data scheduling and storage, including:
[0028] A data receiving module for receiving tree-structured industrial data with hierarchical tags; the tree structure is marked with data source information, and the data format is located through the unique tag corresponding to the data source, and the parsing tag of each data item in the data format is mapped through a defId.
[0029] A data preprocessing module for adding unprocessed raw data to the data queue in the first area of the volatile storage; taking out a piece of raw data from the data queue, locating the srcId of the data source, and after filtering and classifying according to the custom policy of each data node in this raw data, obtaining several groups of preprocessed data with the same data source and the same processing policy, and each group of preprocessed data is stored in a continuous area in the second area of the volatile storage.
[0030] A persistence module for generating a persistence data index for the data in the second area; the persistence data index is a hash index generated according to the combination of the data source srcId or the agreed preset number of digits therein, the data item tag mapping defId or the agreed preset number of digits therein, and the year-to-minute information of the time stamp corresponding to the data; and writing the entire group of data with the persistence data index that meets the writing condition into the non-volatile storage.
[0031] Preferably, the cloud platform performs data reception and preprocessing through the network architecture of a multi-server cluster.
[0032] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program / instructions stored on the memory and executable on the processor, and when the computer program / instructions are executed by the processor, the steps of the non-collaborative industrial data scheduling and storage method described above are implemented.
[0033] Fourthly, the present invention provides a computer program product, including computer programs / instructions, which when executed by a processor implement the steps of the non-collaborative industrial data scheduling and storage method described above.
[0034] Advantages: Compared with the prior art, the present invention has the following advantages:
[0035] 1. Compared with the traditional cloud-edge collaborative data transfer method, the present invention does not limit the specific format of the uploaded data, as long as the format is a tree-like format with hierarchical indexing similar to "protobuf / Json", greatly improving the compatibility of the system.
[0036] 2. Compared with the ordinary non-collaborative data transfer method, the present invention adopts the strategy of caching data in partitions of the same type and source and persisting batch data in partitions, greatly improving the efficiency of data preprocessing and data persistence, and its performance even exceeds that of the traditional cloud-edge collaborative data transfer method in tests.
[0037] 3. In the data index design of the present invention, a hash index composed of data source + data item definition + time information is adopted. Compared with the random unique index / incremental index + multi-dimensional index method commonly used in common Internet of Things products, the physical storage space occupied by the index is greatly reduced, and the query efficiency of this index is also much higher than that of the latter on the premise of meeting various general query interfaces.
[0038] 4. In the data index design of the present invention, a data item is uniquely determined by the combination of srcId and defId, and the same data index method is adopted for both the original data and the internal system data obtained through secondary and tertiary processing. This extends the life cycle of all data in the system, enabling data owners to easily operate the data using the basic operation methods provided in the system without the intervention of engineering and technical personnel, quickly and losslessly integrating industry experience into the system, saving many costs for production enterprises in secondary development and project construction, and enabling many originally customized business logics to be quickly launched. Description of the Drawings
[0039] Figure 1 It is a reference diagram of the existing conventional collaborative industrial data processing flow.
[0040] Figure 2 It is a reference diagram of the non-collaborative industrial data processing flow of the present invention.
[0041] Figure 3 It is a schematic diagram of a typical data format.
[0042] Figure 4 It is a schematic diagram of another typical data format.
[0043] Figure 5 is the flowchart of the method according to an embodiment of the present invention.
[0044] Figure 6 is a schematic diagram of the markup of a typical data format.
[0045] Figure 7 is a schematic diagram of the markup of another typical data format.
[0046] Figure 8 is a screenshot of the markup interface exemplified in an embodiment of the present invention.
[0047] Figure 9 is a schematic diagram of the cache preprocessing process in an embodiment of the present invention.
[0048] Figure 10 is a schematic diagram comparing the indexing rules with the existing rules in an embodiment of the present invention.
[0049] Figure 11 is a schematic diagram of batch storage in an embodiment of the present invention. Detailed implementation manners
[0050] Next, the technical solution of the present invention will be clearly and completely described in conjunction with the accompanying drawings and specific embodiments.
[0051] Figure 2 Schematically shows the non - collaborative industrial data processing flow designed by the present invention. The non - collaborative industrial data described in the present invention refers to data with the following characteristics:
[0052] (1) It has a nestable tree - like structure, such as data formats common in the Internet of Things field, such as protobuf and Json. After the data in the industrial field is converted by common edge gateways on the market, it generally has such a format.
[0053] (2) There is at least one node in the tree - like structure that can mark the physical source information of the data.
[0054] (3) The tree - like structure of the data has various presentation forms, which are not limited.
[0055] Figure 3 and Figure 4 illustrate two typical structures.
[0056] Embodiment 1
[0057] As Figure 5 shown, a non - collaborative industrial data scheduling and storage method disclosed in an embodiment of the present invention mainly includes the following steps:
[0058] Step S1, the cloud receives tree - like structure industrial data with hierarchical markings.
[0059] In this step, the data source information is marked in the tree structure, and the data format is located through the unique mark corresponding to the data source. Each data item in the data format is mapped by a defId. The structural form of industrial data is a tree structure that includes Json, XML, YAML, Bson, FRON, protobuf, etc. and may have hierarchical marks in the future.
[0060] Specifically, to identify, process, and store data in different formats in a unified manner in the cloud and provide a general and high-quality query interface, it is first necessary to determine the structure of the data and the physical meaning of each data node. Then mark these nodes in the cloud system and set a series of Figure 6 and Figure 7 as shown in the data information. These marked data information can be arbitrarily added, deleted, modified, and queried throughout the entire life cycle of the data. Figure 8 An example of a marked interface display form is shown.
[0061] After marking a piece of data, the system can obtain the following information:
[0062] ① The data source srcId.
[0063] ② The data format format corresponding to the data source. (The system needs to locate the approximate range of the data source based on the data format. Data with the same data format may come from different data sources, so it is also necessary to obtain the data source information from the data; and the same data source may also send data in multiple different data formats, and these data also need to be distinguished and processed within the system)
[0064] ③ The unique mark srcMarkValue corresponding to the data source. As Figure 8 sets the srcMarkValue on the DeviceUid.
[0065] ④ In the data format corresponding to the data source, the attributes dataStats (data type, data meaning, other custom attributes) of each data point, the parsing mark parseMark of each data point (composed of the path of the tree structure, the last-level path is composed of the key of the data, and in the case of no key, it is represented by the order of the node where the data is located), and each parsing mark will be mapped by a random defId in the system.
[0066] ⑤ The mapping relationship set formatDictSet between the data format format corresponding to the data source and the unique mark srcMarkValue corresponding to the data source. Data from different data sources may have the same data format. Therefore, in the system, it is necessary to save which data sources correspond to this data format for subsequent accurate positioning through srcMarkValue. For data with the same data format, the parsing mark parseMark of its unique mark srcMarkValue in the data format may also be different. Therefore, after positioning the approximate range formatDictSet of the data source through the format information, it is necessary to locate the position of srcMarkValue in the format through the parseMark again, and finally determine its true data source by comparing the values of srcMarkValue.
[0067] ⑥ The basic processing strategies for each data node (a series of custom strategies such as whether to persistently store, whether to store with changes, whether it is increment-only measurement data, etc.). These strategies determine how the system stores and statistics these data.
[0068] Publish data from the edge acquisition end to the cloud. This process usually uses the TCP long connection method to transfer data evenly and periodically. To solve the performance problem of non-collaborative data flow, a network architecture of multiple server clusters can be adopted in this link to ensure the real-time reception of data and prevent network congestion.
[0069] Step S2: The cloud preliminarily processes the data, adds the unprocessed multi-source heterogeneous raw data into the data queue in the first area of the volatile storage; takes out a piece of raw data from the data queue, locates the data source srcId, and after filtering and classifying according to the custom strategies of each data node in this raw data, obtains several groups of pre-processed data with the same data source and the same processing strategy, and each group of pre-processed data is stored in a continuous area in the second area of the volatile storage.
[0070] Specifically, since no preprocessing and threshold setting are done in the data reception link in this embodiment, the amount of received data is very large. Although this reduces the pressure on the edge side and network load and shows good performance for the edge side and network load, it also greatly increases the pressure on cloud data processing. Therefore, the cloud has done the following processing after receiving the data:
[0071] ① Add the unprocessed data into the data queue in the volatile storage. This queue exists in the first area (denoted as area A) of the volatile memory.
[0072] ②All servers in the cloud simultaneously retrieve data from the data queue in the volatile storage, locate the data source srcId, and execute their basic processing strategies for each data node in this data. For example, if it is set that a certain data node does not need to be persistently stored, then the current data node does not need to continue with subsequent processing operations, which can, to a certain extent, reduce the computing pressure on the server. By storing data based on whether it has changed, the storage space of some high-frequency data can be compressed. When the data remains unchanged, the system only records the time of the data and does not repeatedly record the value of the data. This step can also handle new requirements put forward by end customers. For example, all collected data can be treated as positive values. Typically, for the spindle speed of a machine tool, whether it rotates forward or backward is not concerned, only the speed itself is concerned. Then, an option "whether to process as an absolute value" can be added during processing here. When marked to be processed as an absolute value, an additional copy of its absolute value will be stored persistently in the system when the value is persisted.
[0073] ③After filtering and classifying the data in ② according to the basic processing strategy, several groups of data with the same data source and the same processing strategy can be obtained. As Figure 11 shown, these data are batch-imported into the second area (denoted as area B) of the volatile memory according to the classification. Each classification occupies a continuous area therein. The volatile memory mentioned above generally refers to the memory module of the server. The read and write speed of this module is the fastest hardware memory in the current computer field, but the information it stores is easily lost in the case of power failure or loss of storage index. Therefore, a storage snapshot mechanism is also introduced to ensure data security. Any mature memory snapshot software product on the market can be used to implement this here.
[0074] Step S3: Generate a persistent data index for the data in the second area; the persistent data index is a hash index generated based on the data source srcId or a preset number of digits therein, the data item marker mapping defId or a preset number of digits therein, and the year-to-minute information of the time stamp corresponding to the data.
[0075] The cloud persistent data starts. In step S2, several data sources of data to be persisted have been stored in area B of the volatile memory according to the classification respectively. Some of these data can be directly persisted, and for those with custom rules set, in addition to being persisted according to the system default persistent rules, they also need to be persisted in a specific format. The custom rules can be horizontally extended according to the application scenario and will not be elaborated in this invention.
[0076] Specifically, the core of this step is to generate the default persistent data index, that is, generate as Figure 10For the hash index shown, obtain the specified data source srcId and take the specified n bits therefrom (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios); obtain the specified data item marker mapping defId and take the specified m bits therefrom (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios). Combine the year, month, day, hour, and minute information of the timestamp corresponding to the data with the former two to generate an index, and store the data according to this index. Compared with the traditional multi-dimensional index data persistence method, it greatly reduces the number of indexes and the length of the indexes. The size of the storage space occupied by the index is determined by the number of different indexes in each piece of data, and for each additional index, several times more persistent storage space will be occupied. Therefore, using this data persistence method of the present invention can reduce the storage space occupancy by at least one order of magnitude compared with the traditional method; from the perspective of data redundancy, in the data persistence method of the present invention, useless information such as "random value + hardware number" is removed, and data such as "industrial equipment / asset number, data item name, other attributes of the data item" are uniformly stored in the data format format corresponding to the data source, without the need to be bound to the data itself. Only the hash index, formatted data, and data type enumeration need to be stored. The data type enumeration is used to indicate the type of data, including text type (string), numeric types (int, float, double, bool, byte, etc.), and other general data structure types in the computer field (map, range, list, set, etc.). Such a design reduces the storage occupancy of the data by more than half, and the longer the system runs, the more space is reduced. After actual measurement, compared with the traditional solution, the solution of the present invention saves more than 75% of the persistent storage space after long-term use.
[0077] Step S4, batch write to the non-volatile memory, and write the entire group of data with the persistent data index that meets the write condition into the non-volatile memory.
[0078] Specifically, as Figure 9After the above step S3, several types of data sets with default persistent data indexes (hash indexes) can be obtained. The system will monitor these sets. When their size reaches the specified size set by the system or reaches the set persistence period, the entire set that meets the write condition will be written into the non-volatile memory. This writing method can ensure the writing order, and data with similar attributes or data with similar logical relationships will be written into adjacent areas of the non-volatile memory; it needs to be explained here that non-volatile memory is a memory that will not lose data even if the power is off, usually represented by common external memories such as disks, mechanical hard disks, and solid-state hard disks. Among them, mechanical hard disks have the lowest cost and better stability, but because they rely on head seek when reading and writing, their read and write speed is relatively slow among all memories, and their read and write performance bottleneck lies in the head seek time; but since the present invention writes data with similar attributes or data with similar logical relationships into adjacent areas of the non-volatile memory, when querying a single type of data or querying logical correlation data, if the server uses a low-cost mechanical hard disk, the performance will be greatly improved compared to traditional storage methods.
[0079] Example 2
[0080] Based on the format of persistent data stored using hash indexes in the above embodiment, some common data queries and combined queries can be implemented.
[0081] In a non-cooperative industrial data scheduling and storage method disclosed in an embodiment of the present invention, the implementation principles of several general queries are as follows:
[0082] 1. Query the real-time data of the specified data item in the specified data source
[0083] Get the specified data source srcId, take the specified n bits (set by the system, the more bits are set, the lower the probability of duplication, but the larger the storage space occupied, and 4 to 6 bits are generally used in scenarios); get the specified data item tag mapping defId, take the specified m bits (set by the system, the more bits are set, the lower the probability of duplication, but the larger the storage space occupied, and 4 to 6 bits are generally used in scenarios); directly query the latest data of this category in area B of the volatile memory. If it cannot be found, take the current time information and the aforementioned n bits of srcId plus the m bits of defId to form a hash index, and use the hash algorithm preset in the system to calculate and locate the corresponding data block in the non-volatile memory.
[0084] 2. Query the data of the specified data item in the specified data source in the specified time period
[0085] Obtain the specified data source srcId, and take the specified n bits from it (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios); obtain the specified data item marker mapping defId, and take the specified m bits from it (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios); obtain the number of data items cnt required by the queryer; according to the needs of the business scenario, different sampling methods can be used to provide data to the queryer, and the system defaults to the average sampling method; obtain the start time starttime and end time endtime that the queryer needs to query, and generate a hash index for each query. The formula is: Use this index to locate the data.
[0086] 3. Query the duration during which the specified data items in the specified data source meet the specified rules within the specified time period;
[0087] Obtain the specified data source srcId, and take the specified n bits from it (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios); obtain the specified data item marker mapping defId, and take the specified m bits from it (set by the system. The more bits are set, the lower the probability of repetition, but the more storage space will be occupied. Generally, 4 to 6 bits are sufficient for general usage scenarios); generate an index for each minute within the specified time period. The formula is:
[0088] The hash index of the i-th minute = the specified m bits of srcId + the specified n bits of defId + (starttime + i); use the above index to locate a block of data, copy the data to the server memory, filter the data that meets the preset conditions of the queryer, and merge their time information to obtain the final result;
[0089] Other queries: Other types of queries can be implemented by combining the above three queries.
[0090] Embodiment 3
[0091] Secondary processing of data. In the present invention, the management granularity of data goes deep into the combination of each srcId and defId. Therefore, if you want the specified data to participate in the specified calculation, you only need to know the combination of srcId and defId to obtain all the corresponding data. The system has built-in basic operations such as addition, subtraction, multiplication, division, logical operations, mapping tables, and string concatenation. By combining these basic operations, some business scenarios with complex logics can be realized. Among them, the mapping table is used for user-defined operation rules, and the output parameter value can be obtained according to one or more input parameter values.
[0092] For example:
[0093] (1) In the machining scenario, different types of machine tools and machine tool controllers have a variety of alarm codes. These alarm codes are generally combinations of letters and numbers, and often change with the software version of the machine tool controller. This requires actual processing personnel or equipment maintenance personnel to update the rules of these alarm codes in a timely manner (the rules include two parts: ① the true physical meaning corresponding to the alarm code, ② whether it is an ignorable alarm). In this scenario, experienced processing personnel or equipment maintenance personnel can obtain the combination of srcId and defId of the alarm code by marking which data item in the data format of the data source is the alarm code, and then participate this data in the "mapping table" calculation. The system will automatically generate a random combination of srcId1 and defId1 as the index of the calculation result [① the true physical meaning corresponding to the alarm code], and also generate a random combination of srcId2 and defId2 as the index of the calculation result [② whether it is an ignorable alarm]. The storage method of all calculation results is the same as the original data, so these alarm messages can be queried using the aforementioned several query methods.
[0094] (2) In the energy-saving scenario, there are many high-energy-consuming devices in the industrial field, which consume a large amount of water, electricity and gas resources every year. Each component of this type of device can be started and stopped according to the production process. In the production scenarios of each manufacturing enterprise, the timing of these starts and stops is different, and actual production participants and enterprise managers need to continuously adjust the empirical values based on the long-term data accumulated during the production process. This transformation process is a long-term process and cannot be achieved overnight. Taking the start and stop of an air conditioner as an example, when the temperature (the combination of srcId t and defId t ) in a specified monitored production area increases at a rate (obtained by subtracting the collected values of adjacent collection cycles to get a new combination of srcId deltaT and defId deltaT ) greater than the set empirical value (a fixed value within the system, also with srcId staticVal and defId staticVal ), the air conditioner is started (triggering an external control system signal); similarly, the trigger conditions for shutting down the air conditioner can also be set.
[0095] (3) Operational status and efficiency statistics scenarios of large-scale production equipment. There are hundreds of types of equipment in the industrial field, and each type has different models. Moreover, each production equipment has many actuators. Therefore, for each independent production equipment, the definition of its operational status is different. Even for completely identical equipment, in different production enterprises or workshops, different status definitions are used. This requires the managers on the production site to assist in defining these complex statuses. Using the data flow design of the present invention, the real-time operational status of the equipment can be comprehensively defined, considered, output, and recorded by combining the operating parameters of multiple equipment actuators.
[0096] For example: A certain enterprise defines that the operational status of a machine tool can consist of running, idling, idle, and stopped. Among them, the definition of running is [high spindle speed, high table load], the definition of idling is
[0097] [high spindle speed, low table load], the definition of idle is [low spindle speed, low table load], and the definition of stopped is [no spindle speed, no table load]; the system has obtained the spindle speed (the combination of srcId spspeed and defId spspeed ) and the table load (the combination of srcId tbload and defId tbload ). Only need to:
[0098] ① Apply the "mapping table" [data within a specified range is mapped to "high spindle speed", data within a specified range is mapped to "low spindle speed", data with a specified fixed value is mapped to "no spindle speed"] operation to the spindle speed data, and apply the "mapping table" [data within a specified range is mapped to "high table load", data within a specified range is mapped to "low table load", data within a specified range is mapped to "no table load"] operation to the table load data, and two strings ["high spindle speed / low spindle speed / no spindle speed" and "high table load / low table load / no table load"] can be obtained, and the system generates the combinations of srcId str1 and defId str1 , the combinations of srcId str2 and defId str2 for them respectively.
[0099] ② Apply the "string concatenation" operation to the above two strings to obtain a string
[0100]
"High spindle speed + high table load / High spindle speed + low table load / High spindle speed + no table load / Low spindle speed + high table load / Low spindle speed + low table load / Low spindle speed + no table load / No spindle speed + high table load / No spindle speed + low table load / No spindle speed + no table load"
[0101] ③ Apply the "mapping table" to the above string
"High spindle speed + high table load" is mapped to "running", "High spindle speed + low table load" is mapped to "idle running", "Low spindle speed + low table load" is mapped to "idle", "No spindle speed + no table load" is mapped to "shutdown", and others are all mapped to "abnormal"
[0102] ④ The system will take the data corresponding to the combination of srcId state and defId state as the operating state of the machine tool, and use this data for subsequent data statistics.
[0103] In summary: The data secondary processing method of the present invention can make the data generated in the system have the same data index, the same data storage and query methods as the original data. Therefore, the system described in the present invention extends the data life cycle, enables data to be continuously integrated, and this configuration method does not require the participation of engineering and technical personnel, only the participation of data owners with rich production and management experience in the industrial field is required.
[0104] Example 4
[0105] A cloud platform for non - collaborative industrial data scheduling and storage disclosed in an embodiment of the present invention includes:
[0106] A data receiving module, used to receive tree - structured industrial data with hierarchical tags; the data source information is marked in the tree structure, and the data format is located through the unique tag corresponding to the data source, and the parsing tag of each data item in the data format is mapped through a defId;
[0107] A data preprocessing module is used to add raw unprocessed data to a data queue in the first area of volatile storage; take out a piece of raw data from the data queue, locate the srcId of the data source, and filter and classify it according to the custom policy of each data node in this piece of raw data, and then obtain several groups of preprocessed data with the same data source and the same processing policy. Each group of preprocessed data is stored in a continuous area in the second area of volatile storage.
[0108] A persistence module is used to generate a persistence data index for the data in the second area; the persistence data index is a hash index generated according to the combination of the data source srcId or a preset number of digits therein, the data item marker mapping defId or a preset number of digits therein, and the year-to-minute information of the time stamp corresponding to the data; and write the same group of data with the persistence data index that meets the write condition into non-volatile storage as a whole.
[0109] For the specific implementation details of each module, please refer to the foregoing embodiments and will not be elaborated here.
[0110] Embodiment 5
[0111] A computer system disclosed in an embodiment of the present invention includes a memory, a processor, and a computer program / instruction stored on the memory and executable on the processor. When the computer program / instruction is executed by the processor, the steps of a non-cooperative industrial data scheduling and storage method are implemented.
[0112] Embodiment 6
[0113] A computer program product disclosed in an embodiment of the present invention includes a computer program / instruction. When the computer program / instruction is executed by the processor, the steps of a non-cooperative industrial data scheduling and storage method are implemented.
Claims
1. A non-cooperative industrial data scheduling and storage method, characterized in that: The steps include: Receive tree-structured industrial data with hierarchical tags; the tree-structured data is marked with data source information, the data format is located through a unique tag corresponding to the data source, and the parsing tag of each data item in the data format is mapped through a defId; Adding the unprocessed raw data to a data queue in the first area of the volatile storage; taking out a piece of raw data from the data queue, locating the srcId of the data source, and filtering and classifying according to the custom strategy of each data node in the raw data to obtain several groups of pre-processed data with the same data source and the same processing strategy, and storing each group of pre-processed data in a continuous area in the second area of the volatile storage; Generate a persistent data index for the data in the second region; the persistent data index is a hash index generated based on the data source srcId or a predetermined number of digits agreed therein, the data item tag mapping defId or a predetermined number of digits agreed therein, and the year to minute information of the timestamp corresponding to the data; The same group of data with persistent data index that meets the write conditions is written into non-volatile storage.
2. A non-cooperative industrial data scheduling and storage method according to claim 1, characterized in that: When writing data to non-volatile storage, only the hash index, formatted data, and data type enumeration are written; the data type enumeration is used to indicate the type of data, including one or more of text type, number type, list type, range type, collection type, and mapping type.
3. A non-cooperative industrial data scheduling and storage method according to claim 1, characterized in that: The steps for querying the real-time data of a specified data item in a specified data source include: Get the specified data source or the number of bits agreed upon in the data source, or the specified data item tag or the number of bits agreed upon in the data item; Directly query the latest data of the corresponding category in the second area of volatile storage. If it cannot be found, generate a hash index based on the data source, data item tag and current time information to locate the corresponding data in non-volatile storage.
4. A non-cooperative industrial data scheduling and storage method according to claim 1, characterized in that: The steps of querying the data of a specified data item in a specified data source within a specified time period include: Get the specified data source or the number of bits agreed upon in the data source, or the specified data item tag or the number of bits agreed upon in the data item; Based on the number of time points in a specified time period, a hash index is generated for each query, and the hash index is used to locate data in non-volatile storage.
5. A non-cooperative industrial data scheduling and storage method according to claim 1, characterized in that: The steps of querying the time period that lasts when the specified data item in the specified data source meets the specified rule include: Get the specified data source or the number of bits agreed upon in the data source, or the specified data item tag or the number of bits agreed upon in the data item; Generate a hash index for every minute in the specified time period, use the generated hash index to locate a whole block of data, copy the data to the server memory, filter the data that meets the preset conditions of the queryer, and merge the time information of the data that meets the conditions to obtain the final result.
6. A non-cooperative industrial data scheduling and storage method according to claim 1, characterized in that: It also includes basic operations such as built-in addition, subtraction, multiplication, division, logical operations, mapping tables and string concatenation, supports secondary processing of data, and generates corresponding srcId and defId for each data generated by the secondary processing; the mapping table is a user-defined operation used to obtain output values based on input parameters.
7. A cloud platform for non-cooperative industrial data scheduling and storage, characterized in that: include: A data receiving module, used for receiving tree-structured industrial data with hierarchical labels; The tree structure is marked with data source information, and the data format is located through the unique mark corresponding to the data source. The parsing mark of each data item in the data format is mapped through a defId; The data preprocessing module is used to add the unprocessed raw data to the data queue in the first area of the volatile storage; take out a piece of raw data from the data queue, locate the srcId of the data source, and filter and classify according to the custom strategy of each data node in the raw data to obtain several groups of preprocessed data with the same data source and the same processing strategy, and each group of preprocessed data is stored in a continuous area in the second area of the volatile storage; A persistence module is used to generate a persistent data index for the data in the second area; the persistent data index is a hash index generated based on the data source srcId or the agreed preset number of digits therein, the data item tag mapping defId or the agreed preset number of digits therein, and the year to minute information of the timestamp corresponding to the data; and the same group of data with the persistent data index that meets the write conditions is written into the non-volatile storage as a whole.
8. A cloud platform for non-cooperative industrial data scheduling and storage according to claim 7, characterized in that: The network architecture of the multi-server cluster of the cloud platform performs data reception and preprocessing.
9. A computer system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, characterized in that: When the computer program / instructions are executed by a processor, the steps of a non-cooperative industrial data scheduling and storage method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of a non-cooperative industrial data scheduling and storage method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-level bucket hashing index method for searching mass data
CN101782922A
Implementation method and system of file system oriented to nonvolatile memory and medium
CN111221776A
Data loading method, computer equipment and storage medium
CN116644053A
Data storage method and system
CN117827850A
Self-adaptive multi-level mixed index data storage system and storage method
CN118152407A