A balanced management method for data collection in industrial manufacturing
By matching attribute splitting feature tree and data mapping pipeline, generating data acquisition target tree and collecting underlying data, recursive mapping and parallel transmission optimization, the problems of low data acquisition efficiency and high transmission pressure in the existing technology are solved, and efficient data acquisition and transmission are achieved.
Patent Information
- Application Number
- CN202411668448.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The prior art relies on the underlying data configuration in data collection, making it difficult to directly obtain data attributes at high semantic levels, resulting in high data transmission pressure and low acquisition efficiency.
By responding to the data acquisition request input by the user, the attribute split feature tree and the data mapping pipeline are matched, the data acquisition target tree is generated and the underlying data is collected, the target data is acquired recursively, the data results are divided into several data blocks, and parallel transmission equalization optimization is performed.
It reduces the data transmission pressure, improves data acquisition efficiency, and can directly obtain data attributes at high semantic levels.
Smart Images

Figure CN119166706B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data collection, and in particular to a balanced management method for data collection in industrial manufacturing. Background Art
[0002] In the current industrial manufacturing data collection, users input the data attributes to be collected, and the system retrieves and collects the corresponding underlying data from the database. However, if the attributes requested by the user are missing from the database, data collection will not be possible. Since industrial field data is complex and redundant, users' actual needs are often concentrated on data attributes with higher semantic levels, such as equipment performance analysis and production efficiency. However, each time these high-level data are collected, they still need to be configured and collected from the underlying data, which increases the transmission pressure and reduces the efficiency of data collection. This method cannot meet the needs of modern industry for efficient and intelligent data collection. Summary of the invention
[0003] The present application provides a balanced management method for data collection in the industrial manufacturing industry, which is used to solve the technical problems that the existing technology relies on the underlying data configuration in data collection, is difficult to directly obtain data attributes at a high semantic level, and has high data transmission pressure and low collection efficiency.
[0004] In view of the above problems, the present application provides a balanced management method for data collection in the industrial manufacturing industry.
[0005] The present application provides a balanced management method for data collection in industrial manufacturing, the method comprising:
[0006] In response to a data collection request input by a user, a target data attribute and a target collection time zone are obtained; according to the target data attribute, an attribute splitting feature tree and a data mapping pipeline are matched, the root node of the attribute splitting feature tree is the target data attribute, the N-level leaf node data of the attribute splitting feature tree is the mapping source data of the N-1-level leaf node data, and the data mapping pipeline is a mapper for realizing the N-level leaf node data to the N-1-level leaf node data; based on the target collection time zone, in combination with the industrial manufacturing database, the attribute splitting feature tree is cut to obtain a data collection target tree, the The bottom leaf nodes of the data collection target tree are data that can be directly collected by the industrial manufacturing database, and the bottom leaf nodes have no child nodes; according to the data collection target tree, data is collected on the bottom leaf nodes in the industrial manufacturing database to obtain the bottom leaf node data; according to the data mapping pipeline, the data collection target tree is recursively mapped based on the bottom leaf node data to obtain the target data collection results; the target data collection results are divided to obtain a number of data blocks; the parallel transmission balance optimization is performed on the several data blocks to obtain a data balanced transmission scheme to execute data collection.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0008] The present application obtains target data attributes and target collection time zone in response to a data collection request input by a user end; matches an attribute splitting feature tree and a data mapping pipeline according to the target data attributes, wherein the root node of the attribute splitting feature tree is the target data attribute, the N-level leaf node data of the attribute splitting feature tree is the mapping source data of the N-1-level leaf node data, and the data mapping pipeline is a mapper for realizing the conversion of the N-level leaf node data to the N-1-level leaf node data; based on the target collection time zone and in combination with an industrial manufacturing database, the attribute splitting feature tree is cut to obtain a data collection target tree, wherein the bottom leaf nodes of the data collection target tree are data that can be directly collected by the industrial manufacturing database, and the bottom leaf nodes have no child nodes; according to the data collection target tree, data is collected on the bottom leaf nodes in the industrial manufacturing database to obtain the bottom leaf node data; according to the data mapping pipeline, the data collection target tree is recursively mapped based on the bottom leaf node data to obtain the target data collection result; the target data collection result is divided to obtain a number of data blocks; parallel transmission balancing optimization is performed on a number of data blocks to obtain a data balancing transmission scheme to execute data collection. The present invention solves the technical problems that the prior art relies on the underlying data configuration in data collection, is difficult to directly obtain data attributes at a high semantic level, and has high data transmission pressure and low collection efficiency. The present invention responds to the data collection request input by the user, matches the attribute splitting feature tree with the data mapping pipeline, generates a data collection target tree and collects the underlying data, recursively maps to obtain the target data, divides the data results into several data blocks, and performs parallel transmission balanced optimization, thereby achieving the technical effect of reducing data transmission pressure and improving data collection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A flow chart of a balanced management method for data collection in the industrial manufacturing industry provided in an embodiment of the present application;
[0011] Figure 2 A flowchart of matching attribute splitting feature trees and data mapping pipelines in a balanced management method for data collection for industrial manufacturing provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] The present application provides a balanced management method for data collection for the industrial manufacturing industry, which is used to solve the technical problems that the existing technology relies on the underlying data configuration in data collection, it is difficult to directly obtain data attributes at a high semantic level, and there is a large data transmission pressure and low collection efficiency. By responding to the data collection request input by the user, matching the attribute splitting feature tree with the data mapping pipeline, generating a data collection target tree and collecting the underlying data, recursively mapping to obtain the target data, dividing the data results into several data blocks, and performing parallel transmission balanced optimization, the technical effect of reducing data transmission pressure and improving data collection efficiency is achieved.
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0014] It should be noted that any variations of the terms "include" and "have" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or inherent to these processes, methods, products or devices.
[0015] like Figure 1 As shown, the present application provides a balanced management method for data collection in industrial manufacturing, the method comprising:
[0016] Step S100: In response to a data collection request input by a user, target data attributes and a target collection time zone are obtained.
[0017] In an embodiment of the present application, when a user submits a data collection request, the request content is parsed to extract the target data attributes that the user needs to collect. These target data attributes represent the specific data type or information level that the user wants to obtain, such as the status of production equipment, production efficiency, etc. Next, the target collection time zone is extracted from the request, which defines the time range of the data that the user wants to collect. The target collection time zone can be understood as the timestamp interval of the underlying data, that is, the restricted range of the data that the user wants to collect in the time dimension.
[0018] Through the above process, the target data attributes and target collection time zone are obtained.
[0019] Step S200: According to the target data attribute, match the attribute splitting feature tree and the data mapping pipeline, the root node of the attribute splitting feature tree is the target data attribute, the N-level leaf node data of the attribute splitting feature tree is the mapping source data of the N-1-level leaf node data, and the data mapping pipeline is a mapper for realizing the N-level leaf node data to the N-1-level leaf node data.
[0020] In an embodiment of the present application, first, according to the target data attributes, the corresponding attribute splitting feature tree is matched from the preset data model. The root node of the feature tree represents the target data attribute that the user wants to obtain, such as the operating status or production efficiency of the equipment. Next, the data attribute is split down step by step to form multiple levels of leaf nodes until each leaf node represents the underlying data that can actually be collected. These N-level leaf node data depend on their upper-level nodes, that is, the N-1-level leaf node data, and the latter is the mapping source data of the former.
[0021] While constructing the feature tree, the corresponding data mapping pipeline is automatically matched. The data mapping pipeline is used to achieve the step-by-step mapping and conversion of leaf node data. Specifically, the data mapping pipeline maps the data of N-level leaf nodes to N-1-level nodes step by step through a series of mappers. The data at each level is converted into the data required by the node at the previous level through the mapper until high-level data that meets the user's target data attributes is generated.
[0022] In the whole process, the attribute splitting feature tree is responsible for gradually decomposing high-level data attributes into low-level collectible data, while the data mapping pipeline is responsible for realizing the conversion and mapping of each layer of data. The two are closely combined to ensure that the data is recursively mapped upward from the bottom-level original collected data, and finally generate the high-level data results required by the user.
[0023] By matching the attribute splitting feature tree and the data mapping pipeline, the user's target data attributes can be effectively mapped to the underlying data step by step, and then converted through the data pipeline to finally obtain the data collection results that meet the user's needs.
[0024] Further, such as Figure 2 As shown, in the method provided by the embodiment of the application, the matching attribute splitting feature tree and the data mapping pipeline also include:
[0025] Perform mapping source data analysis on the preset data attributes to obtain the first-level mapping source data attributes; traverse the non-bottom-level data attributes of the first-level mapping source data attributes to perform mapping source data analysis to obtain the second-level mapping source data attributes; until the M-level mapping source data attributes are all bottom-level data attributes, construct a preset attribute splitting feature tree according to the preset data attributes, the first-level mapping source data attributes, the second-level mapping source data attributes, and up to the M-level mapping source data attributes; train a preset data mapping pipeline according to the preset attribute splitting feature tree; associate the preset attribute splitting feature tree with the preset data mapping pipeline and the preset data attributes for storage.
[0026] In the embodiment of the present application, when performing mapping source data analysis on the preset data attribute, firstly, a set of generated record data of the preset data attribute is obtained, the trigger frequency of each source data attribute is counted, and the source data record attributes with a frequency higher than a threshold are extracted. Then, based on these frequencies, the source data record attributes are associated, and an association analysis is performed to obtain the first-level mapping source data attributes.
[0027] Then, the non-bottom-level data attributes in the first-level mapping source data attributes, that is, the high-level data attributes that have not been directly collected, are traversed, and the mapping source data analysis is continued. The above steps are repeated to obtain the second-level mapping source data attributes. The same analysis method is used to recursively process layer by layer to obtain deeper mapping source data attributes in turn. This recursive analysis continues until all M-level mapping source data attributes are extracted, and these data attributes are bottom-level data attributes, that is, data that can be directly collected from the industrial manufacturing database.
[0028] After completing the mapping source data analysis, a preset attribute splitting feature tree is constructed based on the mapping source data attributes at all levels. The root node of this feature tree is the preset data attribute, which is decomposed into different levels of mapping source data attributes step by step, and the final leaf node is the bottom data attribute. This feature tree reflects the layer-by-layer mapping relationship from high-level semantic data to the bottom-level collectible data.
[0029] Next, the feature tree is split based on the preset attributes, and the preset data mapping pipeline is trained. Through training, it is ensured that the leaf node data of each layer can be mapped to the node data of the previous layer step by step through the data mapper until it is finally mapped to the target data attribute. Finally, the preset attribute split feature tree is associated with the preset data mapping pipeline and the corresponding preset data attributes for storage. Through this associated storage, the matching feature tree and mapping pipeline can be quickly retrieved in the future to achieve efficient data collection and processing.
[0030] Furthermore, in the method provided in the embodiment of the application, performing mapping source data analysis on the preset data attributes to obtain the primary mapping source data attributes also includes:
[0031] Obtain a first generated record data set of preset data attributes, wherein the first generated record data set includes a source data record attribute set; count the source data record attribute trigger frequency set of the source data record attribute set; extract the source data record attributes whose source data record attribute trigger frequency set is greater than or equal to the record frequency threshold, and set them as frequency-associated source data record attributes; perform mapping association analysis on the frequency-associated source data record attributes based on the preset data attributes to obtain the first-level mapped source data attributes.
[0032] In the embodiment of the present application, data related to the preset data attribute is first extracted from the database of the industrial manufacturing industry to form a first generated record data set. This set includes multiple data records, including a source data record attribute set related to the target attribute, and these source data record attributes represent basic data in the underlying data that has a potential association with the preset attribute.
[0033] Next, the extracted source data record attribute set is counted to determine the occurrence frequency of each attribute in the data set to form a trigger frequency set. Then, according to the preset record frequency threshold, source data record attributes with a frequency greater than or equal to the threshold are screened from the trigger frequency set to obtain frequency-associated source data record attributes.
[0034] After obtaining the frequency-related source data record attributes, a mapping correlation analysis is performed on the frequency-related source data record attributes using a grey correlation analysis method based on preset data attributes to obtain primary mapping source data attributes.
[0035] Furthermore, in the method provided in the embodiment of the application, mapping and associating the attributes of the frequency-associated source data records are performed based on the preset data attributes to obtain the primary mapping source data attributes, and further includes:
[0036] According to the preset data attributes and the frequency-associated source data record attributes, a second generated record data set is collected; a preset data attribute feature value set is extracted from the second generated record data set, and is set as a reference data sequence after normalization; a frequency-associated source data record attribute feature value set is extracted from the second generated record data set, and is set as a comparison data sequence set after normalization; a grey correlation analysis is performed based on the reference data sequence and the comparison data sequence set to obtain the correlation degree of the frequency-associated source data record attributes; the frequency-associated source data record attributes whose correlation degree of the frequency-associated source data record attributes is greater than or equal to the correlation degree threshold are extracted, and are set as the first-level mapping source data attributes.
[0037] In the embodiment of the present application, firstly, according to the acquired preset data attributes and frequency-related source data record attributes, a second generated record data set is collected from the data source by means of database query, and this set includes data records related to the preset data attributes and frequency-related source data record attributes.
[0038] Then, a set of preset data attribute feature values is extracted from the second generated record data set. These feature values represent the numerical performance of the preset data attributes in different records. After extraction, these feature values are normalized to eliminate data dimension differences by using a minimum-maximum normalization algorithm. The normalized feature value set is set as the benchmark data sequence.
[0039] Then, a set of characteristic values of the frequency-related source data record attributes is extracted from the second generated record data set. These characteristic values reflect the specific performance of the frequency-related source data record attributes in the records. The extraction process includes traversing each record and obtaining the field values corresponding to the frequency-related source data record attributes. Data cleaning is performed during the extraction process to process missing values and outliers to ensure data integrity and reliability. After the extraction is completed, these characteristic values are also normalized to obtain a set of comparison data sequences on the same numerical scale as the benchmark data sequence.
[0040] Next, the grey correlation analysis is performed on the benchmark data sequence and the comparison data sequence set. First, the difference between the benchmark data sequence and each comparison data sequence is calculated, then the minimum difference and the maximum difference in the difference sequence are extracted, and finally the grey correlation calculation is performed to obtain the frequency correlation source data record attribute correlation. For the obtained frequency correlation source data record attribute correlation, the frequency correlation source data record attributes with a correlation greater than or equal to the preset correlation threshold, such as 0.7, are screened out, and these record data are set as the first-level mapping source data attributes.
[0041] Furthermore, in the method provided in the embodiment of the application, the feature tree is split according to the preset attribute, and the preset data mapping pipeline is trained, and further includes:
[0042] According to the preset attribute, the feature tree is split, and the preset data attribute, the first-level mapping source data attribute, the second-level mapping source data attribute, and the M-level mapping source data attribute are extracted; the preset data attribute is used as the supervisory attribute, the first-level mapping source data attribute is used as the input data attribute, and the first-level mapper is collected to construct a data set and train the first-level mapper; the first-level mapping source data attribute is traversed as the supervisory attribute, the corresponding second-level mapping source data attribute is used as the input data attribute, the second-level mapper is collected to construct a data set and train the second-level mapper; until the M-1-level mapping source data attribute is traversed as the supervisory attribute, the corresponding M-level mapping source data attribute is used as the input data attribute, the M-level mapper is collected to construct a data set and train the M-level mapper; the M-level mappers are connected in series in sequence until the second-level mapper and the first-level mapper to generate the preset data mapping pipeline.
[0043] In the embodiment of the present application, the feature tree is first split according to the preset attributes, and the mapping source data attributes of each level are extracted from the tree step by step using a tree structure traversal algorithm, such as breadth-first traversal. This includes the preset data attributes at the top level, and then the first-level mapping source data attributes, the second-level mapping source data attributes, and the M-level mapping source data attributes are extracted in sequence. The extraction process is completed through the indexing of the tree nodes, ensuring the recursive organization structure of the data from the high level to the low level.
[0044] Next, the preset data attributes are used as supervisory attributes and the primary mapping source data attributes are used as input data attributes to construct a data set for the primary mapper. The data set is collected by extracting the feature values of the primary mapping source data attributes for modeling. During the training process, a supervised learning method such as a multilayer perceptron is used to train the primary mapper so that it can learn how to generate the preset data attributes through the primary mapping source data attributes. The primary mapper is obtained through this process.
[0045] After completing the training of the first-level mapper, the output of the trained first-level mapper is used as the supervised attribute, and the second-level mapping source data attribute is used as the input attribute. The second-level mapper is constructed and trained through recursive supervised learning training. In this step, the same supervised learning method as the previous step is used to establish a second-level mapper model based on the characteristic values and structure of each level of attributes, so that it can accurately map the second-level mapping source data to the first-level mapping source data. Then continue to train each level of mapper layer by layer according to the above method until M-level mappers are obtained. The training data set of each level is composed of the output of the previous layer mapper as the supervised attribute, and the input data is the mapping source data attribute of this layer. For example, the M-1 level mapping source data attribute is used as the supervised attribute, and the corresponding M-level mapping source data attribute is used as the input data to train the M-level mapper. This recursive training method ensures that each level of mapper can effectively generate the output data of the previous level from the data of the lower level.
[0046] After completing the training of all mappers, use the model concatenation method to connect all mappers in series to form a multi-layer mapping structure. Specifically, the output of the M-level mapper is used as the input of the M-1-level mapper in sequence, and is passed upward in sequence until the first-level mapper is connected to the preset data attribute. Finally, all trained mappers are combined to form a complete preset data mapping pipeline.
[0047] Step S300: Based on the target collection time zone and in combination with the industrial manufacturing database, the attribute splitting feature tree is cut to obtain a data collection target tree, wherein the bottom leaf nodes of the data collection target tree are data that can be directly collected by the industrial manufacturing database, and the bottom leaf nodes have no child nodes.
[0048] In an embodiment of the present application, first, according to the target collection time zone, the specific time range for collecting data is determined. The target collection time zone refers to the time period or time point of interest in data collection, such as the operating time or downtime of a certain device. Next, in conjunction with the industrial manufacturing database, a database query operation is performed. A large amount of operating data, such as equipment status, production parameters, etc., is stored in the industrial manufacturing database, and these data are recorded by timestamps. Therefore, through the query, the relevant data that meets the target collection time zone is filtered out based on the timestamp.
[0049] After combining with the database, the attribute splitting feature tree is pruned using a pruning algorithm based on the retrieved results that match the time zone. The attribute splitting feature tree is a hierarchical structure tree that maps high-level target data attributes to underlying data attributes step by step. The purpose of the pruning process is to remove branch nodes that are not related to the target acquisition time zone and retain data branches related to the time zone. Specifically, the leaf nodes that match the time zone are found through a tree structure traversal algorithm, and then the CART pruning algorithm or the Alpha pruning algorithm is used to remove nodes that are not related to the time zone and retain those branches that meet the time zone conditions. The pruned tree is the data acquisition target tree, and only data nodes related to the target acquisition time zone are retained.
[0050] Through the cutting operation, a data collection target tree is generated, and its bottom leaf nodes are the data that can be directly collected from the industrial manufacturing database. The cut target tree is a simplified feature tree, and the bottom leaf nodes represent the basic data that can be directly collected during the industrial manufacturing process, including sensor data, equipment status, etc. In the generated data collection target tree, all bottom leaf nodes have no further child nodes, that is, these nodes represent the final data that can be directly collected. These data include sensor temperature, pressure value, equipment operation status, etc., which can be directly obtained through the industrial manufacturing database without further data processing or mapping.
[0051] Step S400: According to the data collection target tree, data is collected on the bottom leaf nodes in the industrial manufacturing database to obtain bottom leaf node data.
[0052] In an embodiment of the present application, according to the data collection target tree, when collecting data for the underlying leaf nodes in the industrial manufacturing database, first initiate a database query based on the underlying leaf nodes in the target tree to extract the corresponding data fields. These fields represent sensor data, equipment status, or other basic industrial data related to the underlying leaf nodes. Afterwards, through query statements or aggregation operations, match data records that meet the collection time zone from the database. After the query is completed, the extracted underlying data is formatted to ensure that the data of each leaf node is complete and corresponds to it. Through this process, the underlying leaf node data is obtained.
[0053] Step S500: According to the data mapping pipeline, based on the bottom leaf node data, the data acquisition target tree is recursively mapped to obtain a target data acquisition result.
[0054] In an embodiment of the present application, the bottom leaf node data is used as input, and the entire data acquisition target tree is recursively mapped step by step. In each level of mapping process, the data of the current node is used as input, and the bottom leaf node data is mapped to the upper leaf node data once through the mappers at each level of the data mapping pipeline, and finally recursively to the top level. After multiple layers of recursive mapping, the target data acquisition results are finally obtained at the highest level. These results are comprehensive information generated step by step through the data mapping pipeline based on the bottom leaf node data, representing higher-level data attributes that users are concerned about, such as production efficiency, equipment health status, or overall system performance.
[0055] Step S600: Divide the target data acquisition result to obtain a plurality of data blocks.
[0056] In an embodiment of the present application, when dividing the target data acquisition results, a strategy of dividing by data type is adopted to aggregate data of the same type together to form data blocks. For example, the device status data is divided into one data block, and the sensor data is divided into another data block. When performing the division, a fixed-length block method is used. According to the size requirements of each data block, the target data is divided into fixed lengths. For example, the size of each data block is set to 1MB, and then the target data is divided into units of 1MB in turn until all the data has been divided. Through this process, the target data acquisition results are divided into several independent data blocks.
[0057] Step S700: performing parallel transmission balance optimization on the plurality of data blocks to obtain a data balance transmission solution and execute data collection.
[0058] In the embodiment of the present application, when performing parallel transmission balanced optimization for several data blocks, the load status of the server node is first obtained, and the redundant transmission data volume of the server node is generated according to the resource usage of each node. Then, the first selected server node whose redundant transmission data volume is greater than or equal to the size of a single data block is extracted. Subsequently, parallel transmission optimization is performed based on these selected nodes to balance the transmission load of each node, and finally a data balanced transmission scheme is generated, through which data collection is performed.
[0059] Furthermore, in the method provided in the embodiment of the application, the parallel transmission of the plurality of data blocks is optimized to obtain a data balanced transmission scheme to perform data collection, and further includes:
[0060] Obtain the server node load status and generate the server node redundant transmission data volume; extract the first selected server node whose server node redundant transmission data volume is greater than or equal to the data block size; perform parallel transmission balancing optimization based on the first selected server node, obtain the data balancing transmission scheme and execute data collection.
[0061] In an embodiment of the present application, first, the load status of each server node is obtained through an integrated server monitoring tool. The load status includes key indicators such as CPU usage, memory occupancy, and network bandwidth usage. After obtaining the load status, the amount of redundant transmission data of each node is calculated based on the bandwidth and load of each node. The amount of redundant transmission data represents the amount of data that each node can carry additionally, which is calculated by bandwidth and network usage. When calculating, the amount of redundant transmission data of the server node is obtained by multiplying the total bandwidth by 1 minus the difference between the bandwidth usage.
[0062] Then, based on the calculated redundant transmission data volume, a filtering algorithm is used to select server nodes that process a single data block. The redundant transmission data volume of the selected server nodes is greater than or equal to the data block size, and these nodes are called first selected server nodes. For example, if the size of a data block is 5MB, nodes with redundant transmission data volume greater than or equal to 5MB are selected.
[0063] Next, parallel transmission balance optimization is performed based on the first selected server node. By performing stability analysis on each node, the server stability coefficient is obtained, and then compared with the preset threshold to select the second selected server node. Then, several data blocks are randomly deployed to these second selected server nodes to generate a data block parallel transmission plan. Finally, data collection is performed based on the generated data balanced transmission plan.
[0064] Furthermore, in the method provided in the embodiment of the application, parallel transmission balancing optimization is performed according to the first selected server node to obtain the data balancing transmission scheme to perform data collection, and further includes:
[0065] Traversing the first selected server node to perform stability analysis and obtain a server stability coefficient; extracting the first selected server node whose server stability coefficient is greater than or equal to a server stability coefficient threshold and setting it as a second selected server node; randomly deploying the plurality of data blocks to the second selected server node to obtain a data block parallel transmission scheme; counting the data block deployment quantity variance of the data block parallel transmission scheme; when the data block deployment quantity variance is less than the equilibrium variance, setting the data block parallel transmission scheme to the data equilibrium transmission scheme to perform data collection.
[0066] In the embodiment of the present application, all first selected server nodes are first traversed, and the performance of each node is evaluated by stability analysis to obtain the server stability coefficient. After the stability analysis of all nodes is completed, the nodes whose server stability coefficient is greater than or equal to the preset server stability coefficient threshold are screened out by comparison, and the screened out nodes are used as the second selected server nodes, wherein the preset server stability coefficient threshold is 1, which is set by technical experts.
[0067] Then, all data blocks are randomly allocated to these second selected server nodes through a random allocation algorithm. By randomly deploying data blocks, a data block parallel transmission scheme is generated. Then, the number of data blocks received by each node in the data block parallel transmission scheme is calculated, and the variance of these data is calculated. The calculated data block allocation variance is compared with the preset balanced variance threshold. If the variance is less than the balanced variance threshold, it means that the allocation of data blocks is sufficiently balanced, and the current transmission scheme is confirmed to be the final data balanced transmission scheme. If the variance value is large, random allocation is performed again until the calculated data block allocation variance is less than the preset balanced variance threshold, and the data balanced transmission scheme is determined. Finally, data collection is performed according to the determined data balanced transmission scheme.
[0068] Furthermore, in the method provided in the embodiment of the application, traversing the first selected server node to perform stability analysis to obtain a server stability coefficient also includes:
[0069] The proportion of data loss times and the proportion of transmission crash times of the first selected server node within three months are counted; when the proportion of data loss times is greater than or equal to the loss times proportion threshold, or / and the proportion of transmission crash times is greater than or equal to the transmission crash times proportion threshold, the server stability coefficient is equal to 0; when the proportion of data loss times is less than the loss times proportion threshold, and the proportion of transmission crash times is less than the transmission crash times proportion threshold, the server stability coefficient is equal to 1.
[0070] In an embodiment of the present application, the data loss of each node in the past three months is first extracted from the server log. The number of data loss refers to the number of times the server fails to successfully transmit a data packet due to network problems, hardware failures or other reasons when performing a data transmission task. By analyzing indicators such as ACK missing and timeout retransmission in the transmission log, the proportion of data loss times for each server is calculated. The proportion of data loss times is obtained by dividing the number of data loss times within three months by the total number of transmission times within three months.
[0071] Next, we count the number of transmission crashes for each node in the past three months, that is, the number of times the server was interrupted due to network disconnection or system crash when performing transmission tasks. The percentage of transmission crashes is obtained by dividing the number of transmission crashes in three months by the total number of transmissions in three months.
[0072] Next, the calculated data loss ratio and transmission crash ratio of each node are compared with the pre-set loss ratio threshold and transmission crash ratio threshold. If the data loss ratio of a node is greater than or equal to the loss ratio threshold, or the transmission crash ratio is greater than or equal to the transmission crash ratio threshold, as long as one of the above conditions occurs, the stability coefficient of the node is set to 0. If the data loss ratio is less than the threshold, and the transmission crash ratio is also less than the threshold, the stability coefficient of the node is set to 1.
[0073] By calculating the stability coefficient of each server node, the nodes suitable for participating in the data transmission task are further screened out, that is, the node with a stability coefficient of 1 is set as the second selected server node.
[0074] In the embodiments of the present application, in summary, the embodiments of the present application have at least the following technical effects:
[0075] The present application obtains target data attributes and target collection time zone in response to a data collection request input by a user end; matches an attribute splitting feature tree and a data mapping pipeline according to the target data attributes, wherein the root node of the attribute splitting feature tree is the target data attribute, the N-level leaf node data of the attribute splitting feature tree is the mapping source data of the N-1-level leaf node data, and the data mapping pipeline is a mapper for realizing the conversion of the N-level leaf node data to the N-1-level leaf node data; based on the target collection time zone and in combination with an industrial manufacturing database, the attribute splitting feature tree is cut to obtain a data collection target tree, wherein the bottom leaf nodes of the data collection target tree are data that can be directly collected by the industrial manufacturing database, and the bottom leaf nodes have no child nodes; according to the data collection target tree, data is collected on the bottom leaf nodes in the industrial manufacturing database to obtain the bottom leaf node data; according to the data mapping pipeline, the data collection target tree is recursively mapped based on the bottom leaf node data to obtain the target data collection result; the target data collection result is divided to obtain a number of data blocks; parallel transmission balancing optimization is performed on a number of data blocks to obtain a data balancing transmission scheme to execute data collection. The present invention solves the technical problems that the prior art relies on the underlying data configuration in data collection, is difficult to directly obtain data attributes at a high semantic level, and has high data transmission pressure and low collection efficiency. The present invention responds to the data collection request input by the user, matches the attribute splitting feature tree with the data mapping pipeline, generates a data collection target tree and collects the underlying data, recursively maps to obtain the target data, divides the data results into several data blocks, and performs parallel transmission balanced optimization, thereby achieving the technical effect of reducing data transmission pressure and improving data collection efficiency.
[0076] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. The processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0077] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0078] This specification and drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.
Claims
1. A balanced management method for data collection in industrial manufacturing, characterized in that: In response to a data collection request input by a user terminal, obtaining target data attributes and a target collection time zone; According to the target data attribute, the attribute splitting feature tree and the data mapping pipeline are matched, the root node of the attribute splitting feature tree is the target data attribute, the N-level leaf node data of the attribute splitting feature tree is the mapping source data of the N-1-level leaf node data, and the data mapping pipeline is a mapper for realizing the N-level leaf node data to the N-1-level leaf node data; Based on the target collection time zone and in combination with the industrial manufacturing database, the attribute splitting feature tree is cut to obtain a data collection target tree, wherein the bottom leaf nodes of the data collection target tree are data that can be directly collected by the industrial manufacturing database, and the bottom leaf nodes have no child nodes; According to the data collection target tree, data is collected on the bottom leaf nodes in the industrial manufacturing database to obtain bottom leaf node data; According to the data mapping pipeline, based on the bottom leaf node data, the data acquisition target tree is recursively mapped to obtain a target data acquisition result; Dividing the target data acquisition result to obtain a plurality of data blocks; Performing parallel transmission balance optimization on the plurality of data blocks to obtain a data balance transmission scheme and perform data collection; According to the target data attributes, matching attribute splitting feature tree and data mapping pipeline, including: Perform mapping source data analysis on preset data attributes to obtain primary mapping source data attributes; Traversing the non-bottom-layer data attributes of the first-level mapping source data attributes to perform mapping source data analysis to obtain second-level mapping source data attributes; All the M-level mapping source data attributes are bottom-level data attributes, and a preset attribute splitting feature tree is constructed according to the preset data attribute, the first-level mapping source data attribute, the second-level mapping source data attribute, and the M-level mapping source data attribute; Splitting the feature tree according to the preset attributes and training the preset data mapping pipeline; The preset attribute splitting feature tree is associated with the preset data mapping pipeline and the preset data attribute and stored.
2. The method according to claim 1, characterized in that Perform mapping source data analysis on the preset data attributes to obtain the first-level mapping source data attributes, including: Obtaining a first generated record data set of preset data attributes, wherein the first generated record data set includes a source data record attribute set; Counting a source data record attribute trigger frequency set of the source data record attribute set; Extracting source data record attributes whose source data record attribute trigger frequency set is greater than or equal to the record frequency threshold, and setting them as frequency-associated source data record attributes; Based on the preset data attributes, mapping association analysis is performed on the frequency-associated source data record attributes to obtain the primary mapping source data attributes.
3. The method according to claim 2, characterized in that Performing mapping association analysis on the frequency-associated source data record attribute based on the preset data attribute to obtain the primary mapping source data attribute includes: According to the preset data attribute and the frequency-associated source data record attribute, collecting a second generated record data set; Extracting a preset data attribute feature value set from the second generated record data set, and setting it as a reference data sequence after normalization processing; Extracting a frequency-related source data record attribute feature value set from the second generated record data set, and setting it as a comparison data sequence set after normalization processing; Performing grey correlation analysis on the reference data sequence and the comparison data sequence set to obtain the correlation degree of the attribute of the frequency-correlated source data record; The frequency-correlated source data record attributes whose correlation degree of the frequency-correlated source data record attributes is greater than or equal to a correlation degree threshold are extracted and set as the first-level mapping source data attributes.
4. The method according to claim 2, characterized in that Splitting the feature tree according to the preset attributes and training the preset data mapping pipeline include: Splitting the feature tree according to the preset attribute, extracting the preset data attribute, the first-level mapping source data attribute, the second-level mapping source data attribute until the M-level mapping source data attribute; Taking the preset data attribute as the supervision attribute and the primary mapping source data attribute as the input data attribute, collecting the primary mapper to construct the data set, and training the primary mapper; Traversing the primary mapping source data attributes as supervision attributes, taking the corresponding secondary mapping source data attributes as input data attributes, collecting a secondary mapper to construct a data set, and training the secondary mapper; Until the M-1 level mapping source data attributes are traversed as the supervision attributes, the corresponding M level mapping source data attributes are used as the input data attributes, the M level mappers are collected to construct the data sets, and the M level mappers are trained; The M-level mappers are sequentially connected in series until the second-level mapper and the first-level mapper to generate the preset data mapping pipeline.
5. The method according to claim 1, characterized in that Performing parallel transmission balance optimization on the plurality of data blocks to obtain a data balance transmission scheme and perform data collection, including: Obtain the server node load status and generate the server node redundant transmission data volume; Extracting a first selected server node whose redundant transmission data volume of the server node is greater than or equal to the data block size; Parallel transmission balance optimization is performed according to the first selected server node to obtain the data balance transmission solution and execute data collection.
6. The method according to claim 5, characterized in that Performing parallel transmission balancing optimization according to the first selected server node to obtain the data balancing transmission solution and perform data collection includes: Traversing the first selected server node to perform stability analysis and obtain a server stability coefficient; Extracting a first selected server node whose server stability coefficient is greater than or equal to a server stability coefficient threshold value, and setting it as a second selected server node; Randomly deploy the plurality of data blocks to the second selected server node to obtain a data block parallel transmission scheme; and count the data block deployment quantity variance of the data block parallel transmission scheme; When the data block deployment quantity variance is smaller than the balanced variance, the data block parallel transmission scheme is set to the data balanced transmission scheme to perform data collection.
7. The method according to claim 6, characterized in that Traversing the first selected server node to perform stability analysis to obtain a server stability coefficient includes: Counting the percentage of data loss and transmission crash within three months of the first selected server node; When the data loss times ratio is greater than or equal to the loss times ratio threshold, or / and the transmission crash times ratio is greater than or equal to the transmission crash times ratio threshold, the server stability coefficient is equal to 0; When the proportion of the number of data losses is less than the threshold of the number of data losses, and the proportion of the number of transmission crashes is less than the threshold of the number of transmission crashes, the server stability coefficient is equal to 1.
Citation Information
Patent Citations
Data query method and system based on Key-Value data blocks
CN105488043A
New energy aggregation data acquisition method and device based on edge calculation
CN118885540A