Data collection and processing method and system in industrial pilot platform tools

By obtaining data collection requirements in the industrial pilot platform tool, performing data preprocessing and classification structure processing and encryption protection, and finally storing the data, the problem of unreasonable data management is solved, efficient data collection and processing is achieved, and data accuracy and management capabilities are improved.

CN119272086BActive Publication Date: 2025-09-19SHENZHEN ZHENGZHEN METAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411212808.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-09-19
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

The existing technology has unreasonable data management in the data collection and processing process of industrial pilot platform tools, resulting in inaccurate data collection and processing, especially in the case of large data volumes, which is difficult to meet actual needs.

Method used

By obtaining data collection requirements, performing data preprocessing, identifying data features and class centers, analyzing the optimal attribute splitting points and spatial local deviation rates, building a data grid for data screening, performing classification structure processing and encryption protection, and finally storing data, target data is formed.

Benefits of technology

It improves the validity and practicality of data, ensures that data is consistent with actual needs, removes noise and outliers, discovers potential patterns, improves data access efficiency and management capabilities, and provides a reliable data source for subsequent analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272086B_ABST
    Figure CN119272086B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and discloses a data collection and processing method in an industrial pilot platform tool, comprising: collecting industrial pilot data to obtain collected data; preprocessing the collected data to obtain preprocessed data, and identifying the data class center of the preprocessed data; analyzing the optimal attribute splitting point of the preprocessed data to identify the spatial local deviation rate of the preprocessed data, sorting the preprocessed data to obtain distributed data, identifying spatial access points of the distributed data, constructing a data grid of the preprocessed data based on the spatial access points, using the data grid to perform data screening on the distributed data to obtain screened data; identifying the data structure of the screened data, performing classification structure processing on the screened data, and performing data encryption protection on the structure-processed data to obtain target data. The present invention can improve the rationality of data management in the industrial pilot platform tool.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data collection and processing method and system in an industrial pilot platform tool. Background Art

[0002] Industrial pilot platform tools play a key role in driving industrial innovation and development, and data collection and processing methods are crucial. Building an efficient data collection and processing system for use in industrial pilot platform tools can help companies accurately record, analyze, and manage all types of data during the pilot process, thereby improving the quality and efficiency of industrial pilots. At the same time, technical personnel can monitor changes in pilot parameters in real time and conduct in-depth data analysis, making timely adjustments to pilot plans to provide more accurate pilot services. Furthermore, the system facilitates information sharing and communication between different departments, between personnel in different technical positions, and between companies and external partners, thereby reducing pilot risks and improving team collaboration efficiency.

[0003] Currently, data collection and processing within industrial pilot platform tools typically relies on traditional data acquisition equipment and software to build a system with basic data storage and simple analysis capabilities to manage pilot data. However, this approach requires manual intervention for much of the data collection and processing. Due to the large amount of multi-source, heterogeneous data present in industrial pilot processes, users face inaccurate data collection and processing when the data volume is large, leading to inadequate data management within the industrial pilot platform tools. Summary of the Invention

[0004] The present invention provides a data collection and processing method and system in an industrial pilot platform tool, the main purpose of which is to improve the rationality of data management in the industrial pilot platform tool.

[0005] To achieve the above objectives, the present invention provides a data collection and processing method in an industrial pilot platform tool, comprising:

[0006] Obtaining data collection requirements of the industrial pilot platform tool, and collecting industrial pilot data according to the data collection requirements to obtain collected data;

[0007] Performing data preprocessing on the collected data to obtain preprocessed data, identifying data features of the preprocessed data, calculating feature averages of the data features, and determining a data class center of the preprocessed data based on the feature averages;

[0008] Analyzing the optimal attribute splitting point of the preprocessed data using the data class center, identifying the spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point, sorting the preprocessed data based on the spatial local deviation rate to obtain distributed data, identifying spatial access points of the distributed data, constructing a data grid of the preprocessed data based on the spatial access points, and filtering the distributed data using the data grid to obtain filtered data;

[0009] Identify the data structure of the screening data, perform classification structure processing on the screening data based on the data structure to obtain structure-processed data, perform data encryption protection on the structure-processed data to obtain secure data, perform data storage on the secure data to obtain target data.

[0010] Optionally, calculating the feature average of the data feature includes:

[0011] Construct a data set corresponding to the data feature, and based on the data set, calculate the feature average of the data feature using the following formula:

[0012]

[0013] Among them, α represents the feature mean, N represents the number of data sets, and x uy Represents the u-th data feature of the y-th data set.

[0014] Optionally, determining the data class center of the preprocessed data according to the feature average value includes:

[0015] determining an initial class center of the preprocessed data based on the feature average;

[0016] Performing data point allocation on the initial class center to obtain an allocated class center;

[0017] Updating the data of the distribution class center to obtain the target data class center;

[0018] Performing a center evaluation on the target data class center to obtain an evaluation result;

[0019] When the evaluation result meets the preset result, the target data class center is used as the data class center of the preprocessed data.

[0020] Optionally, analyzing the optimal attribute splitting point of the pre-processed data by using the data class center includes:

[0021] Determining attribute evaluation indicators of the preprocessed data based on the data class center;

[0022] Calculate the index value of each attribute of the preprocessed data under different values ​​according to the attribute evaluation index;

[0023] Based on the indicator value, an optimal attribute splitting point of the preprocessed data is determined.

[0024] Optionally, identifying the spatial local deviation rate of the pre-processed data based on the optimal attribute splitting point includes:

[0025] Based on the optimal attribute splitting point, calculating the similarity of adjacent data points in the preprocessed data;

[0026] Based on the similarity, the spatial local deviation rate of the preprocessed data is calculated using the following formula:

[0027]

[0028] Among them, SLDF(d i ) represents the spatial local deviation rate, d j represents the jth data point in the preprocessed data, d i represents the i-th data point in the preprocessed data, N(d i ) represents the data point d i The set of neighborhood data points, w ij represents the similarity between data points i and j, m represents the number of spatial dimensions of preprocessed data, x ik represents the eigenvalue of data point i in the kth spatial dimension, x jk Represents the eigenvalue of data point j in the kth spatial dimension.

[0029] Optionally, performing classification structure processing on the screening data based on the data structure to obtain structure-processed data includes:

[0030] Based on the data structure, the screening data is divided into structured data and unstructured data;

[0031] Constructing a relational data table of the structural data;

[0032] Performing data integration on the structured data using the relational data table to obtain integrated data;

[0033] performing classification processing on the unstructured data to obtain classified processing data;

[0034] The integrated data and the classified processed data are merged to obtain structure processed data.

[0035] Optionally, the classifying the unstructured data to obtain the classified data includes:

[0036] Identifying text data, image data, and audio data in the unstructured data;

[0037] Performing natural language processing on the text data to obtain processed text data;

[0038] performing image processing on the image data to obtain image processed data;

[0039] Performing audio processing on the audio data to obtain audio processed data;

[0040] The processed text data, the image processed data and the audio processed data are fused to obtain classified processed data.

[0041] Optionally, performing data encryption protection on the structure processing data to obtain secure data includes:

[0042] identifying data sensitivity of data processed by the structure;

[0043] Based on the data sensitivity, user definition is performed on the access role corresponding to the structure processing data to obtain a defined user;

[0044] querying independent attributes of said structure-processed data;

[0045] Associating the independent attribute with the defined user role to obtain data-role information;

[0046] Setting permissions on the stored data based on the data-role information to obtain permission data;

[0047] The authority data is encrypted to obtain secure data.

[0048] Optionally, encrypting the authority data to obtain secure data includes:

[0049] Calculating a data characteristic value of the permission data to obtain a permission data characteristic value;

[0050] Based on the characteristic value of the permission data, the permission data is encrypted using the following formula to obtain secure data:

[0051]

[0052] Among them, E represents security data, XOR(.) represents exclusive OR operation, T e represents the permission data feature value of the e-th permission data, M represents the feature dimension of the permission data, Represents the intermediate summation result after linear processing of all feature dimensions of all permission data.

[0053] The embodiment of the present invention obtains target data by storing the security data, which can improve data access efficiency and management capabilities, and provide a reliable data source for subsequent data analysis and application. The security data can be stored by utilizing the data repository in the industrial pilot platform tool.

[0054] In order to solve the above problems, the present invention also provides a data acquisition and processing system in an industrial pilot platform tool, the system comprising:

[0055] A data collection module is used to obtain data collection requirements of the industrial pilot platform tool, so as to collect industrial pilot data according to the data collection requirements and obtain collected data;

[0056] a data class center identification module, configured to perform data preprocessing on the collected data to obtain preprocessed data, identify data features of the preprocessed data, calculate feature averages of the data features, and determine the data class center of the preprocessed data based on the feature averages;

[0057] a data screening module, configured to analyze an optimal attribute splitting point of the preprocessed data using the data class center, identify a spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point, sort the preprocessed data based on the spatial local deviation rate to obtain distributed data, identify spatial access points of the distributed data, construct a data grid of the preprocessed data based on the spatial access points, and screen the distributed data using the data grid to obtain screened data;

[0058] A data storage module is used to identify the data structure of the screening data, classify and structure the screening data based on the data structure to obtain structure-processed data, encrypt and protect the structure-processed data to obtain secure data, and store the secure data to obtain target data.

[0059] The present invention is by obtaining the data collection demand of industrial pilot platform tool, with according to the data collection demand, collect industrial pilot data, obtain collected data and can clarify the target and direction of data collection, avoid the blind collection of data, improve the validity and practicality of data, ensure that the data collected is consistent with the actual demand of industrial pilot platform, and then provide original data basis for subsequent data processing and analysis;And data preprocessing is carried out to the collected data, obtain preprocessed data and can remove the interference factors such as noise, abnormal values, improve the quality of data, in industrial pilot data, there may be various errors and irregular data, and data can be made more accurate and reliable by preprocessing;Further, the present invention is by utilizing the data class center, and the optimal attribute splitting point of analyzing the preprocessed data can facilitate data to be classified according to specific attributes, contributes to the potential pattern and law in the data, and can provide more fine-grained information for subsequent data analysis and processing;Further, the present invention is by carrying out data storage to the security data, obtains target data and can improve the access efficiency and management ability of data, and provides reliable data source for subsequent data analysis and application. Therefore, the present invention can promote the rationality of data management in industrial pilot platform tool. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 A flow chart of a data collection and processing method in an industrial pilot platform tool provided by one embodiment of the present invention;

[0061] Figure 2 A functional module diagram of a data acquisition and processing system in an industrial pilot platform tool provided by one embodiment of the present invention;

[0062] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0063] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0064] The embodiment of the present application provides a data collection and processing method in an industrial pilot platform tool. The execution subject of the data collection and processing method in the industrial pilot platform tool includes but is not limited to at least one of the electronic devices such as the server, the terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the data collection and processing method in the industrial pilot platform tool can be executed by software or hardware installed on the terminal device or the server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.

[0065] Reference Figure 1 FIG. 1 is a flow chart of a data collection and processing method in an industrial pilot platform tool provided by an embodiment of the present invention. In this embodiment, the data collection and processing method in the industrial pilot platform tool includes:

[0066] S1. Obtaining data collection requirements of the industrial pilot platform tool, and collecting industrial pilot data according to the data collection requirements to obtain collected data.

[0067] The embodiment of the present invention obtains the data collection requirements of the industrial pilot platform tool to collect industrial pilot data according to the data collection requirements. The collected data can clarify the goal and direction of data collection, avoid blind data collection, improve the validity and practicality of the data, ensure that the collected data is consistent with the actual needs of the industrial pilot platform, and provide the original data basis for subsequent data processing and analysis.

[0068] Among them, the industrial pilot platform tool refers to a comprehensive platform and related tool set built in the process of industrial development to realize the transition of scientific and technological achievements from laboratory to industrialization. The data collection demand refers to the data generated in the industrial pilot platform tool, such as physical parameters, chemical parameters, electrical parameters and other data requirements.

[0069] Optionally, the data collection requirements can be obtained by obtaining the data requirements of relevant staff of the industrial pilot platform, such as engineers, researchers and managers, to understand their specific data requirements during the pilot process, such as which parameters of data are required, the accuracy requirements of the data, the frequency of data collection, etc. The industrial pilot data can be collected through data analysis tools such as Excel.

[0070] S2. Perform data preprocessing on the collected data to obtain preprocessed data, identify data features of the preprocessed data, calculate feature averages of the data features, and determine a data class center of the preprocessed data based on the feature averages.

[0071] The embodiment of the present invention performs data preprocessing on the collected data to obtain preprocessed data, which can remove interference factors such as noise and outliers and improve the quality of the data. In industrial pilot data, there may be various errors and non-standard data. Preprocessing can make the data more accurate and reliable.

[0072] The pre-processed data can be obtained by performing data cleaning, deduplication and outlier processing on the collected data.

[0073] Furthermore, the embodiment of the present invention can help users understand the distribution and characteristics of data by identifying the data features of the pre-processed data, thereby providing a basis for subsequent data analysis and processing.

[0074] Optionally, the data features of the pre-processed data may be obtained by identifying attribute values ​​of data tags of the pre-processed data.

[0075] Among them, the data characteristics refer to the unique attributes of the data in the industrial pilot platform tool, such as numerical type (such as temperature, pressure, size, etc.), category type (such as product type, equipment status, etc.) and other related characteristics.

[0076] Furthermore, the embodiment of the present invention can quickly understand the overall central position of the group of data features by calculating the feature average of the data features, thereby improving data processing efficiency.

[0077] As an embodiment of the present invention, the calculating the feature average of the data feature includes: constructing a data set corresponding to the data feature, and calculating the feature average of the data feature based on the data set using the following formula:

[0078]

[0079] Among them, α represents the feature mean, N represents the number of data sets, and x uy Represents the u-th data feature of the y-th data set.

[0080] Optionally, the data set can be constructed into different sets based on the attributes corresponding to different data features, such as a numerical set, a categorical set, etc.

[0081] Furthermore, the embodiment of the present invention can perform preliminary classification and summary of the data by determining the data class center of the preprocessed data based on the feature average value, providing a basis for further analysis.

[0082] Among them, the data class center refers to data with special performance in a certain data feature. For example, there are three types of product type features: A, B, and C. Through statistics, it is found that type A appears the most times in the data set, so type A can be used as the class center representative of the product type feature.

[0083] As an embodiment of the present invention, determining the data class center of the preprocessed data based on the feature average value includes: determining the initial class center of the preprocessed data based on the feature average value, allocating data points to the initial class center to obtain an allocated class center, updating data on the allocated class center to obtain a target data class center, performing center evaluation on the target data class center to obtain an evaluation result, and when the evaluation result meets a preset result, using the target data class center as the data class center of the preprocessed data.

[0084] Optionally, the initial class center of the preprocessed data can be obtained by selecting the data point closest to the feature average value, the assigned class center can be obtained by calculating the distance between each data point and its respective class center and assigning it to the class to which the nearest class center belongs, the target data class center can be obtained by calculating the average value of each data point in the assigned class center on each feature as the new class center, and the center evaluation of the target data class center can be performed by calculating the intra-class distance and inter-class distance of the target data class center to evaluate the quality of the class center, wherein the smaller the intra-class distance and the larger the inter-class distance, the better the effect of the class center. The preset result can be set to that the intra-class distance is not greater than 0.2 and the inter-class distance is not less than 0.8, and can also be set according to the actual application scenario.

[0085] S3. Analyze the optimal attribute splitting point of the preprocessed data using the data class center; identify the spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point; sort the preprocessed data based on the spatial local deviation rate to obtain distributed data; identify the spatial access points of the distributed data; construct a data grid of the preprocessed data based on the spatial access points; and use the data grid to filter the distributed data to obtain filtered data.

[0086] The embodiment of the present invention utilizes the data class center to analyze the optimal attribute splitting point of the preprocessed data, which can facilitate the classification of data according to specific attributes, help discover potential patterns and laws in the data, and provide more fine-grained information for subsequent data analysis and processing.

[0087] The optimal attribute splitting point refers to a point constructed when classifying data and is used to divide the data set.

[0088] As an embodiment of the present invention, the use of the data class center to analyze the optimal attribute splitting point of the preprocessed data includes: determining the attribute evaluation index of the preprocessed data based on the data class center, calculating the index value of each attribute of the preprocessed data under different values ​​according to the attribute evaluation index, and determining the optimal attribute splitting point of the preprocessed data based on the index value.

[0089] The attribute evaluation index is a criterion used to measure the effectiveness of each attribute (feature) in the preprocessed data for data partitioning. Common attribute evaluation metrics include information gain and the Gini index. The index value is the specific numerical value calculated for each attribute in the preprocessed data under different values ​​based on the selected attribute evaluation index.

[0090] Optionally, the attribute evaluation index of the pre-processed data based on the data class center can be determined by analyzing the characteristics of the data class center and the nature and needs of the pre-processed data. If the data class center is relatively clear and the classification of data is more important, you can consider selecting information gain as the attribute evaluation index, because information gain can better measure the contribution of attributes to classification. If you are more concerned about the purity of the data and the balance of the division, you can choose the Gini index. For example, in the industrial pilot platform tool, if the pre-processed data is about the operating parameters of different equipment, and you want to classify the equipment status according to these parameters, you can choose information gain. Information gain is used as an attribute evaluation indicator to determine which parameters are most helpful in distinguishing different device states; according to the attribute evaluation indicator, the index value of each attribute of the preprocessed data at different values ​​can be calculated by dividing the preprocessed data into two parts, calculating the information entropy or Gini index of each part respectively, and then calculating the index value according to the information entropy or Gini index before and after the division; based on the index value, the optimal attribute splitting point of the preprocessed data is determined by comparing the index values ​​of all attributes at different values, and selecting the attribute and value with the optimal index value (such as the maximum information gain or the minimum Gini index) as the optimal attribute splitting point.

[0091] Furthermore, the embodiment of the present invention can discover abnormal points or outliers in the data by identifying the spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point. These outliers may represent errors, abnormal conditions or potentially valuable information in the data. By identifying the spatial local deviation rate, the distribution and characteristics of the data can be better understood.

[0092] The spatial local deviation rate is an indicator used to measure the degree of deviation of each data point in the preprocessed data in the subspace to which it belongs.

[0093] As an embodiment of the present invention, identifying the spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point includes: calculating the similarity of adjacent data points in the preprocessed data based on the optimal attribute splitting point, and calculating the spatial local deviation rate of the preprocessed data based on the similarity using the following formula:

[0094]

[0095] Among them, SLDF(d i ) represents the spatial local deviation rate, d j represents the jth data point in the preprocessed data, d i represents the i-th data point in the preprocessed data, N(d i ) represents the data point d i The set of neighborhood data points, w ij represents the similarity between data points i and j, m represents the number of spatial dimensions of preprocessed data, x ik represents the eigenvalue of data point i in the kth spatial dimension, x jk Represents the eigenvalue of data point j in the kth spatial dimension.

[0096] Optionally, the similarity of adjacent data points in the preprocessed data is calculated based on the optimal attribute splitting point by using the optimal attribute splitting point to determine attribute values ​​in the preprocessed data, and calculating attribute weights of the attribute values, calculating the attribute weights, and determining the similarity.

[0097] The higher the weight, the higher the similarity.

[0098] In the embodiment of the present invention, by sorting the preprocessed data based on the spatial local deviation rate, the obtained distributed data can more intuitively display the distribution of the data and the differences in the degree of deviation. This helps to further analyze the characteristics of the data and identify the data areas that require special attention.

[0099] Optionally, the data sorting of the preprocessed data based on the spatial local deviation rate can be performed by integrating the local deviation rates of the data points in the preprocessed data in each feature dimension to obtain a comprehensive spatial local deviation rate. For example, the average or weighted average of the local deviation rates in each dimension is taken as the comprehensive deviation rate, and the preprocessed data is sorted according to the comprehensive deviation rate. The sorting can be from small to large or from large to small, depending on the analysis requirements.

[0100] The embodiment of the present invention can determine key locations or nodes that can be accessed and operated in the distributed data by identifying the spatial access points of the distributed data.

[0101] Among them, the spatial access point refers to the position of the key data point determined from the distributed data according to specific rules. The key data points can be selected from the distributed data as spatial access points according to certain rules, such as selecting a specific proportion of data points with a larger or smaller local deviation rate, or selecting points that are representative in the data distribution.

[0102] Furthermore, by constructing a data grid of the pre-processed data based on the spatial access points, embodiments of the present invention can help users better organize and manage data, making the spatial distribution of data more clear. Furthermore, the data grid can be used to quickly locate and query data, thereby improving data processing efficiency.

[0103] Optionally, the data grid of the pre-processed data is constructed based on the spatial access points by arbitrarily selecting a point that has not been visited among the spatial access points, using a set dimensional interval distance value as the grid basic unit size, and constructing the data grid of the pre-processed data based on the grid basic unit size.

[0104] Furthermore, the embodiment of the present invention performs data screening on the distributed data to obtain screened data, thereby removing unnecessary data or selecting a specific data subset, thereby extracting valuable data.

[0105] Optionally, the data can be filtered by setting data filtering conditions according to specific analysis requirements, such as selecting data in a specific grid unit, or selecting data points that meet a certain deviation rate range for filtering.

[0106] S4. Identify the data structure of the screening data, perform classification and structural processing on the screening data based on the data structure to obtain structural processing data, perform data encryption and protection on the structural processing data to obtain secure data, perform data storage on the secure data to obtain target data.

[0107] The embodiment of the present invention can identify the data structure by identifying the data structure of the screening data, thereby better understanding the organization and characteristics of the data, and being more conducive to data analysis.

[0108] The term "data structure" refers to the way computers store and organize data. It refers to a collection of data elements that have one or more specific relationships with each other. Data structure includes both the logical and physical structures of data.

[0109] Optionally, the data structure of the screened data can be determined by analyzing the characteristics of the data: observing the element types of the data, the relationship between the elements, the organization of the data, and other characteristics to determine the data structure type.

[0110] Furthermore, embodiments of the present invention, by performing classification and structural processing on the filtered data based on the data structure, obtain structured data that can organize and store the data according to a specific structure, making data management more orderly and efficient. This facilitates operations such as querying, updating, and deleting data, improving the speed and accuracy of data processing.

[0111] As an embodiment of the present invention, the classified structural processing is performed on the filtered data based on the data structure to obtain structural processing data, including: based on the data structure, the filtered data is divided into structural data and unstructured data, a relational data table of the structural data is constructed, the structural data is integrated using the relational data table to obtain integrated data, the unstructured data is classified and processed to obtain classified processing data, and the integrated data and the classified processing data are merged to obtain structural processing data.

[0112] Structured data refers to data with a well-defined structure and format, typically presented in a table format, with each column representing a specific attribute and each row representing a data record. This can be seen in databases and spreadsheets. Unstructured data refers to data without a fixed structure or format, making it difficult to represent using traditional tables, such as images, audio, and video. A data relationship table is a tabular representation used to describe the relationships between individual data elements in structured data. For example, in the industrial pilot platform tool, if an electronic product with model A was tested on August 1, 2024, at a temperature of 25°C, a voltage of 5V, and a current of 2A, this record would be reflected as a row in the data relationship table. If the product is tested again on August 2, new data would be added to the table. Different product models, such as model B, would also have their own test records in the table. This data relationship table clearly shows the relationships between various performance parameters of different product models at different points in time.

[0113] Optionally, the relational data table of the structural data can be constructed by using a spreadsheet tool such as Excel, and obtaining the integrated data can be achieved by normalizing the structural data.

[0114] Furthermore, as an optional embodiment of the present invention, the classification processing of the unstructured data to obtain classified processed data includes: identifying text data, image data and audio data in the unstructured data, performing natural language processing on the text data to obtain processed text data, performing image processing on the image data to obtain image processed data, performing audio processing on the audio data to obtain audio processed data, and fusing the processed text data, the image processed data and the audio processed data to obtain classified processed data.

[0115] Among them, the text data refers to information expressed in text form, such as books, articles, web pages, etc., the image data refers to information expressed in image form, such as photos, graphics, charts, etc., the audio data refers to information expressed in sound form, such as music, voice, sound effects, etc., and the data fusion refers to the fusion of data represented by multiple structures into a comprehensive data representation, which can be achieved by inputting data values ​​into the same relational database.

[0116] Optionally, the process of performing natural language processing on the text data to obtain processed text data is: dividing the text data into participles, performing text translation on the participles to obtain processed text data, performing image processing on the image data to obtain image processing data, performing image processing on the image data using a deep learning model, performing audio processing on the audio data to obtain audio processing data, performing audio feature extraction on the audio data using a deep learning model to obtain audio features, and performing text translation on the audio features.

[0117] Furthermore, the embodiment of the present invention encrypts the structured processing data to obtain secure data, thereby protecting the security and privacy of the data and preventing the data from being illegally accessed and tampered with.

[0118] As an embodiment of the present invention, the data encryption protection of the structure processing data to obtain secure data includes: identifying the data sensitivity of the structure processing data, based on the data sensitivity, user-defining the access role corresponding to the structure processing data to obtain a defined user, querying the independent attributes of the structure processing data, associating the independent attributes with the defined user to obtain data-role information, setting permissions for the stored data based on the data-role information to obtain permission data, and encrypting the permission data to obtain secure data.

[0119] Among them, the data sensitivity refers to the importance and sensitivity of the data to the organization or individual. For example, in the industrial pilot process, if cost accounting, investment evaluation and other links are involved, financial information will be generated. For example, the pilot project budget, actual expenditure details, expected returns and other data of a new product. These financial information are crucial to the enterprise and are highly sensitive data. The independent attribute refers to a characteristic of the data, that is, it does not depend on other data or external factors, and can be determined or calculated independently of other data. For example, in the industrial pilot platform for electronic products, the product model is an independent attribute. The product model is usually determined by the manufacturer during the product design stage, and it does not depend on other performance parameters or test results of the product.

[0120] Alternatively, data sensitivity can be obtained by classifying the sensitivity of the structured processing data, such as

[0121] In a material industry pilot platform, the following types of data are stored: physical property data such as material color and hardness can be divided into one category, indicating low sensitivity; some more conventional parts of material production cost and production process parameters can be divided into two categories, indicating general sensitivity; performance test results of materials under specific environments are divided into three categories, indicating medium sensitivity; core formula data of materials are in the fourth category, indicating relatively sensitive; special material data under development for specific high-end fields are in the fifth category, indicating high sensitivity. The user definition can be defined by defining the data access administrator corresponding to the structure processing data, for example, low-sensitivity data can be accessed by market research and sales personnel; the permission setting of the stored data based on the data-role information to obtain permission data can be achieved by setting the access level of the access role.

[0122] Furthermore, as an optional embodiment of the present invention, encrypting the permission data to obtain secure data includes: calculating a data characteristic value of the permission data to obtain a permission data characteristic value, and encrypting the permission data based on the permission data characteristic value using the following formula to obtain secure data:

[0123]

[0124] Among them, E represents security data, XOR(.) represents exclusive OR operation, T e represents the permission data feature value of the e-th permission data, M represents the feature dimension of the permission data, Represents the intermediate summation result after linear processing of all feature dimensions of all permission data.

[0125] The embodiment of the present invention obtains target data by storing the security data, which can improve data access efficiency and management capabilities, and provide a reliable data source for subsequent data analysis and application. The security data can be stored by utilizing the data repository in the industrial pilot platform tool.

[0126] like Figure 2 1 is a functional module diagram of a data acquisition and processing system in an industrial pilot platform tool provided by an embodiment of the present invention.

[0127] The data acquisition and processing system 200 in the industrial pilot platform tool described in the present invention can be installed in an electronic device. Depending on the functionality implemented, the data acquisition and processing system 200 in the industrial pilot platform tool can include a data acquisition module 201, a data class center identification module 202, a data screening module 203, and a data storage module 204. A module, also referred to as a unit, refers to a series of computer program segments that can be executed by an electronic device processor and perform a fixed function, and is stored in the electronic device's memory.

[0128] In this embodiment, the functions of each module / unit are as follows:

[0129] The data collection module 201 is used to obtain the data collection requirements of the industrial pilot platform tool, and to collect industrial pilot data according to the data collection requirements to obtain collected data;

[0130] The data class center identification module 202 is used to perform data preprocessing on the collected data to obtain preprocessed data, identify data features of the preprocessed data, calculate feature averages of the data features, and determine the data class center of the preprocessed data based on the feature averages;

[0131] The data screening module 203 is configured to analyze the optimal attribute splitting point of the pre-processed data using the data class center, identify the spatial local deviation rate of the pre-processed data based on the optimal attribute splitting point, sort the pre-processed data based on the spatial local deviation rate to obtain distributed data, identify the spatial access points of the distributed data, construct a data grid of the pre-processed data based on the spatial access points, and use the data grid to screen the distributed data to obtain screened data;

[0132] The data storage module 204 is used to identify the data structure of the screening data, perform classification structure processing on the screening data based on the data structure to obtain structure processing data, perform data encryption protection on the structure processing data to obtain security data, and perform data storage on the security data to obtain target data.

[0133] In detail, the modules described in the data collection and processing system 200 in the industrial pilot platform tool described in the embodiment of the present invention adopt the same technical means as the data collection and processing method in the industrial pilot platform tool described in the accompanying drawings when used, and can produce the same technical effects, which will not be repeated here.

[0134] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.

[0135] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0136] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A data collection and processing method in an industrial pilot platform tool, characterized in that: The method comprises: Obtaining data collection requirements of the industrial pilot platform tool, and collecting industrial pilot data according to the data collection requirements to obtain collected data; Performing data preprocessing on the collected data to obtain preprocessed data, identifying data features of the preprocessed data, and calculating feature averages of the data features, and determining a data class center of the preprocessed data based on the feature averages; Utilizing the data class center, analyzing the optimal attribute splitting point of the preprocessed data, identifying the spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point, sorting the preprocessed data based on the spatial local deviation rate to obtain distributed data, identifying the spatial access points of the distributed data, constructing a data grid of the preprocessed data based on the spatial access points, and filtering the distributed data using the data grid to obtain filtered data. The splitting point based on the optimal attribute includes: Based on the optimal attribute splitting point, calculating the similarity of adjacent data points in the preprocessed data; Based on the similarity, the spatial local deviation rate of the preprocessed data is calculated using the following formula: in, represents the spatial local deviation rate, represents the jth data point in the preprocessed data, represents the i-th data point in the preprocessed data, Represents a data point The set of neighborhood data points, Represents the similarity between data points i and j, m represents the number of spatial dimensions of preprocessed data, represents the eigenvalue of data point i in the kth spatial dimension, Represents the eigenvalue of data point j in the kth spatial dimension; Identifying a data structure of the screening data, performing classification and structural processing on the screening data based on the data structure to obtain structurally processed data, performing data encryption and protection on the structurally processed data to obtain secure data, and performing data storage on the secure data to obtain target data, wherein performing data encryption and protection on the structurally processed data to obtain secure data includes: identifying data sensitivity of data processed by the structure; Based on the data sensitivity, user definition is performed on the access role corresponding to the structure processing data to obtain a defined user; querying independent attributes of said structure-processed data; Associating the independent attribute with the defined user role to obtain data-role information; Setting permissions on the stored data based on the data-role information to obtain permission data; Calculating a data characteristic value of the permission data to obtain a permission data characteristic value; Based on the characteristic value of the authority data, the authority data is encrypted using the following formula to obtain security data, including: in, Indicates safety data, represents the exclusive OR operation, represents the permission data feature value of the e-th permission data, M represents the feature dimension of the permission data, Represents the intermediate summation result after linear processing of all feature dimensions of all permission data.

2. The data collection and processing method in the industrial pilot platform tool according to claim 1, characterized in that: The identifying data features of the pre-processed data and calculating feature averages of the data features includes: constructing a data set of the pre-processed data according to the data characteristics; classifying the pre-processed data using the data set; The following formula is used to calculate the average value of the numerical features corresponding to the pre-processed data in the classified data set: in, represents the feature mean, N represents the number of data sets, Represents the u-th eigenvalue of the y-th data set.

3. The data collection and processing method in the industrial pilot platform tool according to claim 1, characterized in that: The step of determining the data class center of the preprocessed data according to the feature average value further includes: determining an initial class center of the preprocessed data based on the feature average; Performing data point allocation on the initial class center to obtain an allocated class center; Updating the data of the distribution class center to obtain the target data class center; Performing a center evaluation on the target data class center to obtain an evaluation result; When the evaluation result meets the preset result, the target data class center is used as the data class center of the preprocessed data.

4. The data collection and processing method in an industrial pilot platform tool according to claim 1, characterized in that: The analyzing the optimal attribute splitting point of the pre-processed data by using the data class center includes: Determining attribute evaluation indicators of the preprocessed data based on the data class center; Calculate the index value of each attribute of the preprocessed data under different values ​​according to the attribute evaluation index; Based on the indicator value, an optimal attribute splitting point of the preprocessed data is determined.

5. The data collection and processing method in an industrial pilot platform tool according to claim 1, characterized in that: The performing classification structure processing on the screening data based on the data structure to obtain structure-processed data includes: Based on the data structure, the screening data is divided into structured data and unstructured data; Constructing a relational data table of the structural data; Performing data integration on the structured data using the relational data table to obtain integrated data; performing classification processing on the unstructured data to obtain classified processing data; The integrated data and the classified processed data are merged to obtain structure processed data.

6. The data collection and processing method in the industrial pilot platform tool according to claim 5, characterized in that: The classifying and processing the unstructured data to obtain the classified processed data includes: Identifying text data, image data, and audio data in the unstructured data; Performing natural language processing on the text data to obtain processed text data; performing image processing on the image data to obtain image processed data; Performing audio processing on the audio data to obtain audio processed data; The processed text data, the image processed data and the audio processed data are fused to obtain classified processed data.

7. A data acquisition and processing system in an industrial pilot platform tool, characterized in that: A system for executing a data collection and processing method in an industrial pilot platform tool according to any one of claims 1 to 6, comprising: A data collection module is used to obtain data collection requirements of the industrial pilot platform tool, so as to collect industrial pilot data according to the data collection requirements and obtain collected data; a data class center identification module, configured to perform data preprocessing on the collected data to obtain preprocessed data, identify data features of the preprocessed data, calculate feature averages of the data features, and determine the data class center of the preprocessed data based on the feature averages; a data screening module, configured to analyze an optimal attribute splitting point of the preprocessed data using the data class center, identify a spatial local deviation rate of the preprocessed data based on the optimal attribute splitting point, sort the preprocessed data based on the spatial local deviation rate to obtain distributed data, identify spatial access points of the distributed data, construct a data grid of the preprocessed data based on the spatial access points, and screen the distributed data using the data grid to obtain screened data; A data storage module is used to identify the data structure of the screening data, classify and structure the screening data based on the data structure to obtain structure-processed data, encrypt and protect the structure-processed data to obtain secure data, and store the secure data to obtain target data.

Citation Information

Patent Citations

  • High-dimensional feature data classification method and system based on distributed parallel decision tree

    CN111259933A

  • Equipment interactive information service system and method based on big data

    CN118228283A

  • Facility heterogeneous data acquisition method and system for building intellectualization

    CN118276793A