Classification and grading method based on industrial control protocol field content
Through the classification and grading method based on the content of the industrial control protocol field, the data of the industrial control system is classified, which solves the problem of low efficiency in industrial data use and analysis in the prior art, and realizes more efficient data management and utilization.
Patent Information
- Application Number
- CN202510145583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The prior art does not classify the data of industrial control systems for industrial control protocols, resulting in inefficient use and analysis of industrial data.
A classification and grading method based on the content of the industrial control protocol field is provided. By obtaining the original data flow of industrial equipment, preliminary data processing and protocol analysis are performed, classification type is determined based on the field content of the data frame, key fields are extracted, feature vector groups are defined, and supervised learning algorithms are selected according to the data positioning trend for classification.
It improves the utilization value and management efficiency of industrial control data, ensures the accuracy and efficiency of data use and analysis, and meets the differentiated needs of different business scenarios to data importance and sensitivity.
Smart Images

Figure CN119621971B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a classification and grading method based on the content of industrial control protocol fields. Background Art
[0002] With the in-depth development of industrial intelligent manufacturing, the value of industrial data has become increasingly prominent. As the core element of intelligent manufacturing, the security and management of industrial data are becoming the key to determining the competitiveness of industrial enterprises. By classifying and grading industrial data, industrial enterprises can more accurately identify important sensitive data within the organization and formulate corresponding protection measures to balance the contradiction between data circulation and data security. Industrial data classification and grading can help industrial enterprises achieve refined management and control of data assets, effectively monitor the dynamic flow of sensitive data, and ensure the visibility and controllability of data use and data sharing behaviors. This is of great significance for improving the data management capabilities of enterprises, optimizing data processing processes, and supporting business decisions.
[0003] Publication No. CN118820469A discloses a data classification and grading method, which belongs to the technical field of data classification and grading. It includes: step one, the user uploads the data to be classified to the server, and selects the field that does not provide specific data; step two, obtain the data field and content, and execute step three when the data field does not contain template data, otherwise execute step four; step three, use the recognition model to scan the data field, if the output result uniquely corresponds to a certain data in the template, then determine the field level according to the corresponding relationship in the template, otherwise it is considered that the field does not belong to the template range, and execute step four; step four, use the recognition model to scan the data field, and match the output result with the data in the template to form a regular rule array, and the regular rule array represents the matching result; step five, execute the matching process to obtain the classification level; the invention proposes a weight matching function. It can classify and grade data types more accurately. It can be seen that the prior art has the following problems:
[0004] Failure to classify industrial control system data according to industrial control protocols results in inefficient use and analysis of industrial data. Summary of the invention
[0005] To this end, the present invention provides a classification and grading method based on the content of industrial control protocol fields, so as to overcome the problem in the prior art that the data of industrial control systems are not graded according to industrial control protocols, resulting in low efficiency in the use and analysis of industrial data.
[0006] To achieve the above object, the present invention provides a classification and grading method based on the content of industrial control protocol fields, comprising:
[0007] Obtain the original data stream of each industrial equipment;
[0008] Performing preliminary data processing on each of the original data streams to obtain a corresponding standard data stream;
[0009] Parsing the data frames in each of the standard data streams by a protocol parser;
[0010] Determine the classification type according to the field content corresponding to each of the data frames, the classification type including control instruction type, status feedback type, parameter configuration type, communication management type and system maintenance type;
[0011] Extracting key fields of each of the data frames, wherein the key fields include commands, addresses, and parameter values;
[0012] Defining a feature vector group according to business requirements and determining its key features based on the key fields;
[0013] Determine the data location trend based on the classification type of standard data flow and the key characteristics to determine the classification method, including:
[0014] According to the results of the determination of the general trend, the general classification model pre-trained based on the supervised learning algorithm is used for classification;
[0015] Or, according to the determination result of the sensitivity trend, determine to use the sensitivity classification model pre-trained based on the supervised learning algorithm for classification;
[0016] Wherein, the data positioning trend includes sensitive trend and common trend;
[0017] The classification types and grading results are stored in the database and displayed through a visual interface.
[0018] Furthermore, the method of performing preliminary data processing on each of the original data streams to obtain a corresponding standard data stream includes:
[0019] Determine invalid data in the original data stream based on the data format and data representation status of each original data in the original data stream;
[0020] Determining redundant data based on a timestamp and a data value of each original data in the original data stream;
[0021] Eliminating the invalid data and the redundant data to obtain an original valid data stream;
[0022] The original valid data stream is converted into a standard format.
[0023] Furthermore, the method for determining invalid data in the original data stream based on the data format and data representation state of each original data in the original data stream includes:
[0024] Determine whether the data format of the original data conforms to the setting format of the corresponding industrial equipment;
[0025] Performing data fitting on each raw data in the raw data stream to obtain a fitting curve;
[0026] Determine the data value difference between the original data and two adjacent original data respectively, and record them as the first difference and the second difference;
[0027] Determine a corresponding data difference ratio according to a ratio of the first difference to the second difference of a single piece of original data;
[0028] Determine a data characterization state of the original data according to the fitting curve and the data difference ratio, wherein the data characterization state includes a data consistency state and a data discrete state;
[0029] According to the data format of the original data, whether it conforms to the setting format of the corresponding industrial equipment is determined and the data characterization state is determined to determine whether the original data is invalid data, wherein:
[0030] If the data format of the original data conforms to the setting format of the corresponding industrial equipment and the data representation state is a data consistent state, then the original data is determined to be valid data;
[0031] If the data format of the original data does not conform to the setting format of the corresponding industrial equipment, and / or the data representation state is a data discrete state, the original data is determined to be invalid data.
[0032] Furthermore, the method for determining the data characterization state of the original data according to the fitting curve and the data difference ratio includes:
[0033] Obtaining the data difference ratio of each original data to determine the average difference ratio;
[0034] Determine the Y-axis distance between the original data and the fitting curve;
[0035] The data characterization state of the original data is determined according to the Y-axis distance and the data difference ratio, wherein:
[0036] If the Y-axis distance is less than or equal to a preset distance and the data difference ratio is within a preset range, determining that the data representation state is a data consistent state;
[0037] Wherein, the preset range is determined according to the average difference ratio.
[0038] Further, redundant data is determined based on the timestamp and data value of each original data in the original data stream, wherein:
[0039] If the time difference between the corresponding timestamps of original data with the same data value is less than or equal to the preset duration, it is determined to be redundant data.
[0040] Furthermore, the method of parsing the data frames in each standard data stream by the protocol parser includes:
[0041] Parsing the data frame by the protocol parser, wherein the protocol parser is pre-set with a protocol parsing library;
[0042] The parsing status is determined based on the parsing time and parsing content, where:
[0043] If the parsing duration is less than or equal to the reference duration and the completeness of the parsed content is greater than or equal to the reference completeness, then the parsing state is determined to be a standard state;
[0044] If the parsing time is longer than the reference time, and / or the completeness of the parsed content is less than the reference completeness, the parsing state is determined to be a stuck state;
[0045] Determine the parsing of data frames in each standard data stream according to the parsing state, wherein:
[0046] If the parsing state is a standard state, each data frame is parsed according to the protocol parser;
[0047] If the parsing state is a stuck state, the protocol parser is optimized according to the cause of the stuck state and then each data frame is parsed.
[0048] Further, the protocol parser is optimized according to the cause of the hindrance, including:
[0049] If the cause of the stagnation is that the parsing time is longer than the reference time, the current protocol parser is optimized to shorten the parsing time;
[0050] If the cause of the stagnation is that the completeness of the parsed content is less than the reference completeness, the current parsing protocol library is updated to increase the completeness of the parsed content.
[0051] Furthermore, the method of defining a feature vector group according to business requirements and determining its key features based on the key fields includes:
[0052] Determine all business aspects, all operating parameters and business objectives of business needs;
[0053] Identify the business links related to the business objectives as important links;
[0054] Determine the operating parameters of the important link and the equipment parameters obtained by the important link as feature vectors to determine a feature vector group;
[0055] The content in the key field that is identical to any feature vector is determined as a key feature.
[0056] Furthermore, the method for determining the data positioning trend according to the classification type and the key features of each data frame in a single standard data stream includes:
[0057] Determine a first proportion of key type data frames in the standard data stream, wherein the key types are control instruction types, status feedback types, and parameter configuration types in the classification types;
[0058] Determine the ratio of the number of key fields containing key features and record it as the second ratio;
[0059] The data positioning trend is determined according to the first proportion and the second proportion, wherein:
[0060] If the first proportion and the second proportion are respectively greater than the first reference proportion and the second reference proportion, determining that the data positioning trend is a sensitive trend;
[0061] If the first proportion is less than or equal to a first reference proportion, and / or the second proportion is less than or equal to a second reference proportion, then the data positioning trend is determined to be a normal trend.
[0062] Furthermore, the number of classification levels in the common classification model is greater than the number of classification levels in the sensitive classification model.
[0063] Compared with the prior art, the beneficial effect of the present invention lies in that the classification and grading method based on the content of the industrial control protocol field first obtains the original data stream from various interfaces of industrial equipment, covering a variety of common industrial protocols, and providing a rich data foundation for subsequent processing; then, the standard data stream is obtained through preliminary data processing, which effectively solves the problems of large amount of original data, diverse formats and noise errors, and improves data quality; after the protocol parser parses the data frame, it accurately divides a variety of classification types according to the field content, which helps to clearly understand the purpose of data; extracts key fields and defines feature vector groups according to business needs to determine key features, so that data processing is more in line with actual needs; determines data positioning trends according to classification types and key features, and then selects appropriate supervised learning classification models for grading, so that grading is more scientific and targeted; finally, the classification types and grading results are stored in the database and displayed through a visual interface, which is convenient for users to intuitively obtain information, facilitates data analysis, decision making and system monitoring and management, and comprehensively improves the utilization value and management efficiency of industrial control data;
[0064] Furthermore, invalid data is determined based on data format and representation status, redundant data is determined based on timestamp and data value, invalid data is eliminated and some redundant data is reasonably retained, and finally the original valid data stream is converted into a custom standard format. This series of operations effectively improves data quality and lays a good foundation for subsequent protocol parsing, classification and grading, making the analysis of the entire industrial control protocol field content more accurate and efficient.
[0065] Furthermore, the data format of the original data is first compared with the setting format of the corresponding industrial equipment, and the original data is fitted using mathematical software to obtain a fitting curve, and the data representation state is determined by comparing the difference and ratio of adjacent original data, and then the original data is judged whether it is invalid data according to the determination result of the data format and the data representation state, which helps to accurately screen out data that does not conform to the equipment setting format and has abnormal data representation, thereby ensuring the validity and consistency of the data;
[0066] Furthermore, the data difference ratio of each original data is obtained to determine the average difference ratio, and then the Y-axis distance between the original data and the fitting curve is determined. Finally, the data characterization state is accurately judged based on the Y-axis distance and the data difference ratio. This method fully considers the statistical characteristics of the data, combines the average difference ratio and the Y-axis distance, and evaluates the original data in a more scientific and reasonable way, which helps to accurately distinguish the data consistency state and the data discrete state, and provides a more reliable basis for the subsequent invalid data judgment.
[0067] Furthermore, the data frames in the standard data stream are parsed by using a pre-set protocol parser, and the parsing status is determined according to the parsing time and the parsing content, and then the subsequent data frame parsing method is determined according to the parsing status; the existing protocol parsing library is used for routine operations, and problems can be discovered in time according to the parsing status, and possible obstructions can be optimized, thereby ensuring accurate parsing of the data frames and improving the efficiency and quality of parsing;
[0068] Furthermore, the feature vector group is defined according to the business needs, and the key features are determined based on the key fields; first, the various elements of the business needs are clarified, including all business links, operating parameters and business goals, and then the business links related to the business goals are screened as important links, and then the operating parameters and equipment parameters involved in these important links are used as feature vectors, and finally the same content in the key fields is found as the key features; this series of operations helps to closely combine business needs with data processing, ensure that the extracted features have a clear business orientation, make subsequent data processing and analysis more in line with actual business needs, and provide a more targeted and practical data feature foundation for the entire industrial control data processing process;
[0069] Furthermore, the first proportion of key type data frames in the standard data stream and the second proportion of the number of key fields containing key features are first determined, and then the data positioning trend is determined based on the comparison of these two proportions with the corresponding reference proportions; this method can effectively judge the sensitivity of the data accurately from the data category and the proportion of key features, and provides a strong basis for the subsequent selection of an appropriate classification model, so that the classification results are more in line with the actual importance and sensitivity of the data, which helps to improve the pertinence and effectiveness of the entire industrial control protocol data processing, and better meet the business needs for data management and analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A step diagram of a classification and grading method based on industrial control protocol field content according to an embodiment of the present invention;
[0071] Figure 2 A method step diagram for obtaining a corresponding standard data stream according to an embodiment of the present invention;
[0072] Figure 3 A step diagram of a method for determining invalid data according to an embodiment of the present invention;
[0073] Figure 4 The present invention is a flowchart of a method for parsing data frames in various standard data streams by using a protocol parser. DETAILED DESCRIPTION
[0074] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0075] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0076] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0077] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0078] See also Figure 1 As shown, it is a step diagram of a classification and grading method based on the content of industrial control protocol fields according to an embodiment of the present invention. An embodiment of the present invention provides a classification and grading method based on the content of industrial control protocol fields, including:
[0079] Step S1, obtaining the original data stream of each industrial device; it can be understood that the original data stream of industrial equipment refers to the unprocessed data sequence directly obtained from various interfaces of industrial equipment (such as sensor interface, controller interface, communication interface, etc.), which exists in the form of a continuous stream and contains various information during the operation of the industrial equipment; the characteristics of the original data stream are large data volume, diverse formats, and may contain noise and error information; in implementation, the original data stream is obtained from the industrial control system through a network packet capture tool or a direct interface, and supports a variety of common industrial protocols, including SCADA, PLC, Modbus, PROFINET and EtherNet / IP, etc.;
[0080] Step S2, performing preliminary data processing on each of the original data streams to obtain a corresponding standard data stream;
[0081] Step S3, parsing the data frames in each of the standard data streams by a protocol parser;
[0082] Step S4, determining the classification type of each data frame according to the field content corresponding to each data frame, the classification type including control instruction type, status feedback type, parameter configuration type, communication management type and system maintenance type;
[0083] Step S5, extracting key fields of each of the data frames, wherein the key fields include commands, addresses and parameter values;
[0084] Step S6, defining a feature vector group according to business requirements and determining its key features based on the key fields;
[0085] Step S7, determining the data positioning trend of the standard data stream according to the classification type of the standard data stream and the key feature to determine the classification method of the standard data stream, including:
[0086] According to the results of the determination of the general trend, the general classification model pre-trained based on the supervised learning algorithm is used for classification;
[0087] Or, according to the determination result of the sensitivity trend, determine to use the sensitivity classification model pre-trained based on the supervised learning algorithm for classification;
[0088] Wherein, the data positioning trend includes sensitive trend and common trend;
[0089] Step S8, storing the classification type and grading results in a database and displaying them through a visual interface.
[0090] It can be understood that the classification and grading method based on the content of industrial control protocol fields provided by the embodiment of the present invention starts with obtaining the original data stream, and performs a series of operations such as preliminary processing, protocol parsing, classification, key field extraction, feature definition and grading, which comprehensively covers all aspects of industrial control protocol data processing and ensures the full utilization and management of industrial control data; secondly, by classifying based on the content of the protocol field, defining feature vectors in combination with business needs to determine key features, and using different supervised learning models for grading according to data positioning trends, the classification and grading are more scientific and targeted, and can meet the needs of different business scenarios for distinguishing the importance and sensitivity of data; at the same time, the method supports a variety of common industrial protocols and has a wide range of applicability, and can be applied to various complex industrial control environments; thirdly, the embodiment of the present invention stores the classification type and grading results in a database and displays them through a visual interface, so that users can intuitively understand the classification and grading of industrial control data, and facilitate data analysis, decision making, and system monitoring and management.
[0091] See also Figure 2 As shown, it is a step diagram of the method for obtaining the corresponding standard data stream according to an embodiment of the present invention. Specifically, in step S2, the method for performing preliminary data processing on each of the original data streams to obtain the corresponding standard data stream includes:
[0092] Step S21, determining invalid data in the original data stream based on the data format and data representation state of each original data in the original data stream;
[0093] Step S22, determining redundant data based on the timestamp and data value of each original data in the original data stream;
[0094] Step S23, removing the invalid data and the redundant data to obtain the original valid data stream; it can be understood that all invalid data are removed and 1 to 3 identical redundant data are retained;
[0095] Step S24, converting the original valid data stream into a standard format; in implementation, the standard format is defined by a custom method.
[0096] In practice, the definition of the standard format needs to include:
[0097] Data length fixed format (byte alignment format, fixed frame length format);
[0098] Data type clear format:
[0099] Numerical format: ① Integer format: The data length (such as 8-bit integer, 16-bit integer, etc.) and representation (signed or unsigned) of integers are clearly specified. For example, in the data transmission of a counting sensor, a 16-bit unsigned integer is used to represent the product count. The receiver can accurately convert the received binary data into a decimal count value according to this format; ② Floating-point format: For floating-point numbers representing physical quantities (such as temperature, pressure, etc.), they are defined according to specific floating-point standards (such as IEEE 754 standards). The receiving device can correctly parse the floating-point data to ensure the accuracy of the data; for example, in an environmental monitoring system, the data of the temperature sensor is transmitted in the IEEE 754 single-precision floating-point format, and the receiver can accurately convert the received data into the actual temperature value;
[0100] Character format: ①ASCII format: If the data is character-based, such as device name, fault description, etc., it may be transmitted in ASCII code format. For example, the fault information "Motor Overload" sent by the device is converted into a corresponding byte sequence according to the ASCII code for transmission. The receiver converts the byte sequence back to characters through the ASCII code table to obtain the fault description information; ②UTF-8 format: In some industrial devices that need to support characters in multiple languages, the UTF-8 format may be used. For example, in an international industrial control system, the user interface configuration information of the device (such as menu options, operation prompts, etc.) may contain multiple languages. The UTF-8 format can correctly transmit and display these character information;
[0101] Boolean format: A Boolean value is usually represented by a binary bit or a byte (such as 0 for False and 1 for True). For example, the switch status (on / off) and alarm status (triggered / untriggered) of a device can be transmitted in Boolean format. The receiver can simply judge the status of the device based on this format.
[0102] (3) Protocol header and tail format:
[0103] Protocol identifier format: In the header of the data frame, there is a fixed protocol identifier format;
[0104] Address format: including the format of source address and destination address;
[0105] Checksum format: At the end of the data frame, there is usually a checksum (such as CRC-cyclic redundancy check) format;
[0106] (4) Data order and arrangement format:
[0107] Big-endian or little-endian format: In the storage and transmission of multi-byte data (such as integers and floating-point numbers), it is necessary to clarify whether it is big-endian or little-endian. Big-endian stores the high-order bytes of the data at the low address, while little-endian does the opposite. For example, for a 32-bit integer 0x12345678, if it is big-endian, the byte order in memory or transmission is 0x12, 0x34, 0x56, 0x78; if it is little-endian, it is 0x78, 0x56, 0x34, 0x12. The industrial control protocol needs to clarify this data order format so that the receiving device can correctly parse the data.
[0108] Data grouping and arrangement format: Data may be grouped and arranged according to certain functions or device types; for example, in a complex industrial control system, a data frame may first arrange a group of sensor data (such as temperature sensor, pressure sensor data), and then a group of actuator control instruction data. This grouping arrangement format helps the receiving device to classify and process the data according to the functional modules.
[0109] It can be understood that step S21 determines invalid data based on the data format and data representation status of each original data in the original data stream, and can identify and eliminate data that does not meet the requirements due to data format errors or abnormal representation status, so as to avoid interference of these invalid data with subsequent analysis, thereby improving the accuracy of the data and providing a reliable data basis for subsequent processing; step S22 determines redundant data based on timestamps and data values, and step S23 eliminates redundant data, retaining only 1 to 3 identical redundant data, which not only reduces the burden of data storage, but also avoids information loss caused by excessive elimination; this optimizes the data volume and improves data processing efficiency, so that subsequent protocol parsing, classification and grading and other operations can be performed on more streamlined data, reducing the waste of computing resources; step S24 converts the original valid data stream into a custom standard format, so that original data from different sources and in different formats have a unified format, eliminating processing obstacles caused by differences in data formats, facilitating subsequent protocol parsers and other tools to uniformly process data, and improving the consistency and efficiency of the entire industrial control protocol field content analysis process.
[0110] See also Figure 3 Specifically, in step S21, the method for determining invalid data in the original data stream based on the data format and data representation state of each original data in the original data stream includes:
[0111] Step S211, determining whether the data format of the original data conforms to the setting format of the corresponding industrial equipment;
[0112] Step S212, performing data fitting on each raw data in the raw data stream to obtain a fitting curve; in implementation, mathematical software (such as Matlab, Origin, etc.) is used to perform data fitting, and the most suitable fitting curve is determined by comparing the goodness of fit of different fitting functions (such as can be judged according to the determination coefficient, the closer to 1, the better the fitting effect);
[0113] Step S213, respectively determining the data value difference between a single original data and two adjacent original data, recorded as a first difference and a second difference; in implementation, the first difference is the data value difference between the current original data and the previous original data, and the second difference is the data value difference between the current original data and the next original data;
[0114] Step S214, determining a corresponding data difference ratio according to the ratio of the first difference to the second difference of the single original data; in implementation, the data difference ratio=|first difference-second difference|÷first difference;
[0115] Step S215, determining a data characterization state of the original data according to the fitting curve and the data difference ratio, wherein the data characterization state includes a data consistency state and a data discrete state;
[0116] Step S216, determining whether the data format of the original data conforms to the setting format of the corresponding industrial device and the data characterization state to determine whether the original data is invalid data, wherein:
[0117] If the data format of the original data conforms to the setting format of the corresponding industrial equipment and the data representation state is a data consistent state, then the original data is determined to be valid data;
[0118] If the data format of the original data does not conform to the setting format of the corresponding industrial equipment, and / or the data representation state is a data discrete state, the original data is determined to be invalid data;
[0119] It is understandable that the data of each industrial device has its corresponding setting format. If the data format of the original data is different from the setting format, it is determined that the original data has a problem (invalid data).
[0120] It can be understood that step S211 determines whether the data format of the original data conforms to the setting format of the corresponding industrial equipment according to the data format of the original data, which can ensure that the data flowing into the subsequent processing link conforms to the inherent format requirements of the industrial equipment; this can avoid parsing errors and data processing anomalies caused by format inconsistency, ensure the accuracy and reliability of the entire data processing process, and provide a data basis with correct format for subsequent data analysis, protocol parsing and other operations; steps S212 to S215 use mathematical software to fit the data, and introduce the concept of data difference ratio, and determine the data characterization state based on the fitting curve and the data difference ratio, which helps to judge whether the data is normal from the distribution and change trend of the data. , which can effectively discover data with abnormal data representation, such as data discreteness caused by sensor failure, communication interference, etc., and enhance the ability to control data quality; step S216 determines whether the original data is invalid data based on the judgment result of the data format and the data representation status, and comprehensively considers the two factors of data format and representation, and realizes accurate identification and elimination of invalid data, which makes the subsequent data processing process only for valid data, improves the efficiency of data processing, avoids the interference of invalid data on subsequent steps, and provides a purer and more reliable data resource for the processing of industrial control protocol field content, ensuring the stable operation of the entire system.
[0121] Specifically, in step S215, the method for determining the data characterization state of the original data according to the fitting curve and the data difference ratio includes:
[0122] Step S2151, obtaining the data difference ratio of each original data to determine the average difference ratio;
[0123] Step S2152, determining the Y-axis distance between the original data and the fitting curve; it can be understood that the abscissa of the fitting curve is usually time, and the ordinate is the data value of each original data; in one implementation, the coordinates of the original data A are (x1, y1), and the coordinates of the corresponding point when the abscissa on the fitting curve is x1 are (x1, y2), then the Y-axis distance of the original data A is |y1-y2|;
[0124] Step S2153, determining the data characterization state of the original data according to the Y-axis distance and the data difference ratio, wherein:
[0125] If the Y-axis distance is less than or equal to the preset distance and the data difference ratio is within a preset range, the data representation state is determined to be a data consistency state; in implementation, the preset distance is usually set to a data value standard deviation of all original data forming the fitting curve to twice the data value standard deviation; the larger the preset distance is set, the farther the distance between the valid original data and the fitting curve may be, that is, the condition for the data consistency state of invalid data is relaxed;
[0126] Among them, the preset range is determined according to the average difference ratio. In implementation, the middle value of the preset range is the average difference ratio. The preset range is usually the difference between the standard deviation of the average difference ratio and the data difference ratio~the sum of the standard deviation of the average difference ratio and the data difference ratio.
[0127] It can be understood that step S215 uses the two key indicators of data difference ratio and Y-axis distance in a comprehensive manner, avoiding the one-sidedness of judging the data characterization state by a single indicator. The average difference ratio is obtained by step S2151, and then the Y-axis distance in step S2152 is calculated, which combines the local difference of the data (data difference ratio) and the overall deviation (Y-axis distance), comprehensively considers the distribution characteristics of the data, and can more accurately reflect the consistency of the data; step S2153 determines the preset range based on the average difference ratio, and at the same time reasonably sets the preset distance based on the standard deviation of the data values of all the original data, providing a scientific and adjustable standard for determining the data characterization state. The preset range and preset distance are dynamically set according to the characteristics of the data itself, so that the judgment process is more in line with the actual data, reducing the possibility of misjudgment and improving the accuracy of data consistency judgment; in implementation, accurately distinguishing whether the data representation state is a data consistency state or a data discrete state helps to screen out data that is more in line with the actual situation from a large amount of original data, which is crucial for subsequent invalid data judgment and the processing of the entire industrial control protocol field content, ensuring the quality of data entering subsequent processes, enhancing the reliability and stability of the data processing system, and avoiding a series of problems caused by abnormal data representation, such as erroneous control instructions or inaccurate system analysis.
[0128] Specifically, in step S22, redundant data is determined based on the timestamp and data value of each original data in the original data stream, wherein:
[0129] If the time difference between the corresponding timestamps of the original data with the same data value is less than or equal to the preset time length, the original data with the same data value are determined to be redundant data.
[0130] See also Figure 4 As shown, it is a flow chart of a method for parsing data frames in each standard data stream by a protocol parser according to an embodiment of the present invention. Specifically, in step S3, the method for parsing data frames in each standard data stream by a protocol parser includes:
[0131] Step S31, parsing the data frame by the protocol parser, wherein the protocol parser is pre-set with a protocol parsing library, which is usually an existing protocol parsing library;
[0132] Step S32, determining the parsing status according to the parsing time and the parsing content, wherein:
[0133] If the parsing duration is less than or equal to the reference duration and the completeness of the parsed content is greater than or equal to the reference completeness, then the parsing state is determined to be a standard state;
[0134] If the parsing time is longer than the reference time, and / or the completeness of the parsed content is less than the reference completeness, the parsing state is determined to be a stuck state;
[0135] It is understood that the completeness of the parsed content is usually expressed as a percentage, with a completeness of ≤100%; in practice, the longer the reference time, the looser the judgment condition of the standard state, which is usually set to within 5 minutes (including 5 minutes); the more complete the parsed content, the more accurate the judgment of the data classification, which is usually set to a larger value, preferably, the reference completeness is ≥95%;
[0136] Step S33, determining to parse the data frames in each standard data stream according to the parsing state, wherein:
[0137] If the parsing state is the standard state, each data frame is parsed according to the current protocol parser;
[0138] If the parsing state is a stuck state, the protocol parser is optimized according to the cause of the stuck state and then each data frame is parsed.
[0139] It can be understood that step S32 determines the parsing status based on the parsing time and the completeness of the parsed content, taking into account the two key factors of time and content integrity, and reasonably setting the standards of reference time and reference completeness, so that the determination of the parsing status is more scientific, and can effectively distinguish between the standard state and the stagnation state, providing a clear basis for subsequent optimization or continued parsing operations, and improving the monitoring and management capabilities of the parsing process; when the parsing status is a stagnation state, in step S33, the protocol parser can be optimized according to the cause of the stagnation, so as to ensure the parsing quality of the data frame, which reflects the adaptability and optimizability of the method, avoids affecting subsequent data processing due to excessive parsing time or incomplete parsing content, ensures the accuracy and completeness of data frame parsing, and provides high-quality parsing data for subsequent classification, grading and other processes.
[0140] Specifically, in step S33, the protocol parser is optimized according to the cause of the hindrance, including:
[0141] If the cause of the stagnation is that the parsing time is longer than the reference time, the current protocol parser is optimized to shorten the parsing time; in implementation, the method of optimizing the protocol parser includes optimizing the storage structure, optimizing the cache mechanism, improving the hardware resource configuration (bandwidth, memory capacity, etc.), and setting a parallel parsing strategy;
[0142] If the cause of the stagnation is that the completeness of the parsed content is less than the reference completeness, the current parsing protocol library is updated to increase the completeness of the parsed content.
[0143] Specifically, in step S6, the method of defining a feature vector group according to business requirements and determining its key features based on the key fields includes:
[0144] Step S61, determining all business links, all operating parameters and business objectives of the business requirements;
[0145] Step S62, determining the business links related to the business goal as important links;
[0146] Step S63, determining the operating parameters of the important link and the equipment parameters obtained by the important link as feature vectors to determine a feature vector group;
[0147] Step S64: determine the content in the key field that is the same as any feature vector as the key feature.
[0148] Specifically, in step S7, the method for determining the data positioning trend according to the classification type and the key features of each data frame in a single standard data stream includes:
[0149] Step S71, determining a first proportion of key type data frames in the standard data stream, wherein the key types are control instruction types, status feedback types, and parameter configuration types in the classification types;
[0150] Step S72, determining the ratio of the number of key fields containing the key features and recording it as the second ratio;
[0151] Step S73, determining a data positioning trend according to the first proportion and the second proportion, wherein:
[0152] If the first proportion and the second proportion are respectively greater than the first reference proportion and the second reference proportion, determining that the data positioning trend is a sensitive trend;
[0153] If the first proportion is less than or equal to a first reference proportion, and / or the second proportion is less than or equal to a second reference proportion, then the data positioning trend is determined to be a normal trend;
[0154] In implementation, the values of the first reference ratio and the second reference ratio are both greater than or equal to 50%. The larger the value, the more accurate the data of the determined sensitive trend.
[0155] It is understandable that, based on the comparison results of the first proportion and the second proportion with the reference proportion, the data positioning trend is determined to be a sensitive trend or a common trend, so as to decide to adopt a sensitive classification model or a common grading model pre-trained based on a supervised learning algorithm for grading, so that the selection of the grading model is more scientific and reasonable, and can be targeted according to the actual sensitivity of the data, avoiding the problem of over-generalization or over-refinement that may be caused by the use of a single model, improving the accuracy and effectiveness of the grading results, and better meeting the grading requirements of data with different sensitivities; by accurately judging the data positioning trend, different grading models are adopted for data with different trends, making the entire industrial control protocol data processing process more targeted: for data with sensitive trends, a more sophisticated sensitive classification model is adopted to better capture the key information and potential risks in the data; for data with common trends, the use of a common grading model can improve processing efficiency and reduce the consumption of computing resources while ensuring the processing effect; this differentiated processing method improves the overall efficiency and quality of data processing, and meets the differentiated management needs of the business for different types of data.
[0156] Specifically, in step S7, the number of classification levels in the common classification model is greater than the number of classification levels in the sensitive classification model; it can be understood that the importance levels in the common classification model include level 1, level 2, level 3 and level 4 from small to large, while the importance levels in the sensitive classification model include level 2, level 3 and level 4 from small to large; since the data using the sensitive classification model has a sensitive trend, the data should not be at the lowest security level before classification;
[0157] In implementation, the grading process of the two grading models is the same, but the output grading results are different; for example, under the same conditions, if the output result of the ordinary grading model is level 1, then the output result of the sensitive grading model is level 2; if the output result of the ordinary grading model is level 2, then the output result of the sensitive grading model is level 3; if the output result of the ordinary grading model is level 3, then the output result of the sensitive grading model is level 4; if the output result of the ordinary grading model is top level, then the output result of the sensitive grading model is also top level (level 4).
[0158] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0159] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A classification and grading method based on the content of industrial control protocol fields, characterized in that: include: Obtain the original data stream of each industrial equipment; Performing preliminary data processing on each of the original data streams to obtain a corresponding standard data stream; The data frames in each of the standard data streams are parsed by a protocol parser, and the specific steps of parsing the data frames in each of the standard data streams by a protocol parser include: Parsing the data frame by the protocol parser, wherein the protocol parser is pre-set with a protocol parsing library; The parsing status is determined based on the parsing time and parsing content, where: If the parsing duration is less than or equal to the reference duration and the completeness of the parsed content is greater than or equal to the reference completeness, then the parsing state is determined to be a standard state; If the parsing time is longer than the reference time, and / or the completeness of the parsed content is less than the reference completeness, the parsing state is determined to be a stuck state; Determine the parsing of data frames in each standard data stream according to the parsing state, wherein: If the parsing state is a standard state, each data frame is parsed according to the protocol parser; If the parsing state is a stuck state, the protocol parser is optimized according to the cause of the stuck state and then each data frame is parsed; Determine the classification type according to the field content corresponding to each of the data frames; Extracting key fields of each of the data frames, wherein the key fields include commands, addresses, and parameter values; Defining a feature vector group according to business requirements and determining its key features based on the key fields; Determine the data location trend based on the classification type of standard data flow and the key characteristics to determine the classification method, including: According to the results of the determination of the general trend, the general classification model pre-trained based on the supervised learning algorithm is used for classification; Or, according to the determination result of the sensitivity trend, determine to use the sensitivity classification model pre-trained based on the supervised learning algorithm for classification; Wherein, the data positioning trend includes sensitive trend and common trend; The classification types and grading results are stored in the database and displayed through a visual interface.
2. The classification and grading method based on the content of industrial control protocol fields according to claim 1 is characterized in that: The method of performing preliminary data processing on each of the original data streams to obtain a corresponding standard data stream includes: Determine invalid data in the original data stream based on the data format and data representation status of each original data in the original data stream; Determining redundant data based on a timestamp and a data value of each original data in the original data stream; Eliminating the invalid data and the redundant data to obtain an original valid data stream; The original valid data stream is converted into a standard format.
3. The classification and grading method based on the content of industrial control protocol fields according to claim 2 is characterized in that: The method for determining invalid data in the original data stream based on the data format and data representation state of each original data in the original data stream includes: Determine whether the data format of the original data conforms to the setting format of the corresponding industrial equipment; Performing data fitting on each raw data in the raw data stream to obtain a fitting curve; Determine the data value difference between the original data and two adjacent original data respectively, and record them as the first difference and the second difference; Determine a corresponding data difference ratio according to a ratio of the first difference to the second difference of a single piece of original data; Determine a data characterization state of the original data according to the fitting curve and the data difference ratio, wherein the data characterization state includes a data consistency state and a data discrete state; According to the data format of the original data, whether it conforms to the setting format of the corresponding industrial equipment is determined and the data characterization state is determined to determine whether the original data is invalid data, wherein: If the data format of the original data conforms to the setting format of the corresponding industrial equipment and the data representation state is a data consistent state, then the original data is determined to be valid data; If the data format of the original data does not conform to the setting format of the corresponding industrial equipment, and / or the data representation state is a data discrete state, the original data is determined to be invalid data.
4. The classification and grading method based on the content of industrial control protocol fields according to claim 3 is characterized in that: The method for determining the data characterization state of the original data according to the fitting curve and the data difference ratio includes: Obtaining the data difference ratio of each original data to determine the average difference ratio; Determine the Y-axis distance between the original data and the fitting curve; The data characterization state of the original data is determined according to the Y-axis distance and the data difference ratio, wherein: If the Y-axis distance is less than or equal to a preset distance and the data difference ratio is within a preset range, determining that the data representation state is a data consistent state; Wherein, the preset range is determined according to the average difference ratio.
5. The classification and grading method based on the content of industrial control protocol fields according to claim 2 is characterized in that: Redundant data is determined based on the timestamp and data value of each original data in the original data stream, wherein: If the time difference between the corresponding timestamps of original data with the same data value is less than or equal to the preset duration, it is determined to be redundant data.
6. The classification and grading method based on the content of industrial control protocol fields according to claim 1 is characterized in that: Optimizing the protocol parser according to the cause of the stagnation includes: If the cause of the stagnation is that the parsing time is longer than the reference time, the current protocol parser is optimized to shorten the parsing time; If the cause of the stagnation is that the completeness of the parsed content is less than the reference completeness, the current parsing protocol library is updated to increase the completeness of the parsed content.
7. The classification and grading method based on the content of industrial control protocol fields according to claim 1 is characterized in that: The method of defining a feature vector group according to business requirements and determining its key features based on the key fields includes: Determine all business aspects, all operating parameters and business objectives of business needs; Identify the business links related to the business objectives as important links; Determine the operating parameters of the important link and the equipment parameters obtained by the important link as feature vectors to determine a feature vector group; The content in the key field that is identical to any feature vector is determined as a key feature.
8. The classification and grading method based on the content of industrial control protocol fields according to claim 7 is characterized in that: The method for determining the data positioning trend according to the classification type and the key features of each data frame in a single standard data stream includes: Determine a first proportion of key type data frames in the standard data stream, wherein the key types are control instruction types, status feedback types, and parameter configuration types in the classification types; Determine the ratio of the number of key fields containing key features and record it as the second ratio; The data positioning trend is determined according to the first proportion and the second proportion, wherein: If the first proportion and the second proportion are respectively greater than the first reference proportion and the second reference proportion, determining that the data positioning trend is a sensitive trend; If the first proportion is less than or equal to a first reference proportion, and / or the second proportion is less than or equal to a second reference proportion, then the data positioning trend is determined to be a normal trend.
9. The classification and grading method based on the content of industrial control protocol fields according to claim 8 is characterized in that: The number of hierarchical levels in the common hierarchical model is greater than the number of hierarchical levels in the sensitive hierarchical model.
Citation Information
Patent Citations
Data classification and grading method
CN118820469A
Metadata grading and classifying method based on machine learning algorithm
CN114676253A
Data security monitoring system based on traffic, electronic equipment and storage medium
CN115037559A