System data acquisition method and system based on OPC UA protocol, and storage medium
By deploying OPC UA server module and AI-driven intelligent classification and binding algorithms in industrial software systems, the existing system data silos and enclosure problems are solved, and the automated collection and integration of equipment data is realized, and the data utilization value and management efficiency are improved.
Patent Information
- Application Number
- CN202510503072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
AI Technical Summary
The interface design of existing industrial software systems is unfriendly, making it difficult to effectively interact and integrate equipment data with other systems, and is highly enclosed, increasing the complexity and cost of equipment maintenance and fault diagnosis.
The system data acquisition method based on the OPC UA protocol is adopted, and through the combination of OPC UA protocol and AI technology, the automated collection, analysis and automatic generation of device trees, measurement points and corresponding configuration configuration of the old system are realized, forming an integrated data analysis of the old system.
It realizes efficient integration and sharing of equipment data, improves the utilization value of data, reduces manual intervention and costs, and improves the accuracy and efficiency of equipment management and operating status analysis.
Smart Images

Figure CN120017680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a system data acquisition method, system and storage medium based on the OPC UA protocol. Background Art
[0002] In large domestic chemical companies, the operation status data collection of key equipment such as compressors usually relies on industrial software such as Bentley's System1 system. However, with the increasing requirements of enterprises for equipment management and the widespread application of information technology in the industrial field, these systems have gradually exposed some serious problems.
[0003] First, the interface design of these systems is extremely unfriendly, and the collected data cannot be sent to other systems through common API forms, which makes it difficult for the equipment data in the system to effectively interact and integrate with other equipment management systems or data analysis platforms within the enterprise. In actual equipment management work, this data isolation has caused many inconveniences. For example, when an enterprise needs to conduct a comprehensive analysis of the equipment operating status on the entire production line, due to the isolation of these system data, it is impossible to uniformly process and compare it with the data of other equipment, which affects the accurate judgment of the overall equipment operating status.
[0004] Secondly, due to the closed nature of these systems, enterprises rely heavily on manual labor when performing equipment maintenance and fault diagnosis. They often need to switch between different systems to find and integrate relevant data, which not only increases the complexity of work but also reduces work efficiency. Moreover, the existence of such data islands also hinders enterprises from deeply mining and utilizing equipment operation data, and cannot give full play to the potential value of data and provide a strong basis for equipment optimization and decision support for enterprises.
[0005] In addition, from the perspective of technological development, with the advancement of the concept of Industry 4.0 and the rise of intelligent manufacturing, enterprises have put forward higher requirements for the interconnection and intelligent management of equipment data. However, the existing architecture and data interface of these systems are obviously unable to meet these new requirements, which seriously restricts the digital transformation and intelligent development process of enterprises.
[0006] The most common existing technical solution is manual configuration and data forwarding based on middleware scripts, which is specifically: manually exporting the device tree, measurement point information and other information of the old system, and through the development of customized middleware, calling the interface of the old system to obtain data, and converting it into Modbus TCP protocol for transmission to other platforms. This solution has the following problems: manual operation is time-consuming and easily invalidated due to the modification of the old system configuration; the Modbus protocol lacks semantic description capabilities and data requires secondary mapping; the middleware development cost is high and it is difficult to adapt to multiple platforms; there is a lack of automated exception handling mechanism.
[0007] In summary, the most obvious and prominent problems in the existing technology are mainly two points: first, insufficient data intelligence: lack of automated analysis and classification technology, relying on manual processing, low efficiency and high cost; second, limited scalability: the existing architecture is difficult to support the real-time collection and processing of large-scale equipment data. In view of this, we propose a system data collection method, system and storage medium based on the OPC UA protocol. Summary of the invention
[0008] The object of the present invention is to provide a system data acquisition method, system and storage medium based on the OPC UA protocol to solve the problems raised in the above background technology.
[0009] In order to solve the above technical problems, one of the purposes of the present invention is to provide a system data collection method based on the OPC UA protocol, which combines the OPC UA protocol with AI technology to realize the automatic collection and analysis of the old system data and automatically generate the device tree, measurement points and corresponding configuration configuration, thereby forming an automatic analysis and integration of the old system data; specifically, it includes the following steps:
[0010] S1. OPC UA communication layer configuration: Deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes; configure the OPC UA standard client on the client, and set the security policy and session timeout mechanism;
[0011] S2, Data parsing and semantic modeling: String parsing based on regular expressions and context matching algorithms, extracting key fields, and mapping the parsed data to the OPC UA information model to generate a structured JSON format;
[0012] S3. AI-driven intelligent classification and binding: Optimize and train the long short-term memory network (LSTM) model. For the identified and classified information, use the intelligent data binding algorithm to accurately associate the measurement points with the corresponding equipment level. This process requires comprehensive consideration of multiple factors such as the type, location, and function of the equipment to ensure the accuracy of the binding.
[0013] S4, Data integration and exception handling: Design data conversion modules and exception monitoring functions, and integrate processed data with the device management platform;
[0014] S5. Distributed architecture design: Use Kubernetes containerized deployment to achieve elastic expansion of computing nodes; set up a data cache layer (Redis) to cope with high concurrency scenarios and ensure low-latency response; set up real-time monitoring and error handling mechanisms throughout the data collection and processing process.
[0015] As a further improvement of the technical solution, in S2, data parsing is performed by setting a string parsing module, and the string parsing module performs string parsing specifically including the following steps:
[0016] S2.1. Data preprocessing: Before performing string parsing, preprocess the original string data collected from the old system; remove redundant information in the string and convert the string into a unified format to facilitate subsequent processing;
[0017] S2.2. Define regular expression rules: define corresponding regular expression patterns according to the key fields to be extracted;
[0018] S2.3. Apply regular expressions to extract preliminary fields: Use the defined regular expressions to match the preprocessed strings. Use the regular expression matching function in the programming language to extract string fragments that match the regular expression pattern.
[0019] S2.4. Context matching to optimize extraction results: Introduce a context matching algorithm to improve the accuracy of extraction. After extracting the preliminary string, make judgments and corrections based on its context information.
[0020] As a further improvement of the present technical solution, in S2.4, the context matching algorithm uses a string similarity algorithm to assist in determining whether the extracted field is accurate during application, and adopts an edit distance algorithm. The edit distance refers to the minimum number of editing operations (insertion, deletion, replacement) required to convert one string into another. Assume that the string and , the edit distance calculation formula is:
[0021]
[0022] in, Represents a string Before Characters and strings Before The edit distance between characters; is a component of the ternary operator, Indicates when Middle characters and Middle If the characters are not equal, take 1, if they are equal, take 0;
[0023] During context matching, the edit distance between the extracted field and the expected context-related string can be calculated. The smaller the distance, the higher the match between the field and the context, and the higher the extraction accuracy.
[0024] As a further improvement of the technical solution, in S2, the parsed data is mapped to the OPC UA information model to generate a structured JSON format, which specifically includes:
[0025] First, the process of mapping data to the OPC UA information model:
[0026] First, determine the OPC UA information model elements: The OPC UA information model contains multiple types of nodes. Before mapping data, it is necessary to clarify which model elements the parsed data corresponds to;
[0027] Secondly, establish mapping rules: according to the meaning of the data and the structure of the OPC UA information model, formulate detailed mapping rules;
[0028] Finally, the mapping operation is performed: according to the mapping rules, the parsed data is inserted into the corresponding node position of the OPC UA information model;
[0029] Second, the process of generating structured JSON format:
[0030] First, construct the JSON data structure: Based on the data mapped to the OPC UA information model, construct the data structure in JSON format;
[0031] Then, use the JSON serialization tool to convert the data into a JSON format string;
[0032] Third, in the process of mapping data to the OPC UA information model and generating structured JSON format, data validation and normalization processing are involved, including the following:
[0033] Data verification algorithm: Before mapping the data to the OPC UA information model, the parsed data needs to be verified to ensure the accuracy and integrity of the data; a checksum algorithm is used to verify whether errors occur during data transmission or processing; the checksum algorithm formula is:
[0034]
[0035] in, It is the sum of the ASCII code values of all characters in the data. is the ASCII code value of the character, Is the data string to be verified, is the length of the string, is the index of the character sequence number;
[0036] Data normalization algorithm: In order to ensure the consistency of data when mapping and generating JSON format, it may be necessary to normalize the data.
[0037] As a further improvement of the technical solution, in S3, the AI-driven intelligent classification and binding process, for the optimization and training of the long short-term memory network LSTM model, includes:
[0038] Before training the LSTM model, a large amount of string data related to the old system was collected and labeled. The labeling work was done by professionals, who marked the strings into different categories according to the meaning they represent.
[0039] When training the LSTM model, the labeled data is divided into training set, validation set and test set according to a certain ratio; the training set is used for model learning and parameter adjustment, the validation set is used to monitor overfitting during training, and the test set is used to evaluate the final performance of the model;
[0040] During the testing phase, the reserved test set data is input into the trained model. By comparing the model's prediction results with the actual annotation results, the model's accuracy, recall rate, F1 value and other indicators are evaluated. If the test results do not meet the requirements, the model parameters are further adjusted or the training data is increased and retraining is performed.
[0041] As a further improvement of the technical solution, in S3, the optimized long short-term memory network LSTM model includes:
[0042] Input layer: The embedding layer vectorizes the string. In the embedding layer of the input layer, we no longer simply vectorize the string, but choose a more appropriate embedding method based on the characteristics of the data.
[0043] Hidden layer: Bidirectional LSTM captures contextual features; the hidden layer uses bidirectional LSTM to process the input sequence from the forward and reverse directions respectively; the forward LSTM processes from the beginning of the string backward in sequence, while the reverse LSTM processes from the end forward; in addition, layer normalization operations are added between or after the bidirectional LSTM layers to stabilize the training process and accelerate model convergence; layer normalization normalizes each hidden unit of each sample, and the calculation formula is:
[0044]
[0045] in, is the normalized sample, is the input data, is the mean, is the variance, is a small constant that prevents division by zero, and is a learnable parameter;
[0046] Output layer: Softmax classifier labels the equipment level and measurement point type; in order to improve the accuracy of classification, the Softmax function is improved; in the calculation of Softmax, the temperature parameter (Temperature) is introduced to adjust the distribution of Softmax output probability; the Softmax function was originally , after introducing the temperature parameter T, it becomes ; When T is larger, the output probability distribution is more uniform; when T is smaller, the probability distribution is more concentrated on the category with the highest probability;
[0047] Hyperparameters: learning rate, batch size, dropout;
[0048] For the training process of the long short-term memory network LSTM model, its training method can be optimized as follows:
[0049] First, data enhancement strategy: when training data is limited, data enhancement technology is used to expand the data set;
[0050] Second, adaptive learning rate adjustment: During the training process, an adaptive learning rate algorithm is used. In the early stage of training, the learning rate is large and the model adjusts parameters quickly. As the training progresses, the learning rate is adaptively adjusted according to the gradient changes of each parameter to avoid oscillation near the local optimal solution.
[0051] Third, model fusion: train multiple LSTM models with different initializations and then perform model fusion.
[0052] As a further improvement of the technical solution, in S3, the AI-driven intelligent classification and binding process specifically includes the following steps:
[0053] S3.1. Data collection and preprocessing:
[0054] S3.1.1, Data collection: Collect a large amount of string data related to the old system, which must cover the entire operation process of the old system;
[0055] S3.1.2 Data cleaning: Remove noise, duplicate values and missing values from the data; for missing values, you can use the mean, median or mode to fill them, or use machine learning algorithms for prediction and filling;
[0056] S3.1.3 Data standardization: converting data into a uniform scale. Data standardization methods include normalization and standardization.
[0057] S3.2, Feature Engineering:
[0058] S3.2.1. Feature selection: Select the most representative and discriminative features from the original features, reduce the feature dimension, and improve the efficiency and performance of the model;
[0059] S3.2.2, Feature extraction: Generate new and more expressive features by transforming or combining original features;
[0060] S3.3, Classification model selection and training:
[0061] S3.3.1. Model selection: According to the characteristics of the data and the requirements of the task, the long short-term memory network LSTM model is selected;
[0062] S3.3.2. Use the training data to train the selected model and adjust the parameters of the model by optimizing the objective function;
[0063] S3.4, Binding rules formulation and implementation:
[0064] S3.4.1. Rule formulation: formulate reasonable binding rules based on business needs and classification results;
[0065] S3.4.2, Rule execution: Match the classification results with the binding rules to achieve data binding.
[0066] As a further improvement of the technical solution, in S4, data integration and exception handling specifically include the following steps:
[0067] S4.1. Data Integration:
[0068] S4.1.1. Data extraction:
[0069] Data source identification: Identify the data sources that need to be integrated, including data sources in different formats and locations;
[0070] Data extraction: Use appropriate tools and methods to extract data based on the type of data source;
[0071] (1) Database data extraction: Use SQL query statements to extract the required data from the database;
[0072] (2) File data extraction: For CSV files, you can use Python's pandas library to read them;
[0073] S4.1.2 Data conversion:
[0074] (1) Data cleaning: remove noise, duplicate values, and missing values from the data;
[0075] (2) Handling missing values: You can use the mean, median, or mode to fill missing values, or delete rows containing missing values;
[0076] (3) Data standardization: converting data into a unified scale, methods include normalization and standardization;
[0077] (4) Data encoding: convert categorical variables into numerical variables. Methods include one-hot encoding and label encoding.
[0078] S4.1.3, Data loading:
[0079] Load the transformed data into the target data store;
[0080] S4.2. Exception handling:
[0081] S4.2.1. Anomaly Detection:
[0082] Statistical anomaly detection: detect outliers by calculating the statistical characteristics of the data; common methods include the Z-score method;
[0083] Machine learning-based anomaly detection: anomaly detection using machine learning algorithms;
[0084] S4.2.2, Exception handling:
[0085] Delete outliers: directly delete detected outliers;
[0086] Correct outliers: Replace outliers with mean, median or other reasonable values;
[0087] Mark outliers: Mark outliers in the data for subsequent analysis;
[0088] Through the above steps, data integration and exception handling can be achieved to ensure data quality and consistency.
[0089] A second object of the present invention is to provide a system data acquisition system based on the OPC UA protocol, which is used to implement the steps of the above-mentioned system data acquisition method based on the OPC UA protocol, including:
[0090] The communication configuration unit is used to deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes, and use the OPC UA standard client to configure the client to ensure communication stability;
[0091] Data parsing and semantic modeling unit, including string parsing module and semantic modeling module; string parsing module extracts key fields based on regular expression and context matching algorithm; semantic modeling module is used to map the parsed data to OPC UA information model and generate structured JSON format;
[0092] Model optimization unit, used to optimize the long short-term memory network LSTM model to achieve AI-driven intelligent classification and binding;
[0093] Data integration and exception handling unit, including data conversion module and exception monitoring module; the data conversion module converts OPC UA data into the target platform format through the development adapter, and supports direct writing to the time series database; the exception monitoring module is used to detect abnormal situations in real time and trigger the reconnection mechanism or data retransmission;
[0094] The distributed architecture design unit adopts Kubernetes containerized deployment to achieve elastic expansion of computing nodes, and sets up a data cache layer to cope with high concurrency scenarios to ensure low-latency response.
[0095] A third object of the present invention is to provide a computer device for loading the above-mentioned system data acquisition system based on the OPC UA protocol, the device comprising a processor, a memory, and a computer program stored in the memory and running on the processor, and the processor is used to implement the steps of the above-mentioned system data acquisition method based on the OPC UA protocol when executing the computer program.
[0096] A fourth object of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned system data acquisition method based on the OPC UA protocol are implemented.
[0097] Compared with the prior art, the present invention has the following beneficial effects:
[0098] 1. The system data acquisition method, system and storage medium based on the OPC UA protocol successfully solved the data island problem caused by the unfriendly interface of the old system, realized the efficient integration and sharing of equipment data, and greatly improved the utilization value of data; through the precise analysis and classification of string information by AI technology, and the intelligent data binding algorithm, the accuracy and completeness of the equipment tree and measurement point information were significantly improved, providing a reliable data foundation for equipment management and operation status analysis; the automatic collection and processing of data greatly reduced manual intervention, reduced labor costs and the possibility of human errors, and improved work efficiency;
[0099] 2. The system data collection method, system and storage medium based on the OPC UA protocol integrate the data of the old system into the equipment management platform, so that the enterprise can fully and real-time understand the equipment operation status, timely discover potential faults and problems, help to take preventive measures in advance, reduce equipment downtime, improve equipment reliability and stability, thereby ensuring the continuity and stability of production; this solution has good scalability and adaptability, can meet the future needs of enterprise equipment management and information development, and provide strong support for the digital transformation and intelligent development of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 is an exemplary method flow chart of the present invention;
[0101] Figure 2 This is an exemplary electronic computer product structure diagram in the present invention. DETAILED DESCRIPTION
[0102] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0103] Example 1
[0104] like Figure 1 As shown, this embodiment provides a system data collection method based on the OPC UA protocol, which combines the OPC UA protocol with AI technology (wherein, for the AI model, a pre-trained language model (such as BERT) can be used to replace the operation of zero-sample classification to reduce the dependence on labeled data), to realize the automatic collection and analysis of the old system data and the automatic generation of the device tree, measurement points and corresponding configuration configuration, thereby forming an automatic analysis and integration of the old system data; specifically, the following steps are included:
[0105] S1. OPC UA communication layer configuration: Deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes; configure the OPC UA standard client (such as UA Expert) on the client, set security policies (such as Sign&Encrypt) and session timeout mechanism.
[0106] S2, Data parsing and semantic modeling: String parsing based on regular expressions and context matching algorithms, extracting key fields, and mapping the parsed data to the OPC UA information model to generate a structured JSON format;
[0107] In this step, data parsing is performed by setting a string parsing module, and the string parsing module performs string parsing specifically including the following steps:
[0108] S2.1. Data preprocessing: Before parsing the string, preprocess the original string data collected from the old system; remove redundant information in the string, such as spaces, special symbols (without affecting the meaning of the data), etc., and convert the string into a unified format for subsequent processing. For example, process "First-stage compressor: vibration amplitude = 10mm / s" into "First-stage compressor: vibration amplitude = 10mm / s";
[0109] S2.2. Define regular expression rules: Define corresponding regular expression patterns according to the key fields to be extracted, such as device level, measurement point attributes, units, etc.; for example:
[0110] For the device level, assuming that its naming rule starts with a Chinese character, followed by numbers, letters, or underscores, the regular expression is defined as ; This expression means matching a string that starts with a Chinese character and is followed by zero or more numbers, letters, or underscores;
[0111] For the point attributes, if they are composed of letters, numbers, and underscores, and their length is within a certain range (assuming 3-20 characters), define the regular expression as ;
[0112] For units, if common units are composed of letters, numbers, and slashes, define the regular expression as ;
[0113] S2.3. Use regular expressions to extract preliminary fields: Use the defined regular expressions to match the preprocessed strings; use the regular expression matching function in the programming language (such as the re.findall function in Python) to extract the string fragments that match the regular expression pattern; take "first-stage compressor: vibration amplitude = 10mm / s" as an example, use the above regular expression to extract separately, and you may get "first-stage compressor" (equipment level), "vibration amplitude" (measurement point attribute), "mm / s" (unit); but at this time, there may be inaccurate or incomplete extraction, such as the string "first-stage compressor_standby: vibration amplitude = 10mm / s, temperature = 25℃", only relying on the regular expression may mistakenly extract "first-stage compressor_standby: vibration amplitude" as the equipment level;
[0114] Among them, the principle of the regular expression matching algorithm is as follows: in most programming languages, regular expression matching is implemented based on the principle of finite automata; taking Python's re module as an example, when the re.findall(pattern, string) function is called, a finite automaton corresponding to the regular expression pattern will be built internally, and then starting from the starting position of the string string, it will match each character one by one according to the state transition rules of the automaton; if the automaton state sequence corresponding to the entire regular expression is successfully traversed during the matching process, a matching string fragment is found; there is no direct calculation formula in this process, but in principle it can be understood as a state transition function δ(state, char), where state represents the current state of the automaton, char represents the input character, and the δ function determines the automaton to transfer to the next state based on the current state and the input character; when the automaton starts from the initial state and reaches the terminal state after a series of state transitions, it means that a match has been found;
[0115] S2.4. Context matching to optimize extraction results: Introduce a context matching algorithm to improve the accuracy of extraction. Taking the device level as an example, after extracting the preliminary device level string, make judgments and corrections based on its context information. If the extracted device level string is followed by “:” and the “:” is followed by a string related to the measurement point attribute, then the extraction result may be correct. If this is not the case, further analysis is required. The extraction of measurement point attributes and units is similar, and verification and adjustment are performed based on their positional relationship in the string and other information that appears before and after. For example, when judging the unit, if a string suspected to be a unit is preceded by a numerical value and there are no other non-numeric or non-unit related characters between the two, then the string is more likely to be a unit. Through this context matching method, the results of regular expression extraction are optimized to obtain more accurate key fields.
[0116] Among them, in S2.4, the context matching algorithm uses a string similarity algorithm to assist in determining whether the extracted fields are accurate during application. The Edit Distance algorithm, also known as the Levenshtein Distance, is used. The Edit Distance refers to the minimum number of editing operations (insertion, deletion, and substitution) required to convert two strings from one to another. Assuming that the string and , the edit distance calculation formula is:
[0117]
[0118] in, Represents a string Before Characters and strings Before The edit distance between characters; It is a component of the ternary operator. The syntax of the ternary operator is "conditional expression value1:value2"; Indicates when Middle characters and Middle When the characters are not equal, it takes 1, and when they are equal, it takes 0; this is used to determine the impact of whether the current character matches or not on the distance value when calculating the edit distance;
[0119] During context matching, the edit distance between the extracted field and the expected context-related string can be calculated. The smaller the distance, the higher the match between the field and the context, and the higher the accuracy of the extraction. For example, when judging whether the extracted device level is correct, the edit distance between the extracted device level string and other known device-related strings in the row where the string is located is calculated. If the distance is less than a certain threshold (such as 3), the extracted device level is considered to be reasonable; otherwise, the extraction result may need to be rechecked or adjusted.
[0120] Furthermore, in S2, the parsed data is mapped to the OPC UA information model to generate a structured JSON format, which specifically includes:
[0121] First, the process of mapping data to the OPC UA information model:
[0122] First, determine the OPC UA information model elements: The OPC UA information model contains multiple types of nodes, such as objects, variables, methods, etc. Before mapping data, it is necessary to clarify which model elements the parsed data corresponds to. For example, device level information may be mapped to the "DeviceType" object node, and measurement point attributes and corresponding values may be mapped to the variable node under the object.
[0123] Secondly, establish mapping rules: according to the meaning of the data and the structure of the OPC UA information model, formulate detailed mapping rules; assuming that the parsed data contains the device level "first-level compressor", the measurement point attribute "vibration amplitude" and the value "10mm / s"; the mapping rules can be set as follows: "first-level compressor" creates a "DeviceType" object node named "first-level compressor"; "vibration amplitude" is used as a variable node name under the object, and "10mm / s" is the value of the variable;
[0124] Finally, perform the mapping operation: according to the mapping rules, insert the parsed data into the corresponding node position of the OPC UA information model; in actual programming, you can use the relevant OPC UA development library to implement this operation; for example, in the Java environment, you can use the Eclipse Milo library; complete the mapping by creating the corresponding node object and setting its properties and values;
[0125] Second, the process of generating structured JSON format:
[0126] First, construct the JSON data structure: Based on the data mapped to the OPC UA information model, construct the data structure in JSON format;
[0127] Then, use the JSON serialization tool: In most programming languages, there are mature JSON serialization libraries to convert data into JSON format strings; for example, in Python, you can use the json library;
[0128] Third, in the process of mapping data to the OPC UA information model and generating structured JSON format, data validation and normalization processing are involved, including the following:
[0129] Data verification algorithm: Before mapping the data to the OPC UA information model, the parsed data needs to be verified to ensure the accuracy and integrity of the data; for example, a checksum algorithm can be used to verify whether errors occur during data transmission or processing. A simple checksum algorithm can be to sum the ASCII code values of all characters in the data and then compare it with the preset checksum value; the checksum algorithm formula is:
[0130]
[0131] in, It is the sum of the ASCII code values of all characters in the data. is the ASCII code value of the character, Is the data string to be verified, is the length of the string, is the index of the character sequence number;
[0132] Data normalization algorithm: In order to ensure the consistency of data when mapping and generating JSON format, it may be necessary to normalize the data; for example, unify all device level names into uppercase or lowercase; in Python, this can be achieved using the lower() or upper() method of the string, which can be regarded as a simple data normalization algorithm.
[0133] S3. AI-driven intelligent classification and binding: Optimize and train the long short-term memory network (LSTM) model. For the identified and classified information, use the intelligent data binding algorithm to accurately associate the measurement points with the corresponding equipment level. This process requires comprehensive consideration of multiple factors such as the type, location, and function of the equipment to ensure the accuracy of the binding.
[0134] In this step, the AI-driven intelligent classification and binding process is aimed at the optimization and training of the long short-term memory network LSTM model, including:
[0135] Before training the LSTM model, a large amount of string data related to the old system needs to be annotated. The annotation work is done by professionals, who mark the strings into different categories such as equipment level and measurement point attributes according to the meaning they represent. For example, "compressor vibration monitoring point" is annotated as a measurement point attribute category, and "primary compression device" is annotated as an equipment level category.
[0136] When training the LSTM model, the labeled data is divided into training set, validation set and test set according to a certain ratio; the training set is used for model learning and parameter adjustment, the validation set is used to monitor overfitting during training, and the test set is used to evaluate the final performance of the model; the model continuously receives input labeled data, learns the characteristics and patterns of the string, and adjusts the internal weights and biases to improve the accuracy of prediction;
[0137] In the testing phase, the reserved test set data is input into the trained model. By comparing the model's prediction results with the actual annotation results, the model's accuracy, recall rate, F1 value and other indicators are evaluated. If the test results do not meet the requirements, the model parameters are further adjusted or the training data is increased and retrained.
[0138] The indicator calculation process includes:
[0139] Accuracy: Accuracy refers to the ratio of the number of samples correctly predicted by the model to the total number of samples, reflecting the model's ability to correctly judge all samples; the calculation formula is:
[0140]
[0141] Recall: Recall is also called recall rate, which refers to the ratio of the number of positive samples correctly predicted by the model to the actual number of positive samples, and measures the coverage of the model for positive samples. The calculation formula is:
[0142]
[0143] F1-Score: F1-Score is the harmonic mean of accuracy and recall, which comprehensively considers the accuracy and recall of the model and can more comprehensively evaluate the model performance. The calculation formula is:
[0144]
[0145] in, is the accuracy, is the recall rate, is the F1 value, TP (True Positive) represents true positive examples, that is, the number of samples predicted by the model to be positive and actually positive; TN (True Negative) represents true negative examples, that is, the number of samples predicted by the model to be negative and actually negative; FP (False Positive) represents false positive examples, that is, the number of samples predicted by the model to be positive but actually negative; FN (False Negative) represents false negative examples, that is, the number of samples predicted by the model to be negative but actually positive;
[0146] Evaluation indicator fusion: In practical applications, a single indicator may not be able to fully evaluate the performance of the model; accuracy, recall, F1 value and other indicators (such as precision, Matthews correlation coefficient (MCC), etc.) can be combined; precision focuses on the proportion of samples predicted to be positive that are actually positive, and the calculation formula is ; The Matthews correlation coefficient comprehensively considers true positives, true negatives, false positives, and false negatives, and can more comprehensively reflect the performance of the model on an unbalanced data set. The calculation formula is ; Assign weights to different indicators according to specific business needs and build comprehensive evaluation indicators to more accurately evaluate model performance.
[0147] Furthermore, in S3, the optimized long short-term memory network LSTM model includes:
[0148] Input layer: The embedding layer vectorizes the string. In the embedding layer of the input layer, we no longer need to simply vectorize the string, but choose a more appropriate embedding method based on the characteristics of the data. For strings related to equipment and measurement points in text data, we use pre-trained word vector models, such as Word2Vec or GloVe. These pre-trained models have learned rich semantic information on large-scale corpora and can better capture the relationship between words in strings. Taking "first-level compressor" as an example, through pre-trained word vector embedding, it can not only be converted into a numerical vector, but also reflect the semantic association between "compressor" and other related equipment vocabulary in the vector space, so that the model can better understand and distinguish different equipment levels in subsequent processing.
[0149] Hidden layer: Bidirectional LSTM captures contextual features; the hidden layer uses bidirectional LSTM to process the input sequence from the forward and reverse directions respectively; the forward LSTM processes from the beginning of the string backward in sequence, while the reverse LSTM processes from the end forward, which enables the model to capture the contextual information before and after the string at the same time; when processing equipment measurement point information, such as the string "[measurement point attribute] of [measurement point location] at [value at [time point]", the forward LSTM can learn the association information between the measurement point location and the attribute, and the reverse LSTM can capture the relationship between the attribute and the time point and the value; the output of the bidirectional LSTM is spliced to provide more comprehensive features for subsequent classification; in addition, between or after the bidirectional LSTM layers, a layer normalization operation is added to stabilize the training process and accelerate model convergence; layer normalization normalizes each hidden unit of each sample, and the calculation formula is:
[0150]
[0151] in, is the normalized sample, is the input data, is the mean, is the variance, is a small constant that prevents division by zero, and is a learnable parameter;
[0152] Output layer: Softmax classifier labels the equipment level and measurement point type; in order to improve the accuracy of classification, the Softmax function is improved; in the calculation of Softmax, the temperature parameter (Temperature) is introduced to adjust the distribution of Softmax output probability; the Softmax function was originally , after introducing the temperature parameter T, it becomes When T is larger, the output probability distribution is more uniform; when T is smaller, the probability distribution is more concentrated on the category with the highest probability; in the early stage of training, set a larger T value to allow the model to learn a wider classification boundary; in the later stage of training, gradually reduce the T value to make the classification results of the model more focused;
[0153] Hyperparameters: learning rate (0.001), batch size (64), dropout (0.2); hyperparameter selection and optimization include:
[0154] (1) Hyperparameter search range adjustment: For hyperparameters such as learning rate (currently set to 0.001), batch size (currently set to 64), and Dropout (currently set to 0.2), we can no longer limit ourselves to fixed values, but determine a more reasonable search range through experiments; for example, the learning rate can be searched in the range of [0.0001, 0.01], the batch size in the range of [32, 128], and the Dropout in the range of [0.1, 0.5];
[0155] (2) Adopt a more efficient hyperparameter tuning algorithm: Use the Bayesian optimization algorithm instead of a simple grid search or random search. Bayesian optimization builds a probabilistic model of the objective function and predicts the next optimal hyperparameter combination based on existing experimental results. This can more efficiently find the optimal hyperparameters and reduce the waste of computing resources and time.
[0156] In addition, for the training process of the long short-term memory network LSTM model, its training method can be optimized as follows:
[0157] First, data enhancement strategy: when training data is limited, data enhancement technology is used to expand the data set; synonym replacement, random insertion and deletion of words are performed on the string data related to equipment and measurement points; for example, "vibration amplitude" can be replaced with "vibration amplitude"; "large" is randomly inserted into "first-stage compressor" to become "first-stage large compressor"; through these operations, more training samples are generated to increase data diversity and improve the generalization ability of the model;
[0158] Second, adaptive learning rate adjustment: During the training process, an adaptive learning rate algorithm is used, such as Adagrad, Adadelta or Adam. Taking the Adam algorithm as an example, it combines the advantages of momentum and adaptive learning rate. In the early stage of training, the learning rate is large and the model adjusts parameters quickly. As the training progresses, the learning rate is adaptively adjusted according to the gradient changes of each parameter to avoid oscillation near the local optimal solution. The update formula using the Adam algorithm is:
[0159]
[0160]
[0161]
[0162]
[0163]
[0164] in, and are the first-order moment estimate and the second-order moment estimate of the gradient, and is the decay rate of the moment estimate (usually , ), and is the corrected moment estimate, is the learning rate, is a small constant that prevents division by zero, yes The model parameters at time t, yes The gradient of the moment;
[0165] Third, model fusion: train multiple LSTM models with different initializations, and then perform model fusion. You can use a simple average fusion method to average the prediction results of multiple models. You can also use weighted fusion to assign different weights to each model according to its performance on the validation set. For example, a model with a higher accuracy on the validation set is given a larger weight, and model fusion is used to improve the stability and accuracy of the classification results.
[0166] In addition, in S3, the AI-driven intelligent classification and binding process specifically includes the following steps:
[0167] S3.1. Data collection and preprocessing:
[0168] S3.1.1, Data collection: Collect a large amount of string data related to the old system, which must cover the entire operation process of the old system;
[0169] S3.1.2 Data cleaning: remove noise, duplicate values and missing values from the data; for missing values, you can use the mean, median or mode to fill, or use machine learning algorithms for prediction and filling; for example, when processing numerical data, if a feature has missing values, you can use the mean of the feature to fill;
[0170] S3.1.3 Data standardization: Convert data to a uniform scale. Data standardization methods include normalization and standardization. Normalization scales data to a uniform scale. The interval is:
[0171]
[0172] in, is the normalized data, is the original data, and are the minimum and maximum values of the data respectively;
[0173] Standardization converts the data into a distribution with a mean of 0 and a standard deviation of 1. The formula is:
[0174]
[0175] in, is the standardized data, is the mean of the data, is the standard deviation of the data;
[0176] S3.2, Feature Engineering:
[0177] S3.2.1. Feature selection: Select the most representative and discriminative features from the original features, reduce the feature dimension, and improve the efficiency and performance of the model; common feature selection methods include correlation analysis, chi-square test, mutual information, etc.; taking correlation analysis as an example, calculate the Pearson correlation coefficient between the feature and the target variable, the formula is:
[0178]
[0179] in, is the Pearson correlation coefficient, and are the features and target variables, respectively. Sample values, is the sample size, and are the means of the feature and target variables respectively;
[0180] S3.2.2 Feature extraction: Generate new and more expressive features by transforming or combining original features. For example, in image classification, principal component analysis (PCA) can be used for feature extraction. The core idea is to find the principal components of the data so that the variance of the data on these principal components is maximized. The calculation steps of PCA are as follows:
[0181] Calculate the covariance matrix of the data :
[0182]
[0183] Covariance matrix Perform eigenvalue decomposition and obtain the eigenvalue and the corresponding eigenvector ;
[0184] Before selection The eigenvectors corresponding to the largest eigenvalues form the projection matrix ;
[0185] The original data Projection to the projection matrix On top, we get new features ;
[0186] S3.3, Classification model selection and training:
[0187] S3.3.1. Model selection: According to the characteristics of the data and the requirements of the task, the long short-term memory network LSTM model is selected;
[0188] S3.3.2. Use the training data to train the selected model and adjust the model parameters by optimizing the objective function. Common objective functions include cross entropy loss function, mean square error loss function, etc. Taking the cross entropy loss function as an example, for the binary classification problem, the formula is:
[0189]
[0190] in, To optimize the target value, It is The true label of each sample (0 or 1), The model predicts The probability that a sample is a positive class;
[0191] S3.4, Binding rules formulation and implementation:
[0192] S3.4.1. Rule formulation: formulate reasonable binding rules based on business needs and classification results;
[0193] S3.4.2, Rule execution: Match the classification results with the binding rules to implement data binding; in actual applications, conditional judgment statements or rule engines can be used to execute binding rules.
[0194] S4, Data integration and exception handling: Design data conversion modules and exception monitoring functions, and integrate processed data with the device management platform;
[0195] In this step, data integration and exception handling specifically include the following steps:
[0196] S4.1. Data Integration:
[0197] S4.1.1. Data extraction:
[0198] Data source identification: Identify the data sources that need to be integrated, including data sources in different formats (such as CSV, JSON, database tables, etc.) and different locations (local file system, remote server, etc.); you can use configuration files to record the connection information of each data source, such as the database connection string, file path, etc.;
[0199] Data extraction: Use appropriate tools and methods to extract data based on the type of data source;
[0200] (1) Database data extraction: Use SQL query statements to extract the required data from the database; for example, use Python's pymysql library to connect to the MySQL database and extract data;
[0201] (2) File data extraction: For CSV files, you can use Python's pandas library to read them;
[0202] S4.1.2 Data conversion:
[0203] (1) Data cleaning: remove noise, duplicate values, and missing values from the data;
[0204] Remove duplicate values: Use pandas’ drop_duplicates() method to remove duplicate rows;
[0205] (2) Handling missing values: You can use the mean, median, or mode to fill missing values, or delete rows containing missing values;
[0206] (3) Data standardization: converting data to a unified scale. Common methods include normalization and standardization.
[0207] (4) Data encoding: convert categorical variables into numerical variables. Common methods include one-hot encoding and label encoding.
[0208] One-hot encoding: Use pandas’ get_dummies() method for one-hot encoding;
[0209] Label encoding: Use sklearn's LabelEncoder for label encoding;
[0210] S4.1.3, Data loading:
[0211] Load the transformed data into the target data storage, such as a database, data warehouse, etc.
[0212] S4.2. Exception handling:
[0213] S4.2.1. Anomaly Detection:
[0214] Statistical anomaly detection: detect outliers by calculating the statistical characteristics of the data (such as mean, standard deviation, etc.); the common method is the Z-score method; the Z-score formula is:
[0215]
[0216] in, is the Z eigenvalue, is a data point, is the mean of the data, is the standard deviation of the data. In general, when When it is greater than a certain threshold (such as 3), the data point is considered an outlier;
[0217] Machine learning-based anomaly detection: Use machine learning algorithms (such as Isolation Forest, One-Class SVM, etc.) for anomaly detection;
[0218] S4.2.2, Exception handling:
[0219] Delete outliers: directly delete detected outliers;
[0220] Correct outliers: Replace outliers with mean, median or other reasonable values;
[0221] Mark outliers: Mark outliers in the data for subsequent analysis;
[0222] Through the above steps, data integration and exception handling can be achieved to ensure data quality and consistency.
[0223] S5. Distributed architecture design: Kubernetes containerized deployment is used to achieve elastic expansion of computing nodes; a data cache layer (Redis) is set up to cope with high-concurrency scenarios to ensure low-latency response; during the entire data collection and processing process, a real-time monitoring and error handling mechanism is set up. When communication is interrupted, data is abnormal or other errors occur, an alarm can be issued in time and corresponding recovery measures can be taken to ensure the stability and reliability of the system.
[0224] In this step, local processing on edge computing nodes can also be used to replace the centralized cloud architecture to reduce network load.
[0225] This embodiment provides a system data acquisition system based on the OPC UA protocol, which is used to implement the steps of the above-mentioned system data acquisition method based on the OPC UA protocol, including:
[0226] The communication configuration unit is used to deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes, and use the OPC UA standard client to configure the client to ensure communication stability;
[0227] Data parsing and semantic modeling unit, including string parsing module and semantic modeling module; the string parsing module extracts key fields such as equipment level (such as "first-stage compressor"), measurement point attributes (such as "vibration amplitude"), and units (such as "mm / s") based on regular expressions and context matching algorithms; the semantic modeling module is used to map the parsed data to the OPC UA information model (such as "DeviceType" and "SensorType") to generate a structured JSON format;
[0228] Model optimization unit, used to optimize the long short-term memory network LSTM model to achieve AI-driven intelligent classification and binding;
[0229] The data integration and exception handling unit includes a data conversion module and an exception monitoring module. The data conversion module converts OPC UA data into the target platform format (such as MQTT, RESTful API) through the development adapter, and supports direct writing to time series databases (such as InfluxDB). The exception monitoring module is used to detect abnormalities such as communication interruption and data overrun in real time, and trigger a reconnection mechanism or data retransmission. This unit is used to convert the data format obtained by the OPC UA protocol into a format that can be recognized and processed by the device management platform, and at the same time process abnormal values and missing values in the data to ensure the quality and integrity of the data.
[0230] The distributed architecture design unit adopts Kubernetes containerized deployment to achieve elastic expansion of computing nodes, and sets up a data cache layer (Redis) to cope with high concurrency scenarios to ensure low-latency response.
[0231] In summary, the key points of this technical solution are:
[0232] (1) A method and system architecture for obtaining configuration data of legacy systems based on the OPC UA protocol. Through the effective application of the OPC UA protocol, stable communication and data acquisition with legacy systems can be achieved.
[0233] (2) Use AI technology to analyze string information and bind data using algorithms and models, and use AI technology to accurately analyze, classify, and bind device and measurement point information in string form;
[0234] (3) Based on the key steps and technical implementations in the entire data collection, processing and integration process, the data conversion and adaptation modules developed are used to ensure the seamless integration of the collected data with the device management platform.
[0235] like Figure 2 As shown, this embodiment also provides a computer device for loading the above-mentioned system data acquisition system based on the OPC UA protocol, and the device includes a processor, a memory, and a computer program stored in the memory and running on the processor.
[0236] The processor includes one or more processing cores. The processor is connected to the memory through a bus. The memory is used to store program instructions. When the processor executes the program instructions in the memory, the steps of the system data acquisition method based on the OPC UA protocol are implemented.
[0237] Optionally, the memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0238] In addition, this embodiment further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned system data acquisition method based on the OPC UA protocol are implemented.
[0239] Optionally, the present invention further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of the above-mentioned system data acquisition method based on the OPC UA protocol.
[0240] A person of ordinary skill in the art can understand that the process of implementing all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk or an optical disk, etc.
[0241] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A system data acquisition method based on the OPC UA protocol, characterized in that: By combining the OPC UA protocol with AI technology, the old system data can be automatically collected, parsed, and the device tree, measurement points, and corresponding configurations can be automatically generated, thereby forming an automated analysis and integration of the old system data. The specific steps include: S1. OPC UA communication layer configuration: Deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes; configure the OPC UA standard client on the client, and set the security policy and session timeout mechanism; S2, Data parsing and semantic modeling: String parsing based on regular expressions and context matching algorithms, extracting key fields, and mapping the parsed data to the OPC UA information model to generate a structured JSON format; S3, AI-driven intelligent classification and binding: Optimize and train the long short-term memory network LSTM model, and use the intelligent data binding algorithm to accurately associate the measurement points with the corresponding equipment level for the identified and classified information; S4, Data integration and exception handling: Design data conversion modules and exception monitoring functions, and integrate processed data with the device management platform; S5, Distributed architecture design: Use Kubernetes containerized deployment to achieve elastic expansion of computing nodes; set up a data cache layer to cope with high concurrency scenarios and ensure low-latency response; During the entire data collection and processing process, set up real-time monitoring and error handling mechanisms.
2. The system data acquisition method based on the OPC UA protocol according to claim 1, characterized in that: In S2, data parsing is performed by setting a string parsing module, and the string parsing module performs string parsing specifically including the following steps: S2.
1. Data preprocessing: Before parsing the string, preprocess the original string data collected from the old system; remove redundant information in the string and convert the string into a unified format; S2.
2. Define regular expression rules: define corresponding regular expression patterns according to the key fields to be extracted; S2.
3. Apply regular expressions to extract preliminary fields: Use the defined regular expressions to perform matching operations on the preprocessed strings; Through the regular expression matching function in the programming language, the string fragments that meet the regular expression pattern are extracted; S2.
4. Context matching to optimize extraction results: Introduce context matching algorithm to improve the accuracy of extraction; after extracting the preliminary string, make judgments and corrections based on its context information.
3. The system data acquisition method based on the OPC UA protocol according to claim 2, characterized in that: In S2.4, the context matching algorithm uses a string similarity algorithm to assist in determining whether the extracted fields are accurate during application, and uses an edit distance algorithm; the edit distance refers to the minimum number of edit operations required to convert one string into another, including insertion, deletion, and replacement; assuming that the string and , the edit distance calculation formula is: in, Represents a string Before Characters and strings Before The edit distance between characters; is a component of the ternary operator, Indicates when Middle characters and Middle If the characters are not equal, take 1, if they are equal, take 0; During context matching, the edit distance between the extracted field and the expected context-related string can be calculated. The smaller the distance, the higher the match between the field and the context, and the higher the extraction accuracy.
4. The system data acquisition method based on the OPC UA protocol according to claim 3 is characterized in that: In S2, the parsed data is mapped to the OPC UA information model to generate a structured JSON format, which specifically includes: First, the process of mapping data to the OPC UA information model: First, determine the OPC UA information model elements: The OPC UA information model contains multiple types of nodes. Before mapping data, it is necessary to clarify which model elements the parsed data corresponds to; Secondly, establish mapping rules: according to the meaning of the data and the structure of the OPC UA information model, formulate detailed mapping rules; Finally, the mapping operation is performed: according to the mapping rules, the parsed data is inserted into the corresponding node position of the OPC UA information model; Second, the process of generating structured JSON format: First, construct the JSON data structure: Based on the data mapped to the OPC UA information model, construct the data structure in JSON format; Then, use the JSON serialization tool to convert the data into a JSON format string; Third, in the process of mapping data to the OPC UA information model and generating structured JSON format, data validation and normalization processing are involved, including the following: Data verification algorithm: Before mapping the data to the OPC UA information model, the parsed data needs to be verified to ensure the accuracy and integrity of the data; a checksum algorithm is used to verify whether errors occur during data transmission or processing; the checksum algorithm formula is: in, It is the sum of the ASCII code values of all characters in the data. is the ASCII code value of the character, Is the data string to be verified, is the length of the string, Is the index of the character sequence number.
5. The system data acquisition method based on the OPC UA protocol according to claim 4 is characterized in that: In S3, the AI-driven intelligent classification and binding process is aimed at the optimization and training of the long short-term memory network LSTM model, including: Before training the LSTM model, a large amount of string data related to the old system was collected and labeled. The labeling work was done by professionals, who marked the strings into different categories according to the meaning they represent. When training the LSTM model, the labeled data is divided into training set, validation set and test set according to a certain ratio; the training set is used for model learning and parameter adjustment, the validation set is used to monitor overfitting during training, and the test set is used to evaluate the final performance of the model; During the testing phase, the reserved test set data is input into the trained model. The model's accuracy, recall, and F1 value indicators are evaluated by comparing the model's prediction results with the actual annotation results. If the test results do not meet the requirements, the model parameters are further adjusted or the training data is increased and retrained.
6. The system data acquisition method based on the OPC UA protocol according to claim 5, characterized in that: In S3, the optimized long short-term memory network LSTM model includes: Input layer: The embedding layer vectorizes the string. In the embedding layer of the input layer, we no longer simply vectorize the string, but choose a more appropriate embedding method based on the characteristics of the data. Hidden layer: Bidirectional LSTM captures contextual features; the hidden layer uses bidirectional LSTM to process the input sequence from the forward and reverse directions respectively; the forward LSTM processes from the beginning of the string backward in sequence, while the reverse LSTM processes from the end forward; in addition, layer normalization operations are added between or after the bidirectional LSTM layers to stabilize the training process and accelerate model convergence; layer normalization normalizes each hidden unit of each sample, and the calculation formula is: in, is the normalized sample, is the input data, is the mean, is the variance, is a small constant that prevents division by zero, and is a learnable parameter; Output layer: Softmax classifier labels the equipment level and measurement point type; in order to improve the accuracy of classification, the Softmax function is improved; in the calculation of Softmax, the temperature parameter is introduced to adjust the distribution of Softmax output probability; the Softmax function was originally , after introducing the temperature parameter T, it becomes ; When T is larger, the output probability distribution is more uniform; when T is smaller, the probability distribution is more concentrated on the category with the highest probability; Hyperparameters: including learning rate, batch size, and Dropout; For the training process of the long short-term memory network LSTM model, its training method can be optimized as follows: First, data enhancement strategy: when training data is limited, data enhancement technology is used to expand the data set; Second, adaptive learning rate adjustment: During the training process, an adaptive learning rate algorithm is used. In the early stage of training, the learning rate is large and the model adjusts parameters quickly. As the training progresses, the learning rate is adaptively adjusted according to the gradient changes of each parameter to avoid oscillation near the local optimal solution. Third, model fusion: train multiple LSTM models with different initializations and then perform model fusion.
7. The system data acquisition method based on the OPC UA protocol according to claim 6, characterized in that: In S3, the AI-driven intelligent classification and binding process specifically includes the following steps: S3.
1. Data collection and preprocessing: S3.1.1, Data collection: Collect a large amount of string data related to the old system, which must cover the entire operation process of the old system; S3.1.2 Data cleaning: Remove noise, duplicate values and missing values from the data; for missing values, you can use the mean, median or mode to fill them, or use machine learning algorithms for prediction and filling; S3.1.3 Data standardization: converting data into a uniform scale. Data standardization methods include normalization and standardization. S3.2, Feature Engineering: S3.2.
1. Feature selection: Select the most representative and discriminative features from the original features, reduce the feature dimension, and improve the efficiency and performance of the model; S3.2.2, Feature extraction: Generate new and more expressive features by transforming or combining original features; S3.3, Classification model selection and training: S3.3.
1. Model selection: According to the characteristics of the data and the requirements of the task, the long short-term memory network LSTM model is selected; S3.3.
2. Use the training data to train the selected model and adjust the parameters of the model by optimizing the objective function; S3.4, Binding rules formulation and implementation: S3.4.
1. Rule formulation: formulate reasonable binding rules based on business needs and classification results; S3.4.2, Rule execution: Match the classification results with the binding rules to achieve data binding.
8. The system data acquisition method based on the OPC UA protocol according to claim 7, characterized in that: In S4, data integration and exception handling specifically include the following steps: S4.
1. Data Integration: S4.1.
1. Data extraction: Data source identification: Identify the data sources that need to be integrated, including data sources in different formats and locations; Data extraction: Use appropriate tools and methods to extract data based on the type of data source; (1) Database data extraction: Use SQL query statements to extract the required data from the database; (2) File data extraction: For CSV files, you can use Python's pandas library to read them; S4.1.2 Data conversion: (1) Data cleaning: remove noise, duplicate values, and missing values from the data; (2) Handling missing values: You can use the mean, median, or mode to fill missing values, or delete rows containing missing values; (3) Data standardization: converting data into a unified scale, methods include normalization and standardization; (4) Data encoding: convert categorical variables into numerical variables. Methods include one-hot encoding and label encoding. S4.1.3, Data loading: Load the transformed data into the target data store; S4.
2. Exception handling: S4.2.
1. Anomaly Detection: Statistical anomaly detection: detect outliers by calculating the statistical characteristics of the data; the method includes the Z-score method; Machine learning-based anomaly detection: anomaly detection using machine learning algorithms; S4.2.2, Exception handling: Delete outliers: directly delete detected outliers; Correct outliers: replace outliers with mean, median or other reasonable values; Mark outliers: Mark outliers in the data for later analysis.
9. A system data acquisition system based on the OPC UA protocol, used to implement the steps of the system data acquisition method based on the OPC UA protocol according to any one of claims 1 to 8, characterized in that: include: The communication configuration unit is used to deploy the OPC UA server module in the old system, encapsulate the device tree and measurement point information as OPC UA nodes, and use the OPC UA standard client to configure the client to ensure communication stability; Data parsing and semantic modeling unit, including string parsing module and semantic modeling module; The string parsing module extracts key fields based on regular expressions and context matching algorithms; The semantic modeling module is used to map the parsed data to the OPC UA information model and generate a structured JSON format; Model optimization unit, used to optimize the long short-term memory network LSTM model to achieve AI-driven intelligent classification and binding; Data integration and exception handling unit, including data conversion module and exception monitoring module; the data conversion module converts OPC UA data into the target platform format through the development adapter, and supports direct writing to the time series database; the exception monitoring module is used to detect abnormal situations in real time and trigger the reconnection mechanism or data retransmission; The distributed architecture design unit uses Kubernetes containerized deployment to achieve elastic expansion of computing nodes, and sets up a data cache layer to cope with high-concurrency scenarios to ensure low-latency response; It also includes a computer device for loading the above-mentioned system data acquisition system based on the OPC UA protocol, the device includes a processor, a memory, and a computer program stored in the memory and running on the processor, and the processor is used to implement the steps of the system data acquisition method based on the OPC UA protocol as described in any one of claims 1-8 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the system data acquisition method based on the OPC UA protocol as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-source data perception, fusion and predictive analysis method and system for aircraft assembly line
CN118348918A
Practical training platform, evaluation method and equipment for aluminum electrolysis process based on multi-modal data
CN119538980A
Cited By
Automatic configuration method and system for DCS data acquisition interface of power plant
CN120871762A