Unification method based on data acquisition interfaces of different SCADA (Supervisory Control And Data Acquisition) production monitoring systems of sewage plant
Through protocol conversion middleware and device mapping mechanism, the proprietary protocols of devices from different manufacturers are converted into a unified format, which solves the data island problem of the SCADA system, realizes standardized data collection and integration, improves data quality and security, and ensures the intelligent management of sewage treatment plants.
Patent Information
- Application Number
- CN202510881129.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
SCADA systems provided by different manufacturers use their own proprietary communication protocols, resulting in incompatible data acquisition interfaces and the formation of data silos. In addition, there are compatibility issues between old systems and new equipment, making data difficult to manage uniformly and use efficiently, and posing security risks and network instability.
Through protocol conversion middleware and device mapping mechanism, the proprietary protocols of devices from different manufacturers are converted into a unified format. Support vector machine algorithm is used for data cleaning and feature extraction, and the corresponding relationship between devices and data is established. Combined with data normalization algorithm and encryption mechanism, standardized data collection and integration are achieved. Incremental adaptation module is used to solve the compatibility issues of old systems, and network status monitoring and adaptive protocol adjustment module are used to ensure the stability of data transmission.
It achieves standardized collection and integration of heterogeneous equipment data, solves the compatibility issues between old systems and new equipment, improves data quality and security, ensures continuous transmission and correct analysis of data in unstable network environments, and supports intelligent management of sewage treatment plants.
Smart Images

Figure CN120711091A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of monitoring, and in particular relates to a unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage plant. Background Art
[0002] SCADA production monitoring systems play a key role in the operation and management of sewage treatment plants. However, current SCADA systems provided by different manufacturers use their own proprietary communication protocols, resulting in incompatible data acquisition interfaces and the formation of data silos. Significant differences in data formats and transmission standards between heterogeneous devices make device integration and data interoperability difficult. Furthermore, legacy SCADA systems face compatibility issues with newer equipment and are unable to effectively identify newly collected data. Furthermore, data transmission faces security risks, and unstable network environments are prone to transmission interruptions and data loss. These issues severely restrict the unified management and efficient utilization of sewage plant production data. A unified data acquisition interface approach is urgently needed to address these technical challenges. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a unified method based on the data acquisition interface of different SCADA production monitoring systems of sewage treatment plants. By utilizing protocol conversion middleware and device mapping mechanism, the proprietary protocols of devices from different manufacturers are converted into a unified format, a corresponding relationship between devices and data is established, and standardized collection and integration of heterogeneous device data is realized.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is: A unified method based on the data acquisition interface of different SCADA production monitoring systems of sewage treatment plants, the steps are as follows: S1. Utilize pre-established protocol conversion middleware to parse and standardize the communication protocols of devices from different manufacturers. Addressing protocol inconsistencies, we convert multiple proprietary protocols into a unified communication format, obtain standardized data streams, and ultimately create a unified data set for subsequent processing.
[0005] S1.1. Using a pre-built middleware framework, the communication protocols of devices from different vendors are parsed and processed to obtain the initial protocol data stream. If the parsed protocol data streams have format differences, pre-set protocol conversion rules are used to convert the proprietary protocols into a unified format, generating standardized communication data. Based on the standardized communication data, data stream label normalization operations are performed to generate a unified communication data set for subsequent analysis. In this step, the middleware parses the device protocol (such as Modbus and OPC UA) to obtain the initial data stream. If there are format differences, the proprietary protocol is converted into a unified JSON format using preset rules. The key features are extracted simultaneously using the support vector machine (SVM) algorithm. The formula is: ; is the data stream feature vector (such as protocol type, data frequency), is the feature weight vector (obtained through training of historical protocol data), is the classification threshold. This formula is used to determine whether the data flow belongs to the target protocol type (output +1 / -1).
[0006] S1.2. If redundant or abnormal fields exist within the unified communications data set, clean it using a pre-defined filtering mechanism to determine a cleaned data set. A support vector machine algorithm is then used to extract features from the cleaned data set to obtain key communication feature values. Based on these extracted key communication feature values, a mapping relationship is constructed between data flow labels and subsequent analysis to determine the suitability of the data set. In this step, after filtering out redundant fields, the SVM feature extraction algorithm is applied to the cleaned data to construct a mapping relationship between data stream features and device types. For example, if the feature vector , it can be determined by formula calculation that it belongs to the PLC device (output + 1), thereby generating a standardized communication data set.
[0007] S1.3. If the suitability of the data set meets the preset threshold, a final standardized data stream is generated and the output processing is completed in a unified format.
[0008] S2. Based on the unified data set, a device mapping mechanism is used to associate the data streams of heterogeneous devices with preset device identifiers. To address the challenges of heterogeneous device integration, a correspondence between devices and data is established, the data source and format of each device are determined, and a device-identified data mapping table is generated.
[0009] S2.1. Build a unified data processing framework to capture data streams from heterogeneous devices and determine the initial data set. S2.2. Based on the preset device identifiers, the initial data set is classified using a mapping mechanism to generate device-associated data groups. If a data stream within the data group does not match the device identifier, the device category is determined by comparing the data source and format, and a preliminary mapping relationship is generated. S2.3. Based on the preliminary mapping relationships, obtain characteristic information about the data flow and determine whether there are any integration challenges. If the characteristic information does not meet the preset threshold, adjust the mapping mechanism to obtain an optimized mapping relationship. Through the optimized correspondence, an identification table between devices and data is constructed to obtain the clear data source and data format of each device, and determine whether the device association is completed; S2.4. Based on the contents of the identification table, use a support vector machine algorithm to verify the data stream, determine the accuracy of the data mapping, and generate the final device identification data mapping table; S2.5. Obtain the complete data integration results for heterogeneous devices through the final device identification data mapping table, determine whether it meets business requirements, and complete the comprehensive association of data and devices.
[0010] S3. Using the device-identified data mapping table, we implement a data normalization algorithm to unify the format of data fields from different sources and fill in missing values to obtain a standardized data structure and a standardized data set suitable for system analysis.
[0011] S3.1. Obtain an initial data set from multiple data sources using device identification and mapping tables. Perform a preliminary classification of data fields from different sources to obtain a classified field set. Based on this classified field set, perform a consistency check on the data fields using the collection standard. If the check reveals inconsistent field formats, apply a normalization algorithm to unify the formats and determine a consistent field data set. S3.2. Detect missing values in a consistent field data set. If missing values are detected, fill them using a pre-defined interpolation method to obtain a complete set of data fields. Based on this complete set of data fields, construct a standardized data structure and apply pre-defined mapping rules to associate the data fields with device identifiers, resulting in a structured, standardized data set. S3.3. Use a structured standard dataset to integrate data based on source differences. Standardize data from different sources using a unified coding method to determine the integrated standard dataset. For the classified field set, the collection standard is used to check the format consistency. If missing values are detected (such as flow meter data gaps), they are filled by linear interpolation. The formula is: ; The time to be supplemented The flow value, and is the known flow value 10 seconds before and after; For example, a flow meter When data is missing, it is known The flow rate is , At that time , substituting into the formula we get: ; S3.4. Perform data quality verification on the integrated canonical dataset. If any abnormal data is found during the verification, filter it using a preset threshold to obtain a final canonical dataset suitable for system analysis. Generate corresponding metadata descriptions for the final canonical dataset, recording the formatting of data fields and the gap filling process, to produce a complete dataset with descriptive information.
[0012] S4. Based on the standard data set, to address compatibility issues with legacy system upgrades, an incremental data adaptation module is used to adapt the newly collected data to the legacy system's interface, obtain a data stream compatible with the legacy system, and ensure that the adapted data can be recognized and stored by the legacy system.
[0013] S4.1. Analyze compatibility issues between legacy systems and newly collected data, build a data adaptation framework, obtain a preliminary interface matching solution, and ensure that the adaptation framework can handle data differences. S4.2. Based on the output of the adaptation framework, perform incremental adaptation on the newly collected data in layers to generate initial compatible data. Determine whether the initial compatible data meets the interface matching requirements. S4.3. If the initial compatible data does not meet the interface matching requirements, adjust the data fields using the preset rule base to obtain the adjusted compatible data and confirm that the adjusted data passes the interface verification. S4.4. For the adjusted compatible data, use data flow mapping technology to construct a data transmission path, obtain the stability index of the transmission path, and confirm that the transmission path can support the smooth transmission of data flow; S4.5. Apply the support vector machine algorithm to optimize the data stream based on the transmission path stability indicator, generate an optimized data stream, and determine whether the optimized data stream meets the recognition capability requirements. S4.6. If the optimized data stream meets the recognition capability requirements, transfer it to the legacy system's storage mechanism to obtain storage status feedback and determine whether the storage process is complete. S4.7. Generate data storage logs based on storage status feedback. Use log analysis tools to detect anomalies in the storage process. Obtain anomaly detection results to determine whether the legacy system can continue to operate stably.
[0014] S5. To ensure data transmission security through the adapted data stream, a data encryption protection mechanism is implemented. A symmetric encryption algorithm is used to encrypt the data. Prior to transmission, the data is encrypted with a key, and the encrypted data packet is obtained to securely transmit the data content.
[0015] S5.1. Based on the adjusted data stream, encrypt the original data using a symmetric encryption algorithm to generate a preliminary encrypted data unit, in accordance with transmission security and safety requirements. S5.2. Based on the initially encrypted data unit, perform a secondary encryption process on the data unit using a preset key encryption mechanism before transmission to obtain a data block with a higher encryption strength. If an anomaly is detected in the data block during the encryption process, the data block integrity is verified through the preset error checking mechanism to determine whether there is any data corruption or tampering; If the verification result shows that the data block integrity is normal, a data packaging tool is used to convert the data block into a data packet in a standard format and determine the transmission readiness status of the data packet; S5.3. By adapting the data packet to the transmission channel, obtaining the transmission protocol that matches the target channel and obtaining the adapted data packet structure; According to the adapted data packet structure, the preset transmission security policy is used to perform a final security check on the data packet to determine whether it meets the requirements of secure transmission; S5.4. If the detection result meets the preset security threshold, the data packet is transmitted through the target channel to obtain the final secure transmission data content.
[0016] S6. Based on the encrypted data packets and to ensure stable data transmission, a network status monitoring module is used to monitor the impact of network conditions in real time during transmission. If the network bandwidth falls below a preset threshold, the data packet transmission rate is reduced and an adjusted transmission strategy is obtained to ensure that the data packet can be continuously transmitted in an unstable network.
[0017] S6.1. The network status monitoring module continuously monitors network conditions, obtains current bandwidth and latency information, and determines whether the network is stable. If the collected bandwidth data is lower than the preset threshold, the rate adjustment mechanism is triggered to reduce the transmission rate of the data packet and obtain the adjusted transmission parameters; S6.2. Reconfigure the packet distribution logic based on the adjusted transmission parameters, obtain the updated transmission strategy, and determine how to group the packets. S6.3. Based on the updated transmission policy, use the packet prioritization method to determine the transmission order of high-priority packets and determine whether the transmission queue arrangement is reasonable. If the transmission queue arrangement complies with the preset rules, the data packet transmission is executed according to the adjusted strategy, and real-time feedback data is obtained during the transmission process; S6.4 Analyze real-time feedback data and use the support vector machine algorithm to predict network status trends, obtain prediction results, and determine whether to adjust the transmission strategy; S6.5. If the network status is likely to deteriorate further based on the prediction results, perform a secondary optimization of the transmission strategy to determine the final transmission plan.
[0018] S7. To achieve video surveillance fusion through the adjusted transmission strategy, a multi-source data integration framework is employed to align the timestamps and associate the content of video surveillance data with production data. This generates a fused multimodal data stream and a comprehensive dataset containing both video and production information.
[0019] S7.1. Complete preliminary data collection by building a multi-source data integration framework to obtain raw streams from video surveillance and production data, generating an unprocessed mixed data set. S7.2. Based on the mixed data set, synchronize the time information of the video surveillance data and production data using a timestamp alignment method. If the timestamp deviation exceeds a preset threshold, interpolate the data to determine the aligned synchronized data set. S7.3. Based on the synchronized data sets and content association requirements, extract keyframe features from the video surveillance data and event records from the production data. Use a support vector machine algorithm to perform feature matching, identify highly correlated data pairs, and form a correlated data set. S7.4. Generate a multimodal data stream based on the associated data set, fuse the feature information of the video surveillance data and the production data, obtain structured multimodal stream data, and obtain a preliminary fused data stream. S7.5. Based on the initial fused data stream, integrate the video information and production information from the multimodal streams to meet the requirements of the comprehensive dataset. If the missing data ratio exceeds the preset range, supplement it with historical data to determine a complete comprehensive dataset. S7.6. Based on the comprehensive dataset, perform in-depth data fusion processing using cluster analysis methods to group and integrate the video surveillance data and production data, obtaining classified fused data groups and the final multi-source fusion results. S7.7. Based on the final multi-source fusion results, construct a data storage structure to meet the needs of information integration, format the classified fusion data groups, and determine a standardized storage data set.
[0020] S8. Based on the comprehensive data set and to address the adaptability of diverse protocol environments, an adaptive protocol adjustment module is used to dynamically detect the protocol requirements of the target system during data transmission. If the target system protocol does not match the current protocol, the module switches to the corresponding protocol format and obtains the adapted transmission data to ensure that the data can be correctly parsed in the diverse protocol environment.
[0021] S8.1. Build a protocol detection tool to monitor the protocol environment during data transmission in real time, obtain protocol requirement information from the target system, and obtain a preliminary set of protocol features. S8.2. Based on the preliminary set of protocol features, a comparison is performed using a pre-defined protocol matching rule library. If the comparison results indicate a mismatch between the current protocol and the target system's protocol requirements, an adjustment module is triggered to identify the specific protocol fields that do not match. S8.3. For mismatched protocol fields, obtain the pre-established protocol format conversion template, dynamically adjust the data transmission format, and obtain the adapted transmission data content; S8.4. Detect the target system's environmental parsing capabilities using the adapted transmitted data content. If any exceptions occur during the parsing process, record the exception fields and obtain the corresponding parsing error log. S8.5. Based on the parsed error log, use a support vector machine algorithm to classify abnormal fields, determine whether the abnormal fields are omissions during the protocol format conversion process, and determine the classified abnormality category. S8.6. Based on the classified anomaly categories, retrieve the pre-defined protocol correction rule library, adjust the relevant fields in the transmitted data, and generate the final corrected data to ensure that the data is correctly parsed in a diverse protocol environment. Through the final correction data, the protocol matching rule library and protocol format conversion template are dynamically updated, the updated rules and template data are obtained, and it is determined whether the subsequent data transmission can directly adapt to the target system requirements.
[0022] S9. A data integrity check mechanism is used to verify the adapted transmission data. After the data reaches the target system, a checksum comparison is performed. If the checksums do not match, a data retransmission process is triggered to obtain the complete and correct data content. Finally, the transmitted data is judged to meet the preset integrity standards.
[0023] S9.1. Generate a standardized initial data stream by adapting the transmitted data and determine whether it conforms to the predefined format requirements. If the initial data stream meets the format requirements, it will be transmitted to the target system, and the corresponding check code will be obtained when the data arrives to determine whether the check code is consistent with the preset value; If the checksum comparison result is inconsistent, the retransmission process is triggered to re-acquire the transmission data from the source system to obtain the updated data stream; S9.2. Repeat the integrity check step based on the updated data stream to obtain a new checksum and determine whether it matches the preset value. If the new checksum still does not match, the retransmission process continues in a loop until correct data is obtained and it is determined that it meets the integrity standard; S9.3. Utilize a pre-established verification model and the correct data after cyclic retransmission to verify the integrity of the final transmitted data, obtaining a transmission result that meets the preset standards. S9.4. Record the log information for each checksum and retransmission of the final transmission result. Analyze the log data to identify potential anomalies in the transmission process.
[0024] For the adapted data stream, the CRC cyclic redundancy check algorithm is used to generate the check code. The formula is: ; is a binary sequence of raw data (such as an encrypted JSON data packet). is the check bit length, Generate the polynomial "100000111" for CRC-8.
[0025] If the checksums at the receiving end are inconsistent (e.g. ), triggering the retransmission process. For example, the original data After shifting left 8 bits, it is 1010000000, which is the same as The check code is obtained after XOR. If the check fails after transmission, the data is resent until the check passes.
[0026] The present invention can achieve the following beneficial effects: 1. Utilize protocol conversion middleware and device mapping mechanism to convert proprietary protocols of devices from different manufacturers into a unified format, establish a correspondence between devices and data, and achieve standardized collection and integration of heterogeneous device data, breaking down data silos and providing a foundation for unified management and analysis of sewage treatment plant production data.
[0027] 2. The use of incremental data adaptation modules effectively solves the compatibility issues between old systems and newly collected data, enabling new data to be recognized and stored by old systems, facilitating system upgrades and renovations at sewage treatment plants, protecting the original system investment, and reducing upgrade costs.
[0028] 3. The data normalization algorithm is used to unify the format of data fields and fill in missing values, thereby improving data quality. The symmetric encryption algorithm and integrity verification mechanism are used to ensure the security and integrity of data during transmission, prevent data leakage, tampering and loss, and provide a guarantee for the reliable application of sewage treatment plant production data.
[0029] 4. With the help of the network status monitoring module and the adaptive protocol adjustment module, the transmission strategy can be dynamically adjusted according to network conditions, ensuring continuous data transmission in an unstable network environment, and automatically adapting to diverse protocol environments to ensure correct parsing of data between different systems, thereby improving the reliability and adaptability of the system.
[0030] 5. A multi-source data integration framework is used to integrate video surveillance data with production data to obtain a comprehensive data set, providing richer and more comprehensive data support for the intelligent management and decision-making of the sewage treatment plant, which helps to improve the operational efficiency and management level of the sewage treatment plant. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 Flowchart of the present invention; Figure 2 This is the data flow diagram of the present invention; Figure 3 This is a trend diagram of the time consumption for protocol parsing in S1 of the present invention; Figure 4 The network bandwidth and transmission rate dynamic adjustment curve diagram of S6 of the present invention; Figure 5 This is a trend diagram of the data verification success rate in S9 of the present invention. DETAILED DESCRIPTION
[0032] The preferred solution is Figures 1 to 5 As shown in the figure, a unified method based on the data acquisition interface of different SCADA production monitoring systems of sewage treatment plants is presented. The specific steps are as follows: S1. Utilize pre-established protocol conversion middleware to parse and standardize the communication protocols of devices from different manufacturers. Addressing protocol inconsistencies, we convert multiple proprietary protocols into a unified communication format, obtain standardized data streams, and ultimately create a unified data set for subsequent processing.
[0033] S1.1. Using a pre-built middleware framework, the communication protocols of devices from different vendors are parsed and processed to obtain the initial protocol data stream. If the parsed protocol data streams have format differences, pre-set protocol conversion rules are used to convert the proprietary protocols into a unified format, generating standardized communication data. Based on the standardized communication data, data stream label normalization operations are performed to generate a unified communication data set for subsequent analysis. For example, when implementing protocol conversion middleware, we first build a protocol parsing engine to deeply analyze the communication protocols of devices from different manufacturers. For example, suppose a device from manufacturer A uses a proprietary Modbus-based protocol, with packets containing 16-bit addresses and 32-bit data values. Meanwhile, a device from manufacturer B uses a custom binary protocol, with packets beginning with an 8-bit identifier followed by 24 bits of data. We then design a parsing module that, using a byte stream analysis algorithm, disassembles the data packets from manufacturer A bit by bit, extracting address values such as 0x0010 and data values such as 0x00001234. We then perform bitwise operations on the data packets from manufacturer B, separating identifiers such as 0x0A from data values such as 0x00ABCD. The parsing time is kept within 10 milliseconds, ensuring real-time performance. Next, the parsed heterogeneous data is converted into a unified format, for example, using a JSON structure with standard fields defined such as {"device_id": "A001", "value": 4660.5, "timestamp": 1698201234567}. A mapping table maps the address value of vendor A to the device ID. The data value is normalized and scaled by a factor of 1.2, resulting in a calculated value of 4660.5. The data value of vendor B is scaled by a factor of 0.8 to ensure data consistency and keep the conversion error within 0.01%. A data stream normalization engine is then built, utilizing a queue mechanism to process the converted data stream. Assuming 1,000 data items are processed per second, a FIFO algorithm is used for sorting, resulting in a packet loss rate of less than 0.1%. Timestamp verification is used to remove duplicate data and ensure data stream integrity.
[0034] In this step, the middleware parses the device protocol (such as Modbus and OPC UA) to obtain the initial data stream. If there are format differences, the proprietary protocol is converted into a unified JSON format using preset rules. The key features are extracted simultaneously using the support vector machine (SVM) algorithm. The formula is: ; is the data stream feature vector (such as protocol type, data frequency), is the feature weight vector (obtained through training of historical protocol data), is the classification threshold. This formula is used to determine whether the data flow belongs to the target protocol type (output +1 / -1).
[0035] S1.2. If redundant or abnormal fields exist within the unified communications data set, clean it using a pre-defined filtering mechanism to determine a cleaned data set. A support vector machine algorithm is then used to extract features from the cleaned data set to obtain key communication feature values. Based on these extracted key communication feature values, a mapping relationship is constructed between data flow labels and subsequent analysis to determine the suitability of the data set. In this step, after filtering out redundant fields, the SVM feature extraction algorithm is applied to the cleaned data to construct a mapping relationship between data stream features and device types. For example, if the feature vector , it can be determined by formula calculation that it belongs to the PLC device (output + 1), thereby generating a standardized communication data set.
[0036] S1.3. If the suitability of the data set meets the preset threshold, a final standardized data stream is generated and the output processing is completed in a unified format.
[0037] Standardized data sets are stored in distributed databases, such as Redis, with a 24-hour data expiration period. Hash indexes are used to optimize query speed, reducing average query times to less than 5 milliseconds. Data is also compressed, achieving a 30% compression rate, to support subsequent business analysis. For example, energy consumption statistics can be analyzed from 1 million data points to determine a peak energy consumption of 5,000 kWh, aiding decision optimization. These steps are seamlessly integrated through automated scripts and algorithms, forming a complete technical chain from analysis to storage, ensuring efficient and accurate data processing.
[0038] S2. Based on the unified data set, a device mapping mechanism is used to associate the data streams of heterogeneous devices with preset device identifiers. To address the challenges of heterogeneous device integration, a correspondence between devices and data is established, the data source and format of each device are determined, and a device-identified data mapping table is generated.
[0039] S2.1. Build a unified data processing framework to capture data streams from heterogeneous devices and determine the initial data set. S2.2. Based on the preset device identifiers, the initial data set is classified using a mapping mechanism to generate device-associated data groups. If a data stream within the data group does not match the device identifier, the device category is determined by comparing the data source and format, and a preliminary mapping relationship is generated. S2.3. Based on the preliminary mapping relationships, obtain characteristic information about the data flow and determine whether there are any integration challenges. If the characteristic information does not meet the preset threshold, adjust the mapping mechanism to obtain an optimized mapping relationship. Through the optimized correspondence, an identification table between devices and data is constructed to obtain the clear data source and data format of each device, and determine whether the device association is completed; S2.4. Based on the contents of the identification table, use a support vector machine algorithm to verify the data stream, determine the accuracy of the data mapping, and generate the final device identification data mapping table; S2.5. Obtain the complete data integration results for heterogeneous devices through the final device identification data mapping table, determine whether it meets business requirements, and complete the comprehensive association of data and devices.
[0040] For example, consider three devices: Device A is a temperature sensor, Device B is a humidity sensor, and Device C is a pressure sensor. Each device generates a data stream containing a temperature value (e.g., 25.5°C), a humidity value (e.g., 60.2%), and a pressure value (e.g., 1013.2 Pa) every minute. The system assigns a unique identifier to each device, such as A001, B002, or C003, and binds the data stream to the identifier using a mapping algorithm. This algorithm uses a hash function to calculate the correspondence between the device's hardware address and the identifier, ensuring a one-to-one correspondence between the data stream and the device. Next, to establish a correspondence between the device and the data, the system automatically analyzes the data source and format of each device. For example, by parsing the data packet header from Device A, it can identify the data format as a floating-point number, in degrees Celsius, with a sampling frequency of 1 / minute. This information is then recorded in the database, forming a structured information table. After determining the data source and format for each device, the system generates a device-identified data mapping table. For example, the table records the temperature data for device A001 as a floating-point number, ranging from -40.0 to 85.0 degrees Celsius. If the temperature falls outside this range, the anomaly detection algorithm is triggered, which calculates the deviation and logs it. For example, a temperature of 25.5 degrees Celsius is considered normal, while a temperature of -41.0 degrees Celsius is marked as abnormal with a deviation of 1.0 degrees Celsius. Through this process, the system automatically completes the entire process from data collection to mapping, ensuring data traceability. It also integrates anomaly detection mechanisms to enhance data reliability, forming a complete technical closed loop.
[0041] S3. Using the device-identified data mapping table, we implement a data normalization algorithm to unify the format of data fields from different sources and fill in missing values to obtain a standardized data structure and a standardized data set suitable for system analysis.
[0042] S3.1. Obtain an initial data set from multiple data sources using device identification and mapping tables. Perform a preliminary classification of data fields from different sources to obtain a classified field set. Based on this classified field set, perform a consistency check on the data fields using the collection standard. If the check reveals inconsistent field formats, apply a normalization algorithm to unify the formats and determine a consistent field data set. S3.2. Detect missing values in a consistent field data set. If missing values are detected, fill them using a pre-defined interpolation method to obtain a complete set of data fields. Based on this complete set of data fields, construct a standardized data structure and apply pre-defined mapping rules to associate the data fields with device identifiers, resulting in a structured, standardized data set. In the specific implementation process of data collection standards, the first step is to uniformly identify and manage device data from different sources through a device identification data mapping table. For example, suppose there is device data from two different suppliers. Supplier A's device ID format is "A123-456", while Supplier B's format is "B_789_012". Using the mapping table, both are converted to the standard format "DEV_XXXXXX". For example, "A123-456" is mapped to "DEV_000001", and "B_789_012" is mapped to "DEV_000002". This is then stored in the database for subsequent access. Next, to implement data collection standards, we used a data normalization algorithm to unify the format of data fields from different sources. For example, we unified the time field "2023 / 01 / 01 08:00:00" from Supplier A and "2023-01-01T08:00:00" from Supplier B into "2023-01-01 08:00:00." Regular expression matching and replacement algorithms were used to convert the format and ensure the consistency of the time fields. Furthermore, to fill missing values, if the temperature field for a device is missing, a time series-based linear interpolation algorithm was used. By combining the data points from the two hours before and after (e.g., the temperature was 25.5 degrees Celsius in the previous hour and 26.5 degrees Celsius in the next hour), the missing value was calculated to be approximately 26.0 degrees Celsius. After filling, a complete dataset was formed.
[0043] S3.3. Use a structured standard dataset to integrate data based on source differences. Standardize data from different sources using a unified coding method to determine the integrated standard dataset. For the classified field set, the collection standard is used to check the format consistency. If missing values are detected (such as flow meter data gaps), they are filled by linear interpolation. The formula is: ; The time to be supplemented The flow value, and is the known flow value 10 seconds before and after; For example, a flow meter When data is missing, it is known The flow rate is , At that time , substituting into the formula we get: ; S3.4. Perform data quality verification on the integrated canonical dataset. If any abnormal data is found during the verification, filter it using a preset threshold to obtain a final canonical dataset suitable for system analysis. Generate corresponding metadata descriptions for the final canonical dataset, recording the formatting of data fields and the gap filling process, to produce a complete dataset with descriptive information.
[0044] Finally, the standardized data structure is stored in a unified JSON format, for example, {"device_id":"DEV_000001", "timestamp": "2023-01-01 08:00:00", "temperature": 26.0}, to facilitate system access and analysis. The standardized dataset is analyzed using a data quality verification algorithm to calculate field completeness (for example, 98.5% for the temperature field) and the proportion of outliers (for example, 0.2% of records have temperatures outside the acceptable range [0,50]). An analysis report is generated to confirm that the dataset meets the system's analysis requirements. This process forms a complete logical chain from identification to analysis, ensuring data standardization throughout the entire process, from acquisition to application.
[0045] S4. Based on the standard data set, to address compatibility issues with legacy system upgrades, an incremental data adaptation module is used to adapt the newly collected data to the legacy system's interface, obtain a data stream compatible with the legacy system, and ensure that the adapted data can be recognized and stored by the legacy system.
[0046] S4.1. Analyze compatibility issues between legacy systems and newly collected data, build a data adaptation framework, obtain a preliminary interface matching solution, and ensure that the adaptation framework can handle data differences. S4.2. Based on the output of the adaptation framework, perform incremental adaptation on the newly collected data in layers to generate initial compatible data. Determine whether the initial compatible data meets the interface matching requirements. S4.3. If the initial compatible data does not meet the interface matching requirements, adjust the data fields using the preset rule base to obtain the adjusted compatible data and confirm that the adjusted data passes the interface verification. S4.4. For the adjusted compatible data, use data flow mapping technology to construct a data transmission path, obtain the stability index of the transmission path, and confirm that the transmission path can support the smooth transmission of data flow; S4.5. Apply the support vector machine algorithm to optimize the data stream based on the transmission path stability indicator, generate an optimized data stream, and determine whether the optimized data stream meets the recognition capability requirements. S4.6. If the optimized data stream meets the recognition capability requirements, transfer it to the legacy system's storage mechanism to obtain storage status feedback and determine whether the storage process is complete. S4.7. Generate data storage logs based on storage status feedback. Use log analysis tools to detect anomalies in the storage process. Obtain anomaly detection results to determine whether the legacy system can continue to operate stably.
[0047] When dealing with compatibility issues in upgrading old systems, data adaptation and compatibility can be achieved through a series of technical means. First, for the construction of incremental data adaptation modules, a data conversion algorithm can be used to convert the newly collected data format into a format supported by the old system. For example, assuming that the data collected by the new system is in JSON format, each record contains fields such as "timestamp" and "value", while the old system only supports CSV format and the field order is "date" and "value", we can use scripts to automatically parse JSON data, extract key fields, and reorganize them in CSV format. The error of each line of data in the converted file is set to be within 0.01, and the amount of data before and after the conversion (such as 1,000 records) is compared to ensure that there is no loss. Secondly, when adapting data to legacy system interfaces, a mid-layer adapter can be designed. Using API mapping technology, this technology maps the new data call method to the legacy system's interface specifications. For example, if the new system API returns data once per second, while the legacy system interface requires once every five seconds, the adapter can use a cache mechanism to aggregate the data within five seconds, taking the average value (e.g., five values of 1.2, 1.3, 1.4, 1.1, and 1.5, with an average of 1.3) before pushing it to ensure smooth data flow. Next, when acquiring data streams compatible with the legacy system, data integrity can be verified using a checksum algorithm, such as CRC32. A checksum value is calculated for each batch of data (e.g., 100KB) and compared with the expected value, with an error rate below 0.001%, ensuring the data stream is recognizable. Finally, to ensure that the adapted data can be stored in the legacy system, a simulation test is conducted, inputting 1,000 pieces of adapted data and monitoring the storage success rate, aiming for a 99.9% or higher rate. Storage logs are also analyzed. If anomalies are detected (such as field length exceeding the limit), adaptation rules are automatically adjusted, such as truncating overlong fields to the 50-character limit supported by the legacy system. These technical measures form a complete chain from data conversion to storage verification, ensuring that compatibility issues are resolved.
[0048] S5. To ensure data transmission security through the adapted data stream, a data encryption protection mechanism is implemented. A symmetric encryption algorithm is used to encrypt the data. Prior to transmission, the data is encrypted with a key, and the encrypted data packet is obtained to securely transmit the data content.
[0049] S5.1. Based on the adjusted data stream, encrypt the original data using a symmetric encryption algorithm to generate a preliminary encrypted data unit, in accordance with transmission security and safety requirements. S5.2. Based on the initially encrypted data unit, perform a secondary encryption process on the data unit using a preset key encryption mechanism before transmission to obtain a data block with a higher encryption strength. If an anomaly is detected in the data block during the encryption process, the data block integrity is verified through the preset error checking mechanism to determine whether there is any data corruption or tampering; If the verification result shows that the data block integrity is normal, a data packaging tool is used to convert the data block into a data packet in a standard format and determine the transmission readiness status of the data packet; S5.3. By adapting the data packet to the transmission channel, obtaining the transmission protocol that matches the target channel and obtaining the adapted data packet structure; According to the adapted data packet structure, the preset transmission security policy is used to perform a final security check on the data packet to determine whether it meets the requirements of secure transmission; S5.4. If the test result meets the preset security threshold, the data packet is transmitted through the target channel, and the final secure data content is obtained. To ensure data transmission security, the data stream is first adapted to a standard format. For example, input text data is processed in 128-byte blocks to ensure the efficient operation of the subsequent encryption algorithm. Next, the symmetric encryption algorithm AES-256 is selected for encryption to generate a 256-bit key for data protection.
[0050] S6. Based on the encrypted data packets and to ensure stable data transmission, a network status monitoring module is used to monitor the impact of network conditions in real time during transmission. If the network bandwidth falls below a preset threshold, the data packet transmission rate is reduced and an adjusted transmission strategy is obtained to ensure that the data packet can be continuously transmitted in an unstable network.
[0051] S6.1. The network status monitoring module continuously monitors network conditions, obtains current bandwidth and latency information, and determines whether the network is stable. If the collected bandwidth data is lower than the preset threshold, the rate adjustment mechanism is triggered to reduce the transmission rate of the data packet and obtain the adjusted transmission parameters; S6.2. Reconfigure the packet distribution logic based on the adjusted transmission parameters, obtain the updated transmission strategy, and determine how to group the packets. Using a dynamic rate adjustment algorithm, such as one based on a simplified model of TCP congestion control, the transmission rate is reduced from an initial 2.0 Mbps to 1.5 Mbps. The calculation is: New Rate = Current Rate × (1 - Bandwidth Shortage Ratio), where Bandwidth Shortage Ratio = (3.0 - 2.5) / 3.0 = 0.1667. Adjusted Rate = 2.0 × (1 - 0.1667) = 1.67 Mbps, rounded to 1.5 Mbps for conservatism. Analysis of the transmission performance after the adjustment predicts that at a 1.5 Mbps rate, packet transmission time will be extended from 10 seconds to 13.3 seconds, while the packet loss rate is expected to drop from 1.2% to 0.8%, improving transmission stability.
[0052] S6.3. Based on the updated transmission policy, use the packet prioritization method to determine the transmission order of high-priority packets and determine whether the transmission queue arrangement is reasonable. If the transmission queue arrangement complies with the preset rules, the data packet transmission is executed according to the adjusted strategy, and real-time feedback data is obtained during the transmission process; S6.4 Analyze real-time feedback data and use the support vector machine algorithm to predict network status trends, obtain prediction results, and determine whether to adjust the transmission strategy; S6.5. If the network status is likely to deteriorate further based on the prediction results, perform a secondary optimization of the transmission strategy to determine the final transmission plan.
[0053] If network conditions continue to deteriorate, for example, if the bandwidth further drops to 1.0 Mbps, the algorithm repeats, calculating a new rate of 1.5 × (1-(3.0 - 1.0) / 3.0) = 0.75 Mbps to ensure uninterrupted transmission. Furthermore, to enhance system robustness, a backup transmission channel selection mechanism is introduced. If the primary channel bandwidth falls below 1.0 Mbps for five consecutive minutes, the system automatically switches to the backup channel. Assuming the backup channel bandwidth is 1.8 Mbps, the transmission rate is recalculated and the policy is updated to ensure continuous data transmission. This method enables the system to dynamically adjust its policy in unstable networks, ensuring stable packet transmission.
[0054] S7. To achieve video surveillance fusion through the adjusted transmission strategy, a multi-source data integration framework is employed to align the timestamps and associate the content of video surveillance data with production data. This generates a fused multimodal data stream and a comprehensive dataset containing both video and production information.
[0055] S7.1. Complete preliminary data collection by building a multi-source data integration framework to obtain raw streams from video surveillance and production data, generating an unprocessed mixed data set. S7.2. Based on the mixed data set, synchronize the time information of the video surveillance data and production data using a timestamp alignment method. If the timestamp deviation exceeds a preset threshold, interpolate the data to determine the aligned synchronized data set. A multi-source data integration framework was constructed, and a timestamp alignment algorithm was used to adjust the timestamp accuracy of video surveillance data to the millisecond level to match it with the timestamp of production data. For example, the video frame timestamp is 2023-10-01 14:30:00.123, and the production data timestamp is 2023-10-01 14:30:00.125. The time difference calculation (difference 0.002 seconds) determines that the data is synchronized. The error threshold is set to 0.01 seconds. Data exceeding the threshold will be marked as asynchronous and enter the cache queue for processing. Analysis shows that the synchronization rate can reach 98.7%.
[0056] S7.3. Based on the synchronized data sets and content association requirements, extract keyframe features from the video surveillance data and event records from the production data. Use a support vector machine algorithm to perform feature matching, identify highly correlated data pairs, and form a correlated data set. Content association processing is performed using a feature matching-based algorithm to compare the device operating status detected in the video (for example, determining that the device speed is 1200 rpm through an image recognition algorithm) with the device parameters in the production data (for example, the speed is recorded as 1198 rpm). The error rate is controlled within 1%. Successfully associated data pairs are marked as valid data. Analysis shows that the association accuracy rate reaches 96.3%.
[0057] S7.4. Generate a multimodal data stream based on the associated data set, fuse the feature information of the video surveillance data and the production data, obtain structured multimodal stream data, and obtain a preliminary fused data stream. S7.5. Based on the initial fused data stream, integrate the video information and production information from the multimodal streams to meet the requirements of the comprehensive dataset. If the missing data ratio exceeds the preset range, supplement it with historical data to determine a complete comprehensive dataset. S7.6. Based on the comprehensive dataset, perform in-depth data fusion processing using cluster analysis methods to group and integrate the video surveillance data and production data, obtaining classified fused data groups and the final multi-source fusion results. S7.7. Based on the final multi-source fusion results, construct a data storage structure to meet the needs of information integration, format the classified fusion data groups, and determine a standardized storage data set.
[0058] In the process of achieving video surveillance fusion, data stream transmission efficiency is first optimized through an adjusted transmission strategy. For example, a priority-based scheduling algorithm is used to allocate transmission bandwidth between video surveillance data and production data in a ratio of 7:3. This ensures high real-time performance for video data while also ensuring stable transmission of production data. Assuming a video data rate of 10 Mbps and a production data rate of 4.3 Mbps, calculations show that video data transmission latency is controlled within 50 ms and production data latency is controlled within 100 ms. Analysis results show that this strategy can reduce the packet loss rate to below 0.5%. Finally, a fused multimodal data stream is generated. Data encapsulation technology is used to package video frames and production data into a unified format. For example, each video frame is bound to production data from a corresponding time period (e.g., temperature 25.5 degrees Celsius, humidity 60%) to form a comprehensive dataset. Data storage is stored in a distributed database, processing approximately 2.5 GB of data per hour. Analysis shows that data query efficiency has increased by approximately 30%, supporting subsequent business analysis such as equipment failure prediction. This logically forms a complete chain from data acquisition to fusion to application.
[0059] S8. Based on the comprehensive data set and to address the adaptability of diverse protocol environments, an adaptive protocol adjustment module is used to dynamically detect the protocol requirements of the target system during data transmission. If the target system protocol does not match the current protocol, the module switches to the corresponding protocol format and obtains the adapted transmission data to ensure that the data can be correctly parsed in the diverse protocol environment.
[0060] S8.1. Build a protocol detection tool to monitor the protocol environment during data transmission in real time, obtain protocol requirement information from the target system, and obtain a preliminary set of protocol features. S8.2. Based on the preliminary set of protocol features, a comparison is performed using a pre-defined protocol matching rule library. If the comparison results indicate a mismatch between the current protocol and the target system's protocol requirements, an adjustment module is triggered to identify the specific protocol fields that do not match. S8.3. For mismatched protocol fields, obtain the pre-established protocol format conversion template, dynamically adjust the data transmission format, and obtain the adapted transmission data content; In adaptive processing for diverse protocol environments, dynamic adjustments to data transmission are first implemented by the adaptive protocol adjustment module, which performs real-time detection of the target system's protocol requirements. Assuming the current network environment has three protocol formats: Protocols A, B, and C, the detection module analyzes the protocol identification field returned by the target system (for example, a field value of 1 indicates Protocol A, 2 indicates Protocol B, and 3 indicates Protocol C) to determine the target system's current protocol type. This detection frequency is set to once per second to ensure real-time performance. Next, if the target system detects that it uses Protocol B while the current transmission protocol is Protocol A, the module triggers the protocol switching mechanism and invokes a pre-defined protocol conversion algorithm. For example, the module adjusts the packet header length of Protocol A from a fixed 32 bytes to the dynamic length of Protocol B (calculated based on the data content; for example, if the data content is 100 bytes, the Protocol B header length is 100*0.2 = 20 bytes). The module then repackages the packet to ensure format matching.
[0061] S8.4. Detect the target system's environmental parsing capabilities using the adapted transmitted data content. If any exceptions occur during the parsing process, record the exception fields and obtain the corresponding parsing error log. S8.5. Based on the parsed error log, use a support vector machine algorithm to classify abnormal fields, determine whether the abnormal fields are omissions during the protocol format conversion process, and determine the classified abnormality category. S8.6. Based on the classified anomaly categories, retrieve the pre-defined protocol correction rule library, adjust the relevant fields in the transmitted data, and generate the final corrected data to ensure that the data is correctly parsed in a diverse protocol environment. Through the final correction data, the protocol matching rule library and protocol format conversion template are dynamically updated, the updated rules and template data are obtained, and it is determined whether the subsequent data transmission can directly adapt to the target system requirements.
[0062] The adapted transmission data is obtained and the data integrity is calculated using a checksum algorithm (such as CRC32). Assuming the original data checksum is 0x12345678, the adapted checksum remains 0x12345678, confirming that the data has not been lost. Finally, to ensure that the data can be correctly parsed in a diverse protocol environment, the system simulates the target environment for a parsing test. Assuming a 99.8% success rate for 1,000 data packets, if it falls below the 99.5% threshold, secondary optimization adjustments are triggered, such as adding redundant fields (adding a 2-byte checksum to each data packet) to improve the parsing success rate. The above process forms a complete logical chain from detection to adaptation to verification, ensuring the compatibility and reliability of data transmission in different protocol environments.
[0063] S9. A data integrity check mechanism is used to verify the adapted transmission data. After the data reaches the target system, a checksum comparison is performed. If the checksums do not match, a data retransmission process is triggered to obtain the complete and correct data content. Finally, the transmitted data is judged to meet the preset integrity standards.
[0064] S9.1. Generate a standardized initial data stream by adapting the transmitted data and determine whether it conforms to the predefined format requirements. If the initial data stream meets the format requirements, it will be transmitted to the target system, and the corresponding check code will be obtained when the data arrives to determine whether the check code is consistent with the preset value; If the checksum comparison result is inconsistent, the retransmission process is triggered to re-acquire the transmission data from the source system to obtain the updated data stream; S9.2. Repeat the integrity check step based on the updated data stream to obtain a new checksum and determine whether it matches the preset value. If the new checksum still does not match, the retransmission process continues in a loop until correct data is obtained and it is determined that it meets the integrity standard; S9.3. Utilize a pre-established verification model and the correct data after cyclic retransmission to verify the integrity of the final transmitted data, obtaining a transmission result that meets the preset standards. S9.4. Record the log information for each checksum and retransmission of the final transmission result. Analyze the log data to identify potential anomalies in the transmission process.
[0065] For the adapted data stream, the CRC cyclic redundancy check algorithm is used to generate the check code. The formula is: ; is a binary sequence of raw data (such as an encrypted JSON data packet). is the check bit length, Generate the polynomial "100000111" for CRC-8.
[0066] If the checksums at the receiving end are inconsistent (e.g. ), triggering the retransmission process. For example, the original data After shifting left 8 bits, it is 1010000000, which is the same as The check code is obtained after XOR. If the check fails after transmission, the data is resent until the check passes.
[0067] After retransmission, verification is performed again until the checksum matches. For example, after retransmission, the checksum matches the data, confirming data integrity. Finally, the system makes a judgment based on the preset integrity criteria: the checksum matches and the packet size is equal to 1024 bytes. If these criteria are met, a successful transmission log is recorded, including the packet ID and transmission time. If these criteria are not met, retransmission continues and the number of failures is recorded. After more than three failures, the relevant business system is automatically notified for exception handling, forming a closed-loop logic to ensure data transmission reliability. Through this process, from data encapsulation to verification, retransmission, and final judgment, a complete technical chain is formed to ensure that data transmission meets the expected standards.
[0068] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The scope of protection of the present invention shall be the technical solutions set forth in the claims, including equivalent alternatives to the technical features of the technical solutions set forth in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.
Claims
1. A unified method based on the data acquisition interface of different SCADA production monitoring systems of sewage treatment plants, characterized by: S1. Utilize protocol conversion middleware to parse communication protocols of devices from different vendors, convert proprietary protocols into a unified format, and output standardized data streams and unified data sets after processing. S2. Based on a unified data set, the device mapping mechanism associates heterogeneous device data streams with preset identifiers to establish a corresponding relationship and form a device-identified data mapping table; S3. Based on the data mapping table, apply data normalization algorithms to unify data field formats, fill missing values, build a standardized data structure, and generate a standardized data set. S4. To ensure compatibility with legacy systems, an incremental data adaptation module is used to adapt newly collected data to legacy system interfaces, enabling data identification and storage. S5. To ensure data transmission security, the adapted data stream is encrypted using a symmetric encryption algorithm, generating an encrypted data packet to ensure secure transmission. S6. Based on encrypted data packets, the network status monitoring module dynamically adjusts the transmission rate and formulates transmission strategies to ensure continuous data transmission even under unstable networks. S7. Based on the adjusted transmission strategy, use a multi-source data integration framework to fuse video surveillance and production data to obtain a comprehensive dataset. S8. Based on the comprehensive dataset, use the adaptive protocol adjustment module to dynamically adapt the target system protocol to ensure correct data interpretation in diverse environments; S9. For the adapted transmission data, the integrity check mechanism compares the checksum and triggers retransmission to ensure data integrity and compliance with preset standards.
2. A unified method based on data acquisition interfaces of different SCADA production monitoring systems of sewage treatment plants according to claim 1, characterized in that: The sub-steps of step S1 are: S1.
1. Using a pre-built middleware framework, the communication protocols of devices from different vendors are parsed and processed to obtain the initial protocol data stream. If the parsed protocol data streams have format differences, pre-set protocol conversion rules are used to convert the proprietary protocols into a unified format, generating standardized communication data. Based on the standardized communication data, data stream label normalization operations are performed to generate a unified communication data set for subsequent analysis. S1.
2. If redundant or abnormal fields exist within the unified communications data set, clean it using a pre-defined filtering mechanism to determine a cleaned data set. A support vector machine algorithm is then used to extract features from the cleaned data set to obtain key communication feature values. Based on these extracted key communication feature values, a mapping relationship is constructed between data flow labels and subsequent analysis to determine the suitability of the data set. S1.
3. If the suitability of the data set meets the preset threshold, a final standardized data stream is generated and the output processing is completed in a unified format.
3. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S2 are: S2.
1. Build a unified data processing framework to capture data streams from heterogeneous devices and determine the initial data set. S2.
2. Based on the preset device identifiers, the initial data set is classified using a mapping mechanism to generate device-associated data groups. If a data stream within the data group does not match the device identifier, the device category is determined by comparing the data source and format, and a preliminary mapping relationship is generated. S2.
3. Based on the preliminary mapping relationships, obtain characteristic information about the data flow and determine whether there are any integration challenges. If the characteristic information does not meet the preset threshold, adjust the mapping mechanism to obtain an optimized mapping relationship. Through the optimized correspondence, an identification table between devices and data is constructed to obtain the clear data source and data format of each device, and determine whether the device association is completed; S2.
4. Based on the contents of the identification table, use a support vector machine algorithm to verify the data stream, determine the accuracy of the data mapping, and generate the final device identification data mapping table. S2.
5. Obtain the complete data integration results for heterogeneous devices through the final device identification data mapping table, determine whether it meets business requirements, and complete the comprehensive association of data and devices.
4. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S3 are: S3.
1. Obtain an initial data set from multiple data sources using device identification and mapping tables. Perform a preliminary classification of data fields from different sources to obtain a classified field set. Based on this classified field set, perform a consistency check on the data fields using the collection standard. If the check reveals inconsistent field formats, apply a normalization algorithm to unify the formats and determine a consistent field data set. S3.
2. Detect missing values in a consistent field dataset. If missing values are detected, fill them using a pre-defined interpolation method to obtain a complete set of data fields. Based on the complete set of data fields, a standardized data structure is constructed, and the data fields are associated with device identifiers using preset mapping rules to obtain a structured standard data set; S3.
3. Use a structured standard dataset to integrate data based on source differences. Standardize data from different sources using a unified coding method to determine the integrated standard dataset. S3.
4. Perform data quality verification on the integrated canonical dataset. If any abnormal data is found during the verification, filter it using a preset threshold to obtain a final canonical dataset suitable for system analysis. Generate corresponding metadata descriptions for the final canonical dataset, recording the formatting of data fields and the gap filling process, to produce a complete dataset with descriptive information.
5. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1 is characterized in that: The sub-steps of step S4 are: S4.
1. Analyze compatibility issues between legacy systems and newly collected data, build a data adaptation framework, obtain a preliminary interface matching solution, and ensure that the adaptation framework can handle data differences. S4.
2. Based on the output of the adaptation framework, perform incremental adaptation on the newly collected data in layers to generate initial compatible data. Determine whether the initial compatible data meets the interface matching requirements. S4.
3. If the initial compatible data does not meet the interface matching requirements, adjust the data fields using the preset rule base to obtain the adjusted compatible data and confirm that the adjusted data passes the interface verification. S4.
4. For the adjusted compatible data, use data flow mapping technology to construct a data transmission path, obtain the stability index of the transmission path, and confirm that the transmission path can support the smooth transmission of data flow; S4.
5. Apply the support vector machine algorithm to optimize the data stream based on the transmission path stability indicator, generate an optimized data stream, and determine whether the optimized data stream meets the recognition capability requirements. S4.
6. If the optimized data stream meets the recognition capability requirements, transfer it to the legacy system's storage mechanism to obtain storage status feedback and determine whether the storage process is complete. S4.
7. Generate data storage logs based on storage status feedback. Use log analysis tools to detect anomalies in the storage process. Obtain anomaly detection results to determine whether the legacy system can continue to operate stably.
6. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S5 are: S5.
1. Based on the adjusted data stream, encrypt the original data using a symmetric encryption algorithm to generate a preliminary encrypted data unit, in accordance with transmission security and safety requirements. S5.
2. Based on the initially encrypted data unit, perform a secondary encryption process on the data unit using a preset key encryption mechanism before transmission to obtain a data block with a higher encryption strength. If an anomaly is detected in the data block during the encryption process, the data block integrity is verified through the preset error checking mechanism to determine whether there is any data corruption or tampering; If the verification result shows that the data block integrity is normal, a data packaging tool is used to convert the data block into a data packet in a standard format and determine the transmission readiness status of the data packet; S5.
3. By adapting the data packet to the transmission channel, obtaining the transmission protocol that matches the target channel and obtaining the adapted data packet structure; According to the adapted data packet structure, the preset transmission security policy is used to perform a final security check on the data packet to determine whether it meets the requirements of secure transmission; S5.
4. If the detection result meets the preset security threshold, the data packet is transmitted through the target channel to obtain the final secure transmission data content.
7. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S6 are: S6.
1. The network status monitoring module continuously monitors network conditions, obtains current bandwidth and latency information, and determines whether the network is stable. If the collected bandwidth data is lower than the preset threshold, the rate adjustment mechanism is triggered to reduce the transmission rate of the data packet and obtain the adjusted transmission parameters; S6.
2. Reconfigure the packet distribution logic based on the adjusted transmission parameters, obtain the updated transmission strategy, and determine how to group the packets. S6.
3. Based on the updated transmission policy, use the packet prioritization method to determine the transmission order of high-priority packets and determine whether the transmission queue arrangement is reasonable. If the transmission queue arrangement complies with the preset rules, the data packet transmission is executed according to the adjusted strategy, and real-time feedback data is obtained during the transmission process; S6.4 Analyze real-time feedback data and use the support vector machine algorithm to predict network status trends, obtain prediction results, and determine whether to adjust the transmission strategy; S6.
5. If the network status is likely to deteriorate further based on the prediction results, perform a secondary optimization of the transmission strategy to determine the final transmission plan.
8. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S7 are: S7.
1. Complete preliminary data collection by building a multi-source data integration framework to obtain raw streams from video surveillance and production data, generating an unprocessed mixed data set. S7.
2. Based on the mixed data set, synchronize the time information of the video surveillance data and production data using a timestamp alignment method. If the timestamp deviation exceeds a preset threshold, interpolate the data to determine the aligned synchronized data set. S7.
3. Based on the synchronized data sets and content association requirements, extract keyframe features from the video surveillance data and event records from the production data. Use a support vector machine algorithm to perform feature matching, identify highly correlated data pairs, and form a correlated data set. S7.
4. Generate a multimodal data stream based on the associated data set, fuse the feature information of the video surveillance data and the production data, obtain structured multimodal stream data, and obtain a preliminary fused data stream. S7.
5. Based on the initial fused data stream, integrate the video information and production information from the multimodal streams to meet the requirements of the comprehensive dataset. If the missing data ratio exceeds the preset range, supplement it with historical data to determine a complete comprehensive dataset. S7.
6. Based on the comprehensive dataset, perform in-depth data fusion processing using cluster analysis methods to group and integrate the video surveillance data and production data, obtaining classified fused data groups and the final multi-source fusion results. S7.
7. Based on the final multi-source fusion results, construct a data storage structure to meet the needs of information integration, format the classified fusion data groups, and determine a standardized storage data set.
9. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S8 are: S8.
1. Build a protocol detection tool to monitor the protocol environment during data transmission in real time, obtain protocol requirement information from the target system, and obtain a preliminary set of protocol features. S8.
2. Based on the preliminary set of protocol features, a comparison is performed using a pre-defined protocol matching rule library. If the comparison results indicate a mismatch between the current protocol and the target system's protocol requirements, an adjustment module is triggered to identify the specific protocol fields that do not match. S8.
3. For mismatched protocol fields, obtain the pre-established protocol format conversion template, dynamically adjust the data transmission format, and obtain the adapted transmission data content; S8.
4. Detect the target system's environmental parsing capabilities using the adapted transmitted data content. If any exceptions occur during the parsing process, record the exception fields and obtain the corresponding parsing error log. S8.
5. Based on the parsed error log, use a support vector machine algorithm to classify abnormal fields, determine whether the abnormal fields are omissions during the protocol format conversion process, and determine the classified abnormality category. S8.
6. Based on the classified anomaly categories, retrieve the pre-defined protocol correction rule library, adjust the relevant fields in the transmitted data, and generate the final corrected data to ensure that the data is correctly parsed in a diverse protocol environment. Through the final correction data, the protocol matching rule library and protocol format conversion template are dynamically updated, the updated rules and template data are obtained, and it is determined whether the subsequent data transmission can directly adapt to the target system requirements.
10. The unified method based on data acquisition interfaces of different SCADA production monitoring systems of a sewage treatment plant according to claim 1, characterized in that: The sub-steps of step S9 are: S9.
1. Generate a standardized initial data stream by adapting the transmitted data and determine whether it conforms to the predefined format requirements. If the initial data stream meets the format requirements, it will be transmitted to the target system, and the corresponding check code will be obtained when the data arrives to determine whether the check code is consistent with the preset value; If the checksum comparison result is inconsistent, the retransmission process is triggered to re-acquire the transmission data from the source system to obtain the updated data stream; S9.
2. Repeat the integrity check step based on the updated data stream to obtain a new checksum and determine whether it matches the preset value. If the new checksum still does not match, the retransmission process continues in a loop until correct data is obtained and it is determined that it meets the integrity standard; S9.
3. Utilize a pre-established verification model and the correct data after cyclic retransmission to verify the integrity of the final transmitted data, obtaining a transmission result that meets the preset standards. S9.
4. Record the log information for each checksum and retransmission of the final transmission result. Analyze the log data to identify potential anomalies in the transmission process.
Citation Information
Patent Citations
Method, system, medium and equipment for assimilating complex multi-source heterogeneous data of nuclear power plant
CN112598797A
Method and device for expanding system interface
CN112948306A
Photovoltaic data acquisition system and acquisition method
CN115776525A
Geographic entity generation method and system based on multi-source heterogeneous data
CN116719898A
Internet data acquisition and instruction control middleware for water affair industry
CN119814838A
Cited By
Steel wire rope detection and real-time transmission method and system based on multi-modal data fusion
CN121389045A
Steel wire rope detection and real-time transmission method and system based on multi-modal data fusion
CN121389045B