Sewage purification data processing method
Through the B+ tree index structure and random insertion strategy, the problems of large amount of data and low retrieval efficiency in wastewater purification data processing are solved, and efficient data storage and query performance are achieved.
Patent Information
- Application Number
- CN202510811751.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In traditional sewage purification data processing methods, the data volume is huge and complex, the data indexing and retrieval efficiency are inefficient, and the lack of intelligent node selection strategies leads to uneven data distribution, storage redundancy and query performance.
Using the B+ tree index structure, by generating random numbers and calculating alternating scheduling index, the optimal leaf node is determined for data insertion, and the data distribution and storage utilization are optimized.
It improves data retrieval efficiency, dynamically adapts to data distribution changes, reduces waste of storage space, and ensures the balance and real-time nature of the data structure.
Smart Images

Figure CN120336407A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for processing sewage purification data. Background Art
[0002] In the process of sewage purification, data collection, processing, and analysis are important foundations for ensuring purification effects, optimizing process parameters, and realizing intelligent management. However, traditional methods for processing sewage purification data often face the following challenges: On the one hand, the data volume is huge and complex. The sewage purification system involves multiple monitoring points (such as inlet, outlet, and reaction tank, etc.), and each monitoring point generates a large amount of data in real time (such as pH value, COD concentration, ammonia nitrogen content, and flow rate, etc.). These data not only have a large scale but also have the characteristics of high dimension, high frequency, and multi-source heterogeneity, bringing huge pressure to data storage, management, and analysis. On the other hand, the efficiency of data indexing and retrieval is low. In traditional data processing methods, sewage purification data is usually stored in a linear or simple tree structure, lacking an efficient indexing mechanism. When querying or updating data for a specific time period or specific monitoring point, the system often needs to traverse a large amount of data, resulting in low retrieval efficiency and unable to meet real-time requirements. In addition, as the sewage purification system operates, new data reports will be continuously generated (such as daily reports, real-time monitoring reports, etc.). When traditional methods insert new data into the existing data structure, they often lack an intelligent node selection strategy, which may lead to uneven data distribution, storage redundancy, or a decline in query performance. Summary of the Invention
[0003] In order to solve the above problems, the present invention proposes a method for processing sewage purification data.
[0004] The technical solution of the present invention is: A method for processing sewage purification data includes the following steps:
[0005] S1. Obtain the existing index of the sewage purification data packet. When generating the latest data report, determine the transmission set between the latest data report and each leaf node of the existing index;
[0006] S2. Generate the best leaf node according to the transmission set between the latest data report and each leaf node of the existing index;
[0007] S3. Insert the latest data report into the best leaf node to complete the update of the sewage purification data packet.
[0008] Further, S1 includes the following sub-steps:
[0009] S11. Obtain the existing index of the sewage purification data packet; the existing index is a B+ tree index;
[0010] S12. When inserting the latest data report into the sewage purification data packet, obtain the primary key ID of the latest data report;
[0011] S13. According to the primary key ID of the latest data report, determine the transfer set between the latest data report and each leaf node of the B+ tree index.
[0012] The beneficial effect of the above further solution is: In the present invention, the B+ tree is a multi-way balanced search tree with efficient retrieval and range query capabilities. In sewage purification data processing, the B+ tree index can quickly locate data with a specific primary key ID. By uniquely identifying each data report through the primary key ID (such as timestamp, sensor ID, etc.), the data to be processed can be quickly located, improving the efficiency of data processing. By analyzing the transfer set between the latest data report and the B+ tree index nodes, it is possible to dynamically adapt to the distribution characteristics of the data, select the optimal insertion node, and thus maintain the balance and efficiency of the data structure.
[0013] Further, S13 includes the following sub-steps:
[0014] S131. Generate a random number;
[0015] S132. According to the primary key ID of the latest data report and the random number, determine the transfer set between the latest data report and the root node of the B+ tree index;
[0016] S133. According to the transfer set between the latest data report and the root node of the B+ tree index, determine the transfer set between the latest data report and each leaf node of the B+ tree index.
[0017] The beneficial effect of the above further solution is: In the present invention, by generating a random number, randomness is introduced when determining the transfer set, which helps to avoid excessive concentration of data on certain nodes and further optimizes the data distribution. By first determining the transfer set between the latest data report and the root node, and then deriving the transfer set between the latest data report and the leaf nodes, a step-by-step analysis from the top layer to the bottom layer is achieved, ensuring the accuracy and efficiency of the transfer set.
[0018] Further, in S132, the transfer set X root between the latest data report and the root node of the B+ tree index has the following expression: ; where p represents the random number, M represents the number of key values stored in the root node, L represents the number of characters included in the primary key ID of the latest data report, represents the ceiling operation.
[0019] Further, in S133, the expression for the transfer set between the latest data report and the i-th leaf node of the B+ tree index is:
[0020] ; where N1 represents the number of key values stored in the first leaf node, N2 represents the number of key values stored in the second leaf node, and N i represents the number of key values stored in the i-th leaf node, represents the ceiling operation, p represents a random number, and L represents the number of characters included in the primary key ID of the latest data report.
[0021] The beneficial effects of the above further solution are as follows: In the present invention, by precisely quantifying the transmission set, it is possible to scientifically determine the optimal position for data insertion based on parameters such as the primary key ID, random number, and the number of node key values. Considering the dynamic changes in data distribution (such as changes in the number of node key values), it is possible to dynamically adjust the transmission set according to the actual situation to maintain the balance of the data structure.
[0022] Further, S2 includes the following sub-steps:
[0023] S21. Determine the alternating scheduling index between adjacent leaf nodes according to the transmission set between the latest data report and each leaf node of the existing index;
[0024] S22. Extract the two adjacent leaf nodes corresponding to the maximum alternating scheduling index;
[0025] S23. Among the two adjacent leaf nodes, select the leaf node with the maximum fill factor as the optimal leaf node.
[0026] The beneficial effects of the above further solution are as follows: In the present invention, by calculating the alternating scheduling index between adjacent leaf nodes, the alternating scheduling index takes into account the maximum, minimum, and median of the transmission set, and can more comprehensively reflect the correlation between data and nodes. The design of the alternating scheduling index helps to avoid the over-concentration of data on certain nodes, thereby reducing the data skew problem. By extracting the two adjacent leaf nodes corresponding to the maximum alternating scheduling index, the method can accurately locate the candidate nodes most likely to be the optimal insertion position. Selecting the adjacent leaf node with the maximum alternating scheduling index as the candidate, and by selecting the leaf node with the maximum fill factor as the optimal leaf node, the storage utilization rate can be optimized. The fill factor represents the actual occupancy ratio of key values or data in the node. Selecting a node with a larger fill factor can reduce the waste of storage space. And a high fill factor (such as 80%) can reduce the splitting frequency.
[0027] Further, in S21, the alternating scheduling index w i,i+1 between the i-th leaf node and the (i + 1)-th leaf node is calculated as follows:
[0028] ; where Represents the maximum number of the transfer set between the latest data report and the $i$-th leaf node of the B+ tree index, Represents the minimum number of the transfer set between the latest data report and the $i$-th leaf node of the B+ tree index, Represents the median of the transfer set between the latest data report and the $i$-th leaf node of the B+ tree index, Represents the maximum number of the transfer set between the latest data report and the $(i + 1)$-th leaf node of the B+ tree index, Represents the minimum number of the transfer set between the latest data report and the $(i + 1)$-th leaf node of the B+ tree index, Represents the median of the transfer set between the latest data report and the $(i + 1)$-th leaf node of the B+ tree index, Represents the ceiling operation, Represents the maximum operation, Represents the minimum operation.
[0029] The beneficial effects of the present invention are as follows:
[0030] (1) The present invention uses a B+ tree as the index structure, significantly improving the data retrieval efficiency. The multi-way balanced characteristic of the B+ tree ensures the uniform distribution of data among nodes, which is suitable for the processing requirements of a large amount of real-time data in the sewage purification system;
[0031] (2) By calculating the alternating scheduling index between adjacent leaf nodes, the present invention can dynamically sense the change of data distribution and adjust the data insertion strategy accordingly, capable of coping with the dynamic change of data volume in the sewage purification system;
[0032] (3) When selecting the best leaf node, the present invention considers the fill factor of the node, preferentially selects the node with a higher fill factor to reduce the waste of storage space. At the same time, by balancing the split frequency and the overflow risk, it can improve the utilization rate of storage resources while ensuring the balance of the data structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Is a flowchart of the sewage purification data processing method. DETAILED DESCRIPTION OF THE INVENTION
[0034] The following further describes the embodiments of the present invention with reference to the drawings.
[0035] As Figure 1 shown, the present invention provides a sewage purification data processing method, including the following steps:
[0036] S1. Obtain the existing index of the sewage purification data packet. When generating the latest data report, determine the transfer set between the latest data report and each leaf node of the existing index;
[0037] S2. Generate the optimal leaf node based on the transfer set between the latest data report and each leaf node of the existing index.
[0038] S3. Insert the latest data report into the optimal leaf node to complete the update of the sewage purification data packet.
[0039] The structured storage of the sewage purification data packet adopts a B+ tree index structure, which can include basic attribute fields and dynamic monitoring data, etc. The basic attribute fields include the sewage treatment plant ID (primary key), name, type (above ground / underground), geographical location (longitude, latitude), designed treatment capacity, process category (such as A² / O, MBR), and discharge standard. The dynamic monitoring data includes water quality indicators such as chemical oxygen demand, ammonia nitrogen, total phosphorus, and pH value, etc., and also includes flow data such as influent flow, effluent flow, and sludge return flow, etc.
[0040] The latest data report is associated with the sewage purification data packet through the primary key ID, can adopt formats such as CSV, and can include abnormal event records and trend analysis results.
[0041] In the embodiment of the present invention, S1 includes the following sub-steps:
[0042] S11. Obtain the existing index of the sewage purification data packet; the existing index is a B+ tree index.
[0043] S12. When inserting the latest data report into the sewage purification data packet, obtain the primary key ID of the latest data report.
[0044] S13. Determine the transfer set between the latest data report and each leaf node of the B+ tree index according to the primary key ID of the latest data report.
[0045] In the present invention, the B+ tree is a multi-way balanced search tree with efficient retrieval and range query capabilities. In sewage purification data processing, the B+ tree index can quickly locate the data with a specific primary key ID. Each data report is uniquely identified by the primary key ID (such as timestamp, sensor ID, etc.), and the data that needs to be processed can be quickly located, improving the efficiency of data processing. By analyzing the transfer set between the latest data report and the B+ tree index nodes, the distribution characteristics of the data can be dynamically adapted, and the optimal insertion node can be selected to maintain the balance and efficiency of the data structure.
[0046] In the embodiment of the present invention, S13 includes the following sub-steps:
[0047] S131. Generate a random number.
[0048] S132. Determine the transfer set between the latest data report and the root node of the B+ tree index according to the primary key ID and random number of the latest data report;
[0049] S133. Determine the transfer set between the latest data report and each leaf node of the B+ tree index according to the transfer set between the latest data report and the root node of the B+ tree index.
[0050] In the present invention, by generating a random number, randomness is introduced when determining the transfer set, which helps to avoid the over-concentration of data on certain nodes and further optimizes the data distribution. By first determining the transfer set between the latest data report and the root node, and then deriving the transfer set between the leaf nodes, a step-by-step analysis from the top layer to the bottom layer is realized, ensuring the accuracy and efficiency of the transfer set.
[0051] In the embodiment of the present invention, in S132, the transfer set X root between the latest data report and the root node of the B+ tree index is expressed as:
[0052] ; where p represents the random number, M represents the number of key values stored in the root node, L represents the number of characters included in the primary key ID of the latest data report, represents the ceiling operation.
[0053] In the embodiment of the present invention, in S133, the expression of the transfer set between the latest data report and the i-th leaf node of the B+ tree index is:
[0054] ; where N1 represents the number of key values stored in the first leaf node, N2 represents the number of key values stored in the second leaf node, N i represents the number of key values stored in the i-th leaf node, represents the ceiling operation, p represents the random number, and L represents the number of characters included in the primary key ID of the latest data report.
[0055] In the present invention, by accurately quantifying the transfer set, the optimal position for data insertion can be scientifically determined based on parameters such as the primary key ID, random number, and the number of node key values. Considering the dynamic changes in data distribution (such as changes in the number of node key values), the transfer set can be dynamically adjusted according to the actual situation to maintain the balance of the data structure.
[0056] In the embodiment of the present invention, S2 includes the following sub-steps:
[0057] S21. Determine the alternating scheduling index between adjacent leaf nodes according to the transfer set between the latest data report and each leaf node of the existing index;
[0058] S22. Extract two adjacent leaf nodes corresponding to the maximum alternating scheduling index;
[0059] S23. Among the two adjacent leaf nodes, select the leaf node with the maximum fill factor as the optimal leaf node.
[0060] In the present invention, by calculating the alternating scheduling index between adjacent leaf nodes, the alternating scheduling index takes into account the maximum number, minimum number, and median of the transmission set, and can more comprehensively reflect the correlation between data and nodes. The design of the alternating scheduling index helps to avoid the over-concentration of data on certain nodes, thereby reducing the data skew problem. By extracting two adjacent leaf nodes corresponding to the maximum alternating scheduling index, the method can accurately locate the candidate nodes most likely to be the optimal insertion position. Selecting the adjacent leaf node with the maximum alternating scheduling index as the candidate, and by selecting the leaf node with the maximum fill factor as the optimal leaf node, the storage utilization rate can be optimized. The fill factor represents the actual occupancy ratio of key values or data in the node. Selecting a node with a larger fill factor can reduce the waste of storage space. And a high fill factor (such as 80%) can reduce the splitting frequency.
[0061] In the embodiment of the present invention, in S21, the alternating scheduling index w i,i+1 between the i-th leaf node and the (i + 1)-th leaf node is calculated as follows:
[0062] ; where represents the maximum number of the transmission set between the latest data report and the i-th leaf node of the B+ tree index, represents the minimum number of the transmission set between the latest data report and the i-th leaf node of the B+ tree index, represents the median of the transmission set between the latest data report and the i-th leaf node of the B+ tree index, represents the maximum number of the transmission set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the minimum number of the transmission set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the median of the transmission set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the ceiling operation, represents the maximum value operation, represents the minimum value operation.
[0063] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for processing sewage purification data, characterized in that, It includes the following steps: S1. Obtain the existing index of the sewage purification data packet. When generating the latest data report, determine the transmission set between the latest data report and each leaf node of the existing index; S2. Generate the optimal leaf node according to the transmission set between the latest data report and each leaf node of the existing index; S3. Insert the latest data report into the optimal leaf node to complete the update of the sewage purification data packet.
2. The sewage purification data processing method according to claim 1, characterized in that The S1 includes the following sub-steps: S11. Obtain the existing index of the sewage purification data packet; the existing index is a B+ tree index; S12. When inserting the latest data report into the sewage purification data packet, obtain the primary key ID of the latest data report; S13. Determine the transmission set between the latest data report and each leaf node of the B+ tree index according to the primary key ID of the latest data report.
3. The sewage purification data processing method according to claim 2, characterized in that The S13 includes the following sub-steps: S131. Generate a random number; S132. Determine the transmission set between the latest data report and the root node of the B+ tree index according to the primary key ID of the latest data report and the random number; S133. Determine the transmission set between the latest data report and each leaf node of the B+ tree index according to the transmission set between the latest data report and the root node of the B+ tree index.
4. The sewage purification data processing method according to claim 3, characterized in that In S132, the transmission set X between the latest data report and the root node of the B+ tree index root has the following expression: ; where p represents a random number, M represents the number of key values stored in the root node, L represents the number of characters included in the primary key ID of the latest data report, denotes the ceiling operation.
5. The sewage purification data processing method according to claim 4, wherein In the S133, the expression of the transmission set between the latest data report and the i-th leaf node of the B+ tree index is: ; where N1 represents the number of key values stored in the first leaf node, N2 represents the number of key values stored in the second leaf node, and N i represents the number of key values stored in the i-th leaf node, represents the ceiling operation, p represents a random number, and L represents the number of characters included in the primary key ID of the latest data report.
6. The sewage purification data processing method according to claim 1, wherein The S2 includes the following sub-steps: S21. Determine the alternating scheduling index between adjacent leaf nodes according to the transmission set between the latest data report and each leaf node of the existing index; S22. Extract the two adjacent leaf nodes corresponding to the maximum alternating scheduling index; S23. Select the leaf node with the maximum filling factor among the two adjacent leaf nodes as the optimal leaf node.
7. The sewage purification data processing method according to claim 6, wherein, In S21, the alternating scheduling index w between the i-th leaf node and the (i + 1)-th leaf node i,i+1 is calculated as follows: ; where, represents the maximum number of the transfer set between the latest data report and the i-th leaf node of the B+ tree index, represents the minimum number of the transfer set between the latest data report and the i-th leaf node of the B+ tree index, represents the median of the transfer set between the latest data report and the i-th leaf node of the B+ tree index, represents the maximum number of the transfer set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the minimum number of the transfer set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the median of the transfer set between the latest data report and the (i + 1)-th leaf node of the B+ tree index, represents the ceiling operation, represents the maximum value operation, represents the minimum value operation.
Citation Information
Patent Citations
Scalable database system for querying time-series data
CN110622152A
Character string data read optimization method based on B + tree and learning index
CN119597976A
Method for indexing data
US20230252012A1