Sewage treatment knowledge base construction method and system based on multi-source data analysis

By setting floating sampling points in the sewage test box, recording treatment operations and sewage parameters, and building a sewage treatment knowledge base based on multi-source data analysis, the problem of complex data and the need for manual screening in existing technologies is solved, and efficient and convenient knowledge base display is achieved.

CN120688599AActive Publication Date: 2025-09-23NANTONG JINGYUAN CLOUD COMPUTING TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510785417.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing sewage treatment knowledge base has complex data, poor readability, requires manual screening, and is inconvenient to use.

Method used

By setting floating sampling points in the sewage test box, recording the staff's treatment operations and sewage parameters, and building a sewage treatment knowledge base based on multi-source data analysis, the knowledge base can be verified and updated in real time.

Benefits of technology

It achieves a clear and intuitive display of the sewage treatment knowledge base, greatly improving convenience and readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688599A_ABST
    Figure CN120688599A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent sewage treatment, and particularly discloses a sewage treatment knowledge base construction method and system based on multi-source data analysis, and the method comprises the steps: recording the treatment operation of a worker, obtaining sewage parameters containing the position and time based on a floating sampling point, obtaining a test sample based on the same time scale statistical processing operation and the sewage parameters; segmenting the test sample to obtain a sample which is processed and operated to parameter variation, and naming the sample as a mapping sample; constructing a sewage treatment knowledge base according to the mapping samples; according to the method, processing operations are recorded based on the same time axis, sewage parameters are collected in real time, the variable quantity of the sewage parameters is segmented according to the processing operations, the variable quantity of the sewage parameters corresponding to each processing operation is obtained, the variable quantity of the sewage parameters is clustered and displayed, and a possible result of each processing operation is obtained. As a sewage treatment knowledge base of each treatment operation, the system is extremely high in readability and extremely high in convenience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent sewage treatment, and in particular to a method and system for constructing a sewage treatment knowledge base based on multi-source data analysis. Background Art

[0002] A "wastewater treatment knowledge base" is a database or document that systematically organizes information related to sewage treatment. It is used to record the correspondence between treatment operations and treatment results during the sewage treatment process. On the one hand, it can be used for employee training, and on the other hand, it can also serve as a reference for the sewage treatment process. Therefore, a high-quality sewage treatment knowledge base can bring great benefits.

[0003] However, the existing sewage treatment knowledge bases are actually some simple databases. For example, the statistical processing operations and processing results are in chronological order. The data are extremely complicated and have poor readability. When used, staff are required to perform manual screening, which is not very convenient. How to provide a data statistics solution for sewage treatment data to build a clearer and more intuitive sewage treatment knowledge base to facilitate the work of staff is the technical problem that the technical solution of the present invention wants to solve. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for constructing a sewage treatment knowledge base based on multi-source data analysis to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method and system for constructing a sewage treatment knowledge base based on multi-source data analysis, the method comprising:

[0007] Set floating sampling points in the sewage test tank and simultaneously determine the data type at the floating sampling points;

[0008] Record the treatment operations of the staff, obtain sewage parameters with location and time based on floating sampling points, and obtain test samples by statistically analyzing the treatment operations and sewage parameters on the same time scale;

[0009] The test samples are divided into sections to obtain samples that are processed into parameter changes, which are called mapping samples.

[0010] A sewage treatment knowledge base is constructed according to the mapping samples, the treatment process is verified in real time based on the sewage treatment knowledge base, and update instructions are triggered according to the verification results.

[0011] As a further solution of the present invention, the steps of setting a floating sampling point in the sewage test tank and simultaneously determining the data type at the floating sampling point include:

[0012] Obtain a three-dimensional model of the sewage test tank and locate the inner surface in the three-dimensional model;

[0013] Selecting reference points on each inner surface based on a predetermined density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface;

[0014] Obtain a ray pointing to the interior of the sewage test box that passes through the reference point and is perpendicular to the corresponding inner surface. Intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box;

[0015] Select floating sampling points on the line segment based on a preset step size, calculate the distance between any selected floating sampling point and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold;

[0016] For any data type, query the demand ratio of the data type and allocate floating sampling points according to the demand ratio.

[0017] As a further solution of the present invention, the steps of recording the processing operations of the staff, obtaining sewage parameters including location and time based on floating sampling points, and obtaining test samples based on statistical processing operations and sewage parameters at the same time scale include:

[0018] Record the processing operations of staff at each moment;

[0019] Obtaining the relative coordinates of the floating sampling point in the pollution test box based on the rangefinder at the floating sampling point;

[0020] The sewage parameters are obtained based on the pollution collector at the floating sampling point, and the collection time is recorded synchronously to obtain the sewage parameters containing relative coordinates and collection time;

[0021] Regularize the time based on the preset time step to obtain the time period;

[0022] Obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template;

[0023] Based on the same time axis statistical processing operation and parameter matrix, the test samples are obtained.

[0024] As a further solution of the present invention: the step of recording the processing operations of the staff at each moment includes:

[0025] Receive the operation information containing the time period input by the staff based on the preset information receiving port;

[0026] Real-time positioning of staff based on cameras, triggering recognition instructions when staff are detected;

[0027] Obtaining a trigger time of a trigger identification instruction, and when the trigger time is included in the time period, identifying the behavior of the staff member based on a preset first accuracy and verifying the operation information;

[0028] When the trigger time is not included in the time period, the staff's behavior is identified based on the preset second precision to obtain the operation information;

[0029] In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

[0030] As a further solution of the present invention, the step of segmenting the test sample to obtain samples processed to parameter variation, called mapping samples, includes:

[0031] For any processing operation at any moment, the parameter matrix before processing is obtained according to the preset forward span, and the parameter matrix after processing is obtained according to the preset backward span;

[0032] Calculate the difference matrix between the parameter matrix after processing and the parameter matrix before processing;

[0033] The samples that are constructed by processing operations to the difference matrix are called mapped samples.

[0034] As a further solution of the present invention, the steps of constructing a sewage treatment knowledge base according to the mapping samples, verifying the treatment process in real time based on the sewage treatment knowledge base, and triggering an update instruction according to the verification result include:

[0035] Count all mapping samples, classify the mapping samples according to the processing operations in the mapping samples, and obtain a sample set corresponding to each processing operation;

[0036] For each sample set of processing operation, the samples therein are compared pairwise to calculate the similarity; the similarity adopts the similarity of the difference matrix;

[0037] Cluster the samples in the sample set based on similarity, and calculate the sample ratio and mean sample of each clustering result; the sample ratio is the number of samples in the clustering result divided by the total number of samples in the sample set, and the mean sample is the mean matrix of the difference matrix of all samples in the clustering result;

[0038] The sample ratio and its mean sample are regarded as a data item, and the data items are sorted in descending order based on the sample ratio to obtain a data table for processing operations;

[0039] Statistical data tables of all treatment operations to obtain a knowledge base of sewage treatment;

[0040] In actual application, for any processing operation of the staff, a theoretical prediction result is obtained based on the sewage treatment knowledge base, and the accuracy of the actual processing result is determined based on the theoretical prediction result. When the accuracy is less than the preset accuracy threshold, the processing operation and its actual processing result are used as mapping samples to trigger an update instruction; wherein, the process of using the processing operation and its actual processing result as mapping samples at least includes regularizing the actual processing results based on a matrix template.

[0041] The technical solution of the present invention also provides a sewage treatment knowledge base construction system based on multi-source data analysis, the system comprising:

[0042] Sampling point setting module, used to set floating sampling points in the sewage test tank and simultaneously determine the data type at the floating sampling points;

[0043] The test sample generation module is used to record the processing operations of the staff, obtain sewage parameters containing location and time based on floating sampling points, and obtain test samples based on the statistical processing operations and sewage parameters on the same time scale;

[0044] The test sample segmentation module is used to segment the test sample and obtain samples of the processed operation to the parameter change amount, which are called mapping samples;

[0045] The knowledge base construction module is used to construct a sewage treatment knowledge base according to the mapping samples, verify the treatment process in real time based on the sewage treatment knowledge base, and trigger the update instruction according to the verification results.

[0046] As a further solution of the present invention: the sampling point setting module includes:

[0047] An inner surface positioning unit, used for obtaining a three-dimensional model of the sewage test box and locating the inner surface in the three-dimensional model;

[0048] A reference point selection unit, configured to select reference points on each inner surface based on a preset density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface;

[0049] A line segment interception unit is used to obtain a ray pointing to the interior of the sewage test box, passing through the reference point and perpendicular to the corresponding inner surface, and intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box;

[0050] A sampling point deletion unit is used to select floating sampling points on the line segment based on a preset step size, and for any selected floating sampling point, calculate the distance between it and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold;

[0051] The sampling point allocation unit is used to query the demand ratio of any data type and allocate floating sampling points according to the demand ratio.

[0052] As a further solution of the present invention: the test sample generation module includes:

[0053] Operation recording unit, used to record the processing operations of staff at each moment;

[0054] A coordinate acquisition unit, configured to acquire the relative coordinates of the floating sampling point in the pollution test box based on a rangefinder at the floating sampling point;

[0055] A parameter collection unit is used to obtain sewage parameters based on the pollution collector at the floating sampling point, synchronously record the collection time, and obtain sewage parameters containing relative coordinates and collection time;

[0056] A time warping unit, used to warp time based on a preset time step to obtain a time period;

[0057] A parameter matrix construction unit is used to obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template;

[0058] The operation statistics unit is used to statistically process the operation and parameter matrix based on the same time axis to obtain test samples.

[0059] As a further solution of the present invention: the operation recording unit includes:

[0060] An information receiving subunit, configured to receive operation information including a time period input by a staff member based on a preset information receiving port;

[0061] The personnel positioning subunit is used to locate the staff in real time based on the camera and trigger the recognition command when the staff is detected;

[0062] A first identification subunit is configured to obtain a trigger time of a trigger identification instruction, and when the trigger time is included in a time period, identify the behavior of the staff member based on a preset first accuracy and verify the operation information;

[0063] The second identification subunit is configured to identify the behavior of the staff member based on a preset second accuracy and obtain operation information when the trigger time is not included in the time period;

[0064] In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] The present invention records treatment operations based on the same time axis and collects sewage parameters in real time. The changes in sewage parameters are divided according to the treatment operations to obtain the changes in sewage parameters corresponding to each treatment operation. The changes in sewage parameters are clustered and displayed to obtain the possible results of each treatment operation. As a sewage treatment knowledge base for each treatment operation, the present invention has strong readability and is extremely convenient to use. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention.

[0068] Figure 1 The overall flow chart of the sewage treatment knowledge base construction method based on multi-source data analysis is shown.

[0069] Figure 2 The structural diagram of the sewage treatment knowledge base construction system based on multi-source data analysis is shown. DETAILED DESCRIPTION

[0070] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0071] Figure 1 The following is a general flow chart of a method and system for constructing a sewage treatment knowledge base based on multi-source data analysis. In an embodiment of the present invention, a method for constructing a sewage treatment knowledge base based on multi-source data analysis includes:

[0072] Step S100: setting a floating sampling point in the sewage test tank and simultaneously determining the data type at the floating sampling point;

[0073] The technical solution of the present invention is used to build a sewage treatment knowledge base, therefore, a large amount of testing work is required. The testing work takes place in a sewage test box. Sewage is placed in the pollution test box, and floating sampling points are set in the sewage test box to collect the state of the sewage at each moment. When the staff takes treatment measures, the sewage data before and after treatment are completely recorded and stored in a preset database. It should be noted that the state of the sewage at each moment is actually an upper probability, which is actually a parameter of different data types, such as the concentration of different components.

[0074] Step S200: Record the processing operations of the staff, obtain sewage parameters including location and time based on the floating sampling points, and obtain test samples based on the statistical processing operations and sewage parameters at the same time scale;

[0075] The treatment operation of the staff refers to the sewage treatment process. In actual applications, it mainly refers to when and how much treatment reagents were added. Of course, there are also some physical means, such as when to use filtration and salvage operations. These are collectively referred to as treatment operations. From a computer perspective, treatment operations are some numerical or text data; then, collection equipment is installed at the floating sampling point to obtain sewage parameters containing location and time. A timeline is constructed at time zero, which is the starting time of the test process. Based on the timeline, all treatment operations and sewage parameters in a test process are counted to obtain a sample corresponding to a test process, which is called a test sample.

[0076] It should be noted that the meaning of the floating sampling point in the technical solution of the present invention is that the floating sampling point is a sampling point whose position will change. For example, a floating collection device is connected to the bottom of the sewage test box with a flexible material (line). As the water flow changes, the position of the collection device will continue to change. Therefore, when obtaining sewage parameters, in addition to recording time, it is also necessary to record the position; regarding the position identification process, a rangefinder pointing to three mutually perpendicular directions is installed on the collection device, one of which points to the bottom of the collection device. In this way, the position of the collection device in the sewage test box can be obtained in real time.

[0077] Step S300: Segment the test sample to obtain samples that have been processed and converted into parameter changes, which are called mapping samples.

[0078] The test sample contains processing operations and pollution parameters, and their time relationship is very clear. For each processing operation, the change of pollution parameters before and after it is obtained, which is called parameter change. Then, a sample of processing operations and parameter change is constructed, which is called a mapping sample.

[0079] Step S400: constructing a sewage treatment knowledge base based on the mapping sample, verifying the treatment process in real time based on the sewage treatment knowledge base, and triggering an update instruction based on the verification result;

[0080] Each test can obtain a test sample, and then obtain multiple mapping samples. By continuously conducting multiple tests (these tests can be carried out in multiple plant areas), a large number of mapping samples can be obtained. A sewage treatment knowledge base is constructed based on the mapping samples. The sewage treatment knowledge base is used to characterize the relationship between treatment operations and sewage treatment results. In actual applications, the sewage treatment knowledge base can be used as a processing reference, and during processing, the treatment process is verified in real time based on the sewage treatment knowledge base to obtain the application accuracy of the sewage treatment knowledge base. If the application accuracy is low, the sewage treatment knowledge base needs to be updated, that is, the update instruction is triggered.

[0081] It should be noted that the goal of updating the technical solution of the present invention is to test samples, and the actual application process is also regarded as a test, and the acquired data is used as a new test sample. When the test sample changes, the sewage treatment knowledge base will also change.

[0082] Regarding step S100, the steps of setting floating sampling points in the sewage test tank and simultaneously determining the data types at the floating sampling points include:

[0083] Obtain a three-dimensional model of the sewage test tank and locate the inner surface in the three-dimensional model;

[0084] Selecting reference points on each inner surface based on a predetermined density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface;

[0085] Obtain a ray pointing to the interior of the sewage test box that passes through the reference point and is perpendicular to the corresponding inner surface. Intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box;

[0086] Select floating sampling points on the line segment based on a preset step size, calculate the distance between any selected floating sampling point and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold;

[0087] For any data type, query the demand ratio of the data type and allocate floating sampling points according to the demand ratio.

[0088] The structure of a sewage test tank is generally simple. It is a box for holding sewage. The box is actually a general concept. Some pits that can hold sewage can also be used as a box. The sewage test tank is modeled to obtain a three-dimensional model. The inner surface is located in the three-dimensional model. Based on a preset density, reference points are selected on each inner surface. This process is used to select some points on the inner surface of the sewage test tank, called reference points. Rays passing through the reference points and perpendicular to the inner surface pointing into the sewage test tank are obtained. The portion of the ray contained in the sewage test tank is called a line segment. Floating sampling points are selected on the line segment based on a preset step size. For any selected floating sampling point, its distance to the nearest floating sampling point is calculated. When the distance is less than a preset distance threshold, the floating sampling point is deleted (only one is retained to reduce the number of floating sampling points). For any data type, the demand ratio of the data type is queried and the floating sampling points are allocated according to the demand ratio. The demand ratio indicates the proportion of floating sampling points required for each type of data. The number of floating sampling points corresponding to each data type is obtained by multiplying the demand ratio by the total number. The data type is generally the type of pollutant to be monitored in sewage.

[0089] Regarding step S200, the steps of recording the processing operations of the staff, obtaining sewage parameters including location and time based on the floating sampling points, and obtaining the test samples based on the statistical processing operations and sewage parameters at the same time scale include:

[0090] Record the processing operations of staff at each moment;

[0091] Obtaining the relative coordinates of the floating sampling point in the pollution test box based on the rangefinder at the floating sampling point;

[0092] The sewage parameters are obtained based on the pollution collector at the floating sampling point, and the collection time is recorded synchronously to obtain the sewage parameters containing relative coordinates and collection time;

[0093] Regularize the time based on the preset time step to obtain the time period;

[0094] Obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template;

[0095] Based on the same time axis statistical processing operation and parameter matrix, the test samples are obtained.

[0096] When the staff is performing the operation, the technical solution of the present invention records in real time, recording the processing operations of the staff at each moment, and obtains the relative coordinates of the floating sampling point in the pollution test box based on the rangefinder at the floating sampling point. The relative coordinates are the position, and the sewage parameters are obtained based on the pollution collector at the floating sampling point. The collection time is synchronously recorded to obtain sewage parameters containing relative coordinates and collection time; since the collection time is relatively instantaneous and it will be affected by the transmission process, the technical solution of the present invention needs to regularize the time into units of one minute, 30 seconds or 10 seconds, determine some time periods according to the time step, and regard the sewage parameters belonging to the same time period as the sewage parameters of the same time, obtain the sewage parameters containing relative coordinates in each time period, and construct a parameter matrix based on a preset matrix template. After this operation, the parameter matrix of each time period can be obtained, and the test sample is obtained based on the statistical processing operation and parameter matrix on the same time axis.

[0097] Specifically, since the number of floating sampling points is limited and they must be assigned to different data types, there are very few floating sampling points corresponding to each data type, and the real data obtained is actually very little. For this situation, the technical solution of the present invention introduces a data expansion process. For any position in the parameter matrix, the nearest real data is queried and the real data is filled into the unknown position. In addition, regarding the matrix template, it is a three-dimensional matrix with fixed length, width and height. Each data type corresponds to a matrix template, and each data type can eventually obtain a parameter matrix for different time periods.

[0098] As a preferred embodiment of the technical solution of the present invention, the step of recording the processing operations of the staff at each moment includes:

[0099] Receive the operation information containing the time period input by the staff based on the preset information receiving port;

[0100] Real-time positioning of staff based on cameras, triggering recognition instructions when staff are detected;

[0101] Obtaining a trigger time of a trigger identification instruction, and when the trigger time is included in the time period, identifying the behavior of the staff member based on a preset first accuracy and verifying the operation information;

[0102] When the trigger time is not included in the time period, the staff's behavior is identified based on the preset second precision to obtain the operation information;

[0103] In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

[0104] Before the staff performs the operation, the operation information will be input on the information receiving port. For the execution subject of the technical solution of the present invention, it is only necessary to execute the data receiving process. In an example of the technical solution of the present invention, a verification process is also introduced on the basis of receiving the data. The staff is located in real time based on the camera. When the staff is detected, the recognition instruction is triggered. The process of the camera locating the staff is real-time. Considering the cost issue, its accuracy is low, which is the third accuracy in the above content; further, the trigger time of the trigger recognition instruction is obtained. When the trigger time is included in the time period, the behavior of the staff is identified based on the preset first accuracy, and the operation information is verified. The first accuracy is higher than the third accuracy, but not much higher. This means that the staff is operating according to the preset rules and only needs to be verified; when the trigger time is not included in the time period, the staff's behavior is identified based on the preset second precision to obtain operation information. The second precision is the highest because it means that the staff did not upload the task in advance and appeared in the work scene. At this time, the staff needs to be identified in real time to obtain whether he has performed any operations that have not been uploaded. In fact, when the staff did not upload the task in advance and appeared in the work scene, the accuracy of monitoring is improved. A very important reason is that he may be a non-staff member and needs to be identified in real time by the background. Once a certain risk is found (for example, the person appears in a risky location), a warning message is generated.

[0105] Regarding step S300, the step of dividing the test sample to obtain samples of the processed operation to the parameter variation, which is called the step of mapping the samples, includes:

[0106] For any processing operation at any moment, the parameter matrix before processing is obtained according to the preset forward span, and the parameter matrix after processing is obtained according to the preset backward span;

[0107] Calculate the difference matrix between the parameter matrix after processing and the parameter matrix before processing;

[0108] The samples that are constructed by processing operations to the difference matrix are called mapped samples.

[0109] The forward span and backward span are both time differences, which respectively indicate how long it takes to obtain data forward and how long it takes to obtain data backward. For the processing operation at any moment, the parameter matrix before processing is obtained according to the preset forward span, and the parameter matrix after processing is obtained according to the preset backward span. The difference matrix between the parameter matrix after processing and the parameter matrix before processing is calculated, and the sample from the processing operation to the difference matrix is ​​constructed, which is called a mapping sample.

[0110] Regarding step S400, the steps of constructing a sewage treatment knowledge base based on the mapping sample, verifying the treatment process in real time based on the sewage treatment knowledge base, and triggering an update instruction based on the verification result include:

[0111] Count all mapping samples, classify the mapping samples according to the processing operations in the mapping samples, and obtain a sample set corresponding to each processing operation;

[0112] For each sample set of processing operation, the samples therein are compared pairwise to calculate the similarity; the similarity adopts the similarity of the difference matrix;

[0113] Cluster the samples in the sample set based on similarity, and calculate the sample ratio and mean sample of each clustering result; the sample ratio is the number of samples in the clustering result divided by the total number of samples in the sample set, and the mean sample is the mean matrix of the difference matrix of all samples in the clustering result;

[0114] The sample ratio and its mean sample are regarded as a data item, and the data items are sorted in descending order based on the sample ratio to obtain a data table for processing operations;

[0115] Statistical data tables of all treatment operations are used to obtain a knowledge base of sewage treatment.

[0116] In an example of the technical solution of the present invention, the application and update process of the sewage treatment knowledge base is described, all mapping samples are counted, and the mapping samples are classified according to the processing operations in the mapping samples to obtain a sample set corresponding to each processing operation. For each sample set of processing operation, the samples therein are compared pairwise, and the similarity of the difference matrix is ​​calculated; based on the similarity, the samples in the sample set are clustered to obtain a clustering result, and the clustering result is a set for each type of sample, which belongs to a subset of the sample set. The ratio of the number of samples in each clustering result to the total number of samples in the sample set is calculated, and then the number of samples in each clustering result is divided into the number of samples in the sample set. It is converted into a numerical value in the range of zero to one, which is called the sample proportion. Then, for each clustering result, the mean matrix of all the difference matrices is calculated (a matrix composed of the means at each row and column position), and the mean matrix is ​​arranged according to the sample proportion. The arrangement method adopts the descending order of the sample proportion. After the arrangement is completed, a data table is obtained, which is called the sewage treatment knowledge base. After the above processing, a sewage treatment knowledge base can be obtained for each treatment operation. The actual meaning of the sewage treatment knowledge base is that it represents the possible situations of each treatment operation. The closer the situation is, the more likely it is to occur (the corresponding clustering result has a larger number of samples).

[0117] Furthermore, regarding the update process, in actual application, for any processing operation of the staff, the theoretical prediction results are obtained based on the sewage treatment knowledge base, and the accuracy of the actual processing results is determined based on the theoretical prediction results. When the accuracy is less than the preset accuracy threshold, the processing operation and its actual processing results are used as mapping samples to trigger the update instruction.

[0118] It should be noted that the process of using the processing operation and its actual processing results as mapping samples at least includes regularizing the actual processing results based on the matrix template. The regularization method is very simple, that is, converting the data of the actual processing results into the matrix template. The conversion process adopts some conventional upsampling and downsampling methods. In fact, in the technical solution of the present invention, since matrices need to be frequently operated on, their sizes need to be the same. For matrices of different sizes, upsampling operations (increasing the number of rows and columns) and downsampling operations (reducing the number of rows and columns) are required to adjust the matrix size.

[0119] Figure 2 The structure diagram of the sewage treatment knowledge base construction system based on multi-source data analysis is shown. In a preferred embodiment of the technical solution of the present invention, a sewage treatment knowledge base construction system based on multi-source data analysis is also provided. The system 10 includes:

[0120] The sampling point setting module 11 is used to set floating sampling points in the sewage test tank and simultaneously determine the data type at the floating sampling points;

[0121] The test sample generation module 12 is used to record the processing operations of the staff, obtain sewage parameters including location and time based on the floating sampling points, and obtain test samples based on the statistical processing operations and sewage parameters on the same time scale;

[0122] The test sample segmentation module 13 is used to segment the test sample to obtain samples of parameter changes after processing, which are called mapping samples;

[0123] The knowledge base construction module 14 is used to construct a sewage treatment knowledge base according to the mapping samples, verify the treatment process in real time based on the sewage treatment knowledge base, and trigger an update instruction according to the verification result.

[0124] Furthermore, the sampling point setting module 11 includes:

[0125] An inner surface positioning unit, used for obtaining a three-dimensional model of the sewage test box and locating the inner surface in the three-dimensional model;

[0126] A reference point selection unit, configured to select reference points on each inner surface based on a preset density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface;

[0127] A line segment interception unit is used to obtain a ray pointing to the interior of the sewage test box, passing through the reference point and perpendicular to the corresponding inner surface, and intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box;

[0128] A sampling point deletion unit is used to select floating sampling points on the line segment based on a preset step size, and for any selected floating sampling point, calculate the distance between it and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold;

[0129] The sampling point allocation unit is used to query the demand ratio of any data type and allocate floating sampling points according to the demand ratio.

[0130] Specifically, the test sample generation module 12 includes:

[0131] Operation recording unit, used to record the processing operations of staff at each moment;

[0132] A coordinate acquisition unit, configured to acquire the relative coordinates of the floating sampling point in the pollution test box based on a rangefinder at the floating sampling point;

[0133] A parameter collection unit is used to obtain sewage parameters based on the pollution collector at the floating sampling point, synchronously record the collection time, and obtain sewage parameters containing relative coordinates and collection time;

[0134] A time warping unit, used to warp time based on a preset time step to obtain a time period;

[0135] A parameter matrix construction unit is used to obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template;

[0136] The operation statistics unit is used to statistically process the operation and parameter matrix based on the same time axis to obtain test samples.

[0137] Furthermore, the operation recording unit includes:

[0138] An information receiving subunit, configured to receive operation information including a time period input by a staff member based on a preset information receiving port;

[0139] The personnel positioning subunit is used to locate the staff in real time based on the camera and trigger the recognition command when the staff is detected;

[0140] A first identification subunit is configured to obtain a trigger time of a trigger identification instruction, and when the trigger time is included in a time period, identify the behavior of the staff member based on a preset first accuracy and verify the operation information;

[0141] The second identification subunit is configured to identify the behavior of the staff member based on a preset second accuracy and obtain operation information when the trigger time is not included in the time period;

[0142] In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

[0143] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for constructing a sewage treatment knowledge base based on multi-source data analysis, characterized in that: The method comprises: Set floating sampling points in the sewage test tank and simultaneously determine the data type at the floating sampling points; Record the treatment operations of the staff, obtain sewage parameters with location and time based on floating sampling points, and obtain test samples by statistically analyzing the treatment operations and sewage parameters on the same time scale; The test samples are divided into sections to obtain samples that are processed into parameter changes, which are called mapping samples. A sewage treatment knowledge base is constructed according to the mapping samples, the treatment process is verified in real time based on the sewage treatment knowledge base, and update instructions are triggered according to the verification results.

2. The method for constructing a sewage treatment knowledge base based on multi-source data analysis according to claim 1, characterized in that: The steps of setting floating sampling points in the sewage test tank and simultaneously determining the data types at the floating sampling points include: Obtain a three-dimensional model of the sewage test tank and locate the inner surface in the three-dimensional model; Selecting reference points on each inner surface based on a predetermined density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface; Obtain a ray pointing to the interior of the sewage test box that passes through the reference point and is perpendicular to the corresponding inner surface. Intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box; Select floating sampling points on the line segment based on a preset step size, calculate the distance between any selected floating sampling point and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold; For any data type, query the demand ratio of the data type and allocate floating sampling points according to the demand ratio.

3. The method for constructing a sewage treatment knowledge base based on multi-source data analysis according to claim 1, characterized in that: The steps of recording the processing operations of the staff, obtaining sewage parameters including location and time based on the floating sampling points, and obtaining the test samples based on the statistical processing operations and sewage parameters at the same time scale include: Record the processing operations of staff at each moment; Obtaining the relative coordinates of the floating sampling point in the pollution test box based on the rangefinder at the floating sampling point; The sewage parameters are obtained based on the pollution collector at the floating sampling point, and the collection time is recorded synchronously to obtain the sewage parameters containing relative coordinates and collection time; Regularize the time based on the preset time step to obtain the time period; Obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template; Based on the same time axis statistical processing operation and parameter matrix, the test samples are obtained.

4. The method for constructing a sewage treatment knowledge base based on multi-source data analysis according to claim 3 is characterized in that: The steps of recording the processing operations of the staff at each moment include: Receive the operation information containing the time period input by the staff based on the preset information receiving port; Real-time positioning of staff based on cameras, triggering recognition instructions when staff are detected; Obtaining a trigger time of a trigger identification instruction, and when the trigger time is included in the time period, identifying the behavior of the staff member based on a preset first accuracy and verifying the operation information; When the trigger time is not included in the time period, the staff's behavior is identified based on the preset second precision to obtain the operation information; In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

5. The method for constructing a sewage treatment knowledge base based on multi-source data analysis according to claim 1, characterized in that: The step of dividing the test sample to obtain samples processed into parameter variation values, which is called mapping samples, includes: For any processing operation at any moment, the parameter matrix before processing is obtained according to the preset forward span, and the parameter matrix after processing is obtained according to the preset backward span; Calculate the difference matrix between the parameter matrix after processing and the parameter matrix before processing; The samples that are constructed by processing operations to the difference matrix are called mapped samples.

6. The method for constructing a sewage treatment knowledge base based on multi-source data analysis according to claim 5, characterized in that: The steps of constructing a sewage treatment knowledge base according to the mapping sample, verifying the treatment process in real time based on the sewage treatment knowledge base, and triggering an update instruction according to the verification result include: Count all mapping samples, classify the mapping samples according to the processing operations in the mapping samples, and obtain a sample set corresponding to each processing operation; For each sample set of processing operation, the samples therein are compared pairwise to calculate the similarity; the similarity adopts the similarity of the difference matrix; Cluster the samples in the sample set based on similarity, and calculate the sample ratio and mean sample of each clustering result; the sample ratio is the number of samples in the clustering result divided by the total number of samples in the sample set, and the mean sample is the mean matrix of the difference matrix of all samples in the clustering result; The sample ratio and its mean sample are regarded as a data item, and the data items are sorted in descending order based on the sample ratio to obtain a data table for processing operations; Statistical data tables of all treatment operations to obtain a knowledge base of sewage treatment; In actual application, for any processing operation of the staff, a theoretical prediction result is obtained based on the sewage treatment knowledge base, and the accuracy of the actual processing result is determined based on the theoretical prediction result. When the accuracy is less than the preset accuracy threshold, the processing operation and its actual processing result are used as mapping samples to trigger an update instruction; wherein, the process of using the processing operation and its actual processing result as mapping samples at least includes regularizing the actual processing results based on a matrix template.

7. A sewage treatment knowledge base construction system based on multi-source data analysis, characterized in that: The system comprises: Sampling point setting module, used to set floating sampling points in the sewage test tank and simultaneously determine the data type at the floating sampling points; The test sample generation module is used to record the processing operations of the staff, obtain sewage parameters containing location and time based on floating sampling points, and obtain test samples based on the statistical processing operations and sewage parameters on the same time scale; The test sample segmentation module is used to segment the test sample and obtain samples of the processed operation to the parameter change amount, which are called mapping samples; The knowledge base construction module is used to construct a sewage treatment knowledge base according to the mapping samples, verify the treatment process in real time based on the sewage treatment knowledge base, and trigger the update instruction according to the verification results.

8. The sewage treatment knowledge base construction system based on multi-source data analysis according to claim 7 is characterized in that: The sampling point setting module includes: An inner surface positioning unit, used for obtaining a three-dimensional model of the sewage test box and locating the inner surface in the three-dimensional model; A reference point selection unit, configured to select reference points on each inner surface based on a preset density; the density being the number of reference points per unit area, the density being determined by the number of sides of the inner surface; A line segment interception unit is used to obtain a ray pointing to the interior of the sewage test box, passing through the reference point and perpendicular to the corresponding inner surface, and intercept the ray based on the box area of ​​the sewage test box to obtain a line segment corresponding to each reference point; wherein the box area is not smaller than the space of the sewage test box; A sampling point deletion unit is used to select floating sampling points on the line segment based on a preset step size, and for any selected floating sampling point, calculate the distance between it and the nearest floating sampling point, and delete the floating sampling point when the distance is less than a preset distance threshold; The sampling point allocation unit is used to query the demand ratio of any data type and allocate floating sampling points according to the demand ratio.

9. The sewage treatment knowledge base construction system based on multi-source data analysis according to claim 7 is characterized in that: The test sample generation module includes: Operation recording unit, used to record the processing operations of staff at each moment; A coordinate acquisition unit, configured to acquire the relative coordinates of the floating sampling point in the pollution test box based on a rangefinder at the floating sampling point; A parameter collection unit is used to obtain sewage parameters based on the pollution collector at the floating sampling point, synchronously record the collection time, and obtain sewage parameters containing relative coordinates and collection time; A time warping unit, used to warp time based on a preset time step to obtain a time period; A parameter matrix construction unit is used to obtain sewage parameters with relative coordinates in each time period and construct a parameter matrix based on a preset matrix template; The operation statistics unit is used to statistically process the operation and parameter matrix based on the same time axis to obtain test samples.

10. The sewage treatment knowledge base construction system based on multi-source data analysis according to claim 9 is characterized in that: The operation recording unit includes: An information receiving subunit, configured to receive operation information containing a time period input by a staff member based on a preset information receiving port; The personnel positioning subunit is used to locate the staff in real time based on the camera and trigger the recognition command when the staff is detected; A first identification subunit is configured to obtain a trigger time of a trigger identification instruction, and when the trigger time is included in a time period, identify the behavior of the staff member based on a preset first accuracy and verify the operation information; The second identification subunit is configured to identify the behavior of the staff member based on a preset second accuracy and obtain operation information when the trigger time is not included in the time period; In which, when the camera does not trigger the recognition instruction, it works under a preset third precision condition, the second precision is greater than the first precision, and the first precision is greater than the third precision.

Citation Information

Patent Citations

  • Condition recognition based multi-target optimization control method of sewage treatment process

    CN110262431A

  • Sewage treatment sample data management method and system based on block chain and big data

    CN113919971A

  • Steel structure deformation positioning method and device, computer equipment and storage medium

    CN115655128A

  • Sewage treatment plant diagnosis method and system based on multi-source heterogeneous diagnosis knowledge base

    CN118026308A

  • Industrial sewage treatment method and system based on big data

    CN119359092A