Municipal underground pipeline inspection equipment and method
By segmenting and merging the multi-dimensional environmental data collected by the underground pipeline inspection robot and selecting representative points, the problem of inaccurate abnormal data detection caused by the difference in the position of the inspection robot is solved, and more accurate abnormal data analysis is achieved.
Patent Information
- Application Number
- CN202510933638.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-08
AI Technical Summary
When the existing intelligent inspection robot system conducts abnormal analysis of multi-dimensional environmental data, due to the position difference caused by the motion state of the inspection robot in the pipeline, fixed data thresholds lead to inaccurate abnormal data detection results.
By analyzing the correlation changes between environmental data in different dimensions, segment the single-dimensional environmental data, merge data segments, select representative points, combine trend differences and intra-cluster distances to obtain abnormal data.
It improves the accuracy of abnormal data detection, reduces interference to abnormal data, and reflects the environmental characteristics of the area and time period in which the patrol robot is located.
Smart Images

Figure CN120449059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pipeline inspection data analysis, and in particular to municipal underground pipeline inspection equipment and method. Background Art
[0002] Municipal underground pipeline inspection equipment plays a vital role in the maintenance of modern urban infrastructure. These devices not only improve inspection efficiency but also significantly reduce the labor intensity and potential risks of manual operations. Rail-mounted or wheeled intelligent inspection robots are used to conduct inspections within integrated pipeline corridors. Equipped with high-definition cameras, sensors, and other equipment, these robots can collect pipeline data in real time. Intelligent inspection robots can also enter pipeline corridors through remote control or autonomous navigation for comprehensive inspections. The intelligent inspection robot system enables continuous and complete data collection, requiring no human intervention. The system automatically uploads data to a server for analysis and displays abnormal data in reports, thereby providing alerts for potential hazards and demonstrating the current status of potential risk points.
[0003] When existing intelligent inspection robot systems conduct abnormal data analysis of multi-dimensional environmental data, they generally analyze and judge abnormal hidden dangers by setting data thresholds. However, the intelligent inspection robots are in motion in the pipeline, and their positions in the pipeline are different at different times, which will lead to large differences in environmental data in various dimensions. Fixed data thresholds will lead to inaccurate abnormal data detection results. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a municipal underground pipeline inspection device and method.
[0005] According to a first aspect of an embodiment of the present application, a municipal underground pipeline inspection method is provided, and the technical solution adopted specifically includes: Collect multi-dimensional environmental data inside underground pipelines; Analyze the correlation changes between environmental data of different dimensions, segment the single-dimensional environmental data, and obtain several data segments of environmental data of each dimension; For each dimensional environmental data, the length of the data segment and the change characteristics of the data in the data segment are analyzed, and the adjusted distance between different data segments in each dimensional environmental data is obtained by combining the inter-segment distance between the data segments. The data segments are merged to obtain a merged cluster of each dimensional environmental data; Analyze the correlation between the data segment and the data segments corresponding to each other dimensional environmental data to obtain a correlation description vector for each data segment, and then obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster, and select a representative point of the merged cluster; The trend difference between the representative point and the centroid of the merged cluster is analyzed, and the shrinkage factor of the representative point is obtained in combination with the adjustment of the intra-cluster distance, and the merged clusters are merged to obtain abnormal data.
[0006] In some embodiments of the present invention, the correlation changes between environmental data of different dimensions are analyzed, and the single-dimensional environmental data is segmented to obtain several data segments corresponding to each dimensional environmental data, including: Calculate the Pearson correlation coefficient of the target dimension environment data and any other dimension environment data one by one starting from the first environment data, and record the sign of the difference between the Pearson correlation coefficients of two adjacent target dimension environment data to obtain the inflection point; Using the inflection point as a cutoff point, obtaining an initial data segment corresponding to the target dimension environment data and each other dimension environment data; Using the initial data segment with the smallest segment length among the multiple initial data segments as the first data segment of the target dimension environment data; In the target dimension environment data, the environment data after the first data segment is used as a starting point to obtain the second data segment of the target dimension environment data; And so on, several data segments corresponding to each dimension of environmental data are obtained.
[0007] In some embodiments of the present invention, for each dimension of environmental data, the length of the data segment and the change characteristics of the data within the data segment are analyzed, and the adjustment distance between different data segments in each dimension of environmental data is obtained by combining the inter-segment distance between the data segments, including: Analyzing the length of the data segment and the change characteristics of the data in the data segment to obtain a stability factor of each data segment; For each dimension of environmental data, the inter-segment distances between the data segments are analyzed, and combined with the stability factor, an adjusted distance between different data segments in each dimension of environmental data is obtained.
[0008] In some embodiments of the present invention, analyzing the length of the data segment and the variation characteristics of the data within the data segment to obtain the stability factor of each data segment includes: Counting the amount of environmental data in the data segment to obtain the length of the data segment; Analyzing the fluctuation characteristics of the environmental data within the data segment and the overall distribution characteristics of the environmental data to obtain the variation characteristics of the data within the data segment; The stability factor of each data segment is obtained by combining the length and the change characteristic.
[0009] In some embodiments of the present invention, merging the data segments to obtain a merged cluster of each dimension of environmental data includes: Merging the data segments with the closest distances according to the adjusted distance to obtain a number of initial merged clusters of the environmental data of each dimension; Determining the number of representative points of the environmental data in the merged cluster according to the minimum value of the number of environmental data in all the data segments; Determining whether the amount of environmental data in the initial merged cluster is greater than or equal to 5 times the amount of representative points; If not, calculate the mean of the stability factors of the data segments in each of the initial merged clusters, combine the inter-cluster distances of the initial merged clusters to obtain the adjusted distance, merge the initial merged clusters to obtain a merged cluster of each dimensional environment data.
[0010] In some embodiments of the present invention, analyzing the correlation between the data segment and the data segments corresponding to each other dimension of environmental data to obtain a correlation description vector for each data segment, and then obtaining the adjusted intra-cluster distance between any two environmental data in the merged cluster, includes: For any one of the data segments, analyzing the correlation between the data segment and the data segments corresponding to each of the other dimensional environmental data, sorting the multiple correlations, and obtaining a correlation description vector for each of the data segments; The similarity between the related description vectors of the data segments where any two environmental data in the merged cluster are located is analyzed, and the adjusted intra-cluster distance between any two environmental data in the merged cluster is obtained in combination with the numerical values of the environmental data.
[0011] In some embodiments of the present invention, analyzing the trend difference between the representative point and the centroid of the merged cluster and combining the adjustment of the intra-cluster distance to obtain the shrinkage factor of the representative point includes: respectively obtaining reference data segments of representative points and centroids of the merged cluster; Analyzing the distance relationship between the reference data segment of each representative point in the merged cluster and the reference data segment corresponding to the centroid to obtain the initial trend difference between the reference data segment of the representative point of the merged cluster and the reference data segment of the centroid; Analyze the data difference between the representative point and other data in the corresponding reference data segment to obtain the trend confidence; Combining the initial trend difference degree and the trend confidence, obtaining the trend difference between the representative point and the centroid of the merged cluster; According to the trend difference and in combination with the adjusted intra-cluster distance between the representative point and the centroid, a shrinkage factor of the representative point is obtained.
[0012] In some embodiments of the present invention, merging the merged clusters to obtain abnormal data includes: Shrinking the representative points of the merged cluster according to the shrinkage factor; Calculating the mean stability factor of the data segments in each merged cluster, combining the inter-cluster distances of the two merged clusters, and iteratively merging the merged clusters; Based on the merging of the merged clusters, outliers are screened to obtain abnormal data.
[0013] According to a second aspect of an embodiment of the present application, a municipal underground pipeline inspection device is provided, wherein the technical solution adopted is as follows: the device includes: a memory and a processor, wherein: The memory is used to store program code; The processor is used to read the program code stored in the memory and execute the municipal underground pipeline inspection method as described in the first aspect of the embodiment of the present application.
[0014] In some embodiments of the present invention, the processor includes: Environmental data acquisition module, used to collect multi-dimensional environmental data in municipal underground pipelines; The data segmentation module is used to analyze the correlation changes between environmental data of different dimensions and segment the single-dimensional environmental data to obtain several data segments of each dimension of environmental data; A data segment merging module is used to analyze the length of each data segment and the change characteristics of the data in the data segment for each dimension of environmental data, combine the inter-segment distances between the data segments, obtain the adjusted distances between different data segments in each dimension of environmental data, merge the data segments, and obtain a merged cluster for each dimension of environmental data; a representative point acquisition module, configured to analyze the correlation between the data segment and the data segments corresponding to each other dimensional environmental data, obtain a correlation description vector for each data segment, and further obtain an adjusted intra-cluster distance between any two environmental data in the merged cluster, and select a representative point of the merged cluster; The abnormal data screening module is used to analyze the trend difference between the representative point and the centroid of the merged cluster, combine the distance within the adjusted cluster to obtain the shrinkage factor of the representative point, merge the merged cluster, and obtain abnormal data.
[0015] Compared with the existing technology, the municipal underground pipeline inspection equipment and method provided by the present invention have the following beneficial effects: The present invention segments the single-dimensional environmental data by analyzing the correlation changes between environmental data of different dimensions, and obtains several data segments of each dimensional environmental data; for each dimensional environmental data, the length of the data segment and the change characteristics of the data in the data segment are analyzed, and the distance between the segments is combined to obtain the adjusted distance between different data segments in each dimensional environmental data; in the hierarchical clustering process, the segments are used to replace the clusters of the initial part, and the data segments are merged to obtain the merged clusters of each dimensional environmental data. In the process of cluster merging, the influence of the collected data on the area and time period of the corresponding inspection robot is taken into account; at the same time, the data segments are combined with each other. The correlation changes between the data segments corresponding to the dimensional environmental data are described, so as to adjust the intra-cluster distance between any two environmental data in the merged cluster, and then select the representative point of the merged cluster. During the representative point selection process, the distance is adjusted according to the influence of the area and time period where the inspection robot is located, so that the representative point not only reflects the original data differences in the merged cluster, but also reflects the differences between the change trends; the trend difference between the representative point and the centroid of the merged cluster is analyzed, and combined with the adjustment of the intra-cluster distance, the shrinkage factor of the representative point is obtained, which ensures that the representative point can reflect the overall range of environmental characteristics in similar areas and time periods without being disturbed by abnormal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 A schematic diagram of the basic process of a municipal underground pipeline inspection method provided by one embodiment of the present invention; Figure 2 The present invention provides a schematic diagram of the composition of a municipal underground pipeline inspection device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, provides a detailed description of the specific implementation, structure, features, and effectiveness of a municipal underground pipeline inspection device and method according to the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. Terms such as "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a circuit structure, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such article or device. In the absence of further limitations, the phrase "comprising a ..." to define an element does not preclude the presence of other identical elements in the article or device comprising the element.
[0020] The specific scenario addressed by the embodiments of this invention is that during underground pipeline inspections, an inspection robot collects and analyzes multidimensional environmental data within the pipeline. However, because the inspection robot is in motion, it resides at different locations within the pipeline at different times, resulting in variations and trends in the environmental data across different dimensions. Therefore, the purpose of the embodiments of this invention is to address the issue of inaccurate anomaly analysis results based on fixed data thresholds, caused by differences in environmental data from different locations during underground pipeline inspections.
[0021] The specific scheme of the municipal underground pipeline inspection equipment and method provided by the present invention is described in detail below with reference to the accompanying drawings.
[0022] See also Figure 1 , which shows the basic process of a municipal underground pipeline inspection method provided by an embodiment of the present invention.
[0023] like Figure 1 As shown, an embodiment of the present invention provides a municipal underground pipeline inspection method, comprising: S100: Collect multi-dimensional environmental data inside municipal underground pipelines.
[0024] A rail-mounted inspection robot moves at a constant speed within municipal underground pipelines. Multiple sensors deployed on the robot collect multidimensional environmental data in real time. This data includes temperature, humidity, air pressure, and the concentrations of various gases, including oxygen, carbon monoxide, hydrogen sulfide, and methane. Each dimension of environmental data is collected by its corresponding sensor at a uniform 10-second interval. The collected data is then standardized to generate the final dimensional environmental data.
[0025] At this point, multi-dimensional environmental data within the municipal underground pipeline has been obtained, which is convenient for subsequent analysis.
[0026] S200: Analyze the correlation changes between environmental data of different dimensions, segment the single-dimensional environmental data, and obtain several data segments of environmental data of each dimension.
[0027] As the inspection robot moves through underground pipelines, it collects multidimensional environmental data from different areas. For environmental data of the same dimension within the same area, the data varies little, remaining within a certain range and exhibiting a relatively fixed trend. However, as the inspection robot's movement causes regional changes, the data for the same dimension within different areas will change, and the overall trend will also differ from that of the previous area. For environmental data of different dimensions within the same area, since the overall trend of each dimension is fixed, the correlation between these data does not change significantly over time. If a change occurs, i.e., an inflection point in the correlation occurs, the trend will change, indicating a possible change in the area. For example, humidity and temperature are relatively fixed environmental parameters within the same area, but regional changes can cause significant changes in the data of both areas, leading to a change in the trend correlation.
[0028] Taking the environmental data of any dimension as the target dimension environmental data, the target dimension environmental data will cause the data itself and the change trend to change due to the change of the position of the inspection robot where the sensor is located. Therefore, the target dimension environmental data needs to be segmented to ensure that the difference between the data in the same data segment is small, and the correlation with the change trend between the environmental data of other dimensions is similar, that is, the correlation between the environmental data of different dimensions is in the same change direction, then the corresponding inspection robots may be in the same area during this period.
[0029] Based on the above analysis, in an embodiment of the present invention, by analyzing the correlation changes between different dimensional environmental data, the single-dimensional environmental data is segmented to obtain several data segments of each dimensional environmental data. First, the Pearson correlation coefficient of the target dimension environmental data and any other dimension environmental data is calculated one by one starting from the first environmental data, and the signs of the difference between the Pearson correlation coefficients of the two adjacent target dimension environmental data (the Pearson correlation coefficient corresponding to the latter target dimension environmental data minus the Pearson correlation coefficient corresponding to the former target dimension environmental data) are recorded. When the sign changes, an inflection point occurs, and the inflection point is obtained; and the inflection point is used as the truncation point, that is, the former target dimension environmental data of the two target dimension environmental data corresponding to the sign change is used as the truncation point to obtain the initial data segment corresponding to the target dimension environmental data and each other dimension environmental data.
[0030] Then, the initial data segment with the smallest segment length among multiple initial data segments (the target dimension environment data and each other dimension environment data each obtain an initial data segment, so the target dimension environment data corresponds to multiple initial data segments) is used as the first data segment of the target dimension environment data.
[0031] Then, in the target dimension environment data, taking the first environment data after the first data segment as the starting point, re-acquire the initial data segments corresponding to the target dimension environment data and each dimension environment data, and filter to obtain the second data segment of the target dimension environment data; and so on, obtain several data segments corresponding to each dimension environment data.
[0032] At this point, several data segments of the target dimension environmental data have been obtained. The environmental data in the same data segment is most likely collected by the inspection robot in the same area, its change trend has not changed significantly, and the data itself has little difference.
[0033] S300: For each dimension of environmental data, analyze the length of the data segment and the change characteristics of the data within the data segment, combine the inter-segment distance between the data segments, obtain the adjusted distance between different data segments in each dimension of environmental data, merge the data segments, and obtain a merged cluster of each dimension of environmental data.
[0034] In both cure and traditional hierarchical clustering, initial clusters need to be merged during the initial clustering process until a representative cluster point can be selected. However, there is a difference between cure and hierarchical clustering. Traditional hierarchical clustering only divides and merges initial clusters based on differences in the data itself, while cure clustering replaces initial and merged clusters with data segments, avoiding the traditional clustering approach of ignoring the changing characteristics of the region corresponding to the data collection period. Therefore, it is necessary to merge data segments to obtain merged clusters for each dimension of environmental data.
[0035] For each data segment, the length of the data segment, the difference between the data within the data segment, and the overall distribution of the data can all reflect the stability of the data segment. The longer the data segment and the smaller the data fluctuation, the more stable the data has been for a long time. The kurtosis can reflect the overall concentrated distribution of the data. The more concentrated the data distribution is, the more stable the data distribution is and the smaller the probability of data outliers is.
[0036] In the hierarchical clustering process, clusters need to be continuously merged. In the process of merging data segments as clusters, the cluster merging is judged based on the original calculation of the distance between clusters. That is, the smaller the distance between the data segments and the closer the stability factors are, the closer the stable states of the data segments are, and the values and change trends of the data within them are relatively similar. This may be environmental data collected by the inspection robot in the same area at different times. In the clustering process, the more they should be regarded as the same cluster, that is, the more they should be merged.
[0037] Corresponding to multidimensional environmental data, taking humidity data as an example, the purpose of CURE clustering of humidity data is to ultimately filter out abnormal data. In the process of obtaining data segments, the humidity data collected continuously in the same area corresponding to the data segment have small numerical differences and are easily clustered into the same cluster. For different data segments, the data segment contains several humidity data, and the difference between the humidity data in different data segments is used to quantify the distance between the data segments, that is, the inter-cluster distance. The smaller the difference in humidity data, and thus the smaller the inter-cluster distance between data segments, the data segments may correspond to similar time periods and areas, and the greater the possibility of merging the two data segments.
[0038] Based on the above analysis, in an embodiment of the present invention, for each dimension of environmental data, the length of the data segment and the variation characteristics of the data within the data segment are analyzed. Combined with the inter-segment distance between the data segments, the adjusted distance between different data segments in each dimension of environmental data is obtained, and the data segments are merged to obtain a merged cluster for each dimension of environmental data. For each dimension of environmental data, we analyze the length of the data segment and the change characteristics of the data within the data segment. Combined with the distance between the data segments, we can obtain the adjustment distance between different data segments in each dimension of environmental data. Specifically, we can obtain the following: First, analyze the length of the data segment and the change characteristics of the data in the data segment to obtain the stability factor of each data segment. The specific implementation method is: count the amount of environmental data in the data segment to obtain the length of the data segment; analyze the fluctuation characteristics of the environmental data in the data segment and the overall distribution characteristics of the environmental data to obtain the change characteristics of the data in the data segment; combine the length and change characteristics to obtain the stability factor of each data segment. Construct the first The calculation formula of the stability factor of a data segment is: Where, Indicates the The stability factor of each data segment; Indicates the The amount of environmental data in each data segment; Indicates the The variance of all environmental data in a data segment; Indicates the The standard deviation of all environmental data in a data segment; Indicates the The mean of all environmental data in a data segment; Indicates the In the data segment The value of the environmental data; Represents an infinite decimal greater than 0, used to prevent the denominator from being 0.
[0039] The amount of environmental data in the data segment The larger the value, the longer the length of the data segment, which means the longer the period corresponding to the data segment, that is, the more stable the environmental data is in a longer period of time, which means the greater the stability factor of the data segment; the variance of all environmental data in the data segment , reflects the fluctuation characteristics of environmental data within a data segment. The smaller the value, the smaller the fluctuation of environmental data within the corresponding data segment, indicating that the stability factor of the data segment is greater; It represents the kurtosis of the environmental data in the data segment, which can reflect the overall distribution characteristics of the environmental data in the data segment. The larger the value, the more concentrated the distribution of the environmental data in the data segment, the more stable the environmental data in the data segment, and the larger the stability factor of the data segment.
[0040] Then, for each dimension of environmental data, the inter-segment distance between data segments is analyzed, that is, the inter-cluster distance between data segments is calculated. This is an existing technology and will not be described here. Based on the inter-segment distance and the stability factor, the adjusted distance between different data segments in each dimension of environmental data is obtained. The calculation formula for the adjusted distance between different data segments in the dimensional environment data is: Where, Indicates the The data segment and Adjustment distance between data segments; Indicates the The data segment and The inter-segment distance between data segments (inter-cluster distance); Indicates the The stability factor of each data segment; Indicates the The stability factor of a data segment.
[0041] The inter-cluster distance between two data segments The smaller the value, the closer the stability factor ( The smaller the value is, the closer the stable state of the data segment is, and the values and change trends of the environmental data within it are relatively similar, indicating that the smaller the adjustment distance between the two data segments is, the more likely it is that the environmental data collected by the inspection robot in the same area at different times should be considered as the same cluster in the clustering process and should be merged.
[0042] In the inter-cluster merging process of cure clustering (hierarchical clustering), similar data segments need to be merged. This is reflected in humidity data. The environmental data of the same area in different time periods may be relatively similar. In this case, the corresponding different data segments can be merged. At the same time, if the environmental conditions in different areas are also relatively similar, the corresponding data segments can also be merged. The inter-cluster merging process is to gradually merge the humidity data collected under similar environments into the same cluster. Therefore, in an embodiment of the present invention, by merging data segments, a merged cluster of environmental data of each dimension is obtained, which specifically includes: First, according to the adjusted distance, the data segments with the closest distance are merged to obtain several initial merged clusters of the environmental data of each dimension.
[0043] Then, the number of representative points of the environmental data in the merged cluster is determined based on the minimum number of environmental data in all data segments (the minimum number of environmental data in the data segment may be obtained from outlier or mutant environmental data, and the number of representative points is determined in this way).
[0044] Then, since the number of environmental data in the merged cluster must be at least five times the number of representative points, the selection of representative points begins. Therefore, it is judged whether the number of environmental data in the initial merged cluster is greater than or equal to 5 times the number of representative points; if so, it means that the initial merged cluster has met the standard and representative points can be selected; if not, it means that the initial merged cluster does not meet the standard, and it is necessary to continue merging the initial merged cluster, that is, by calculating the mean of the stability factor of the data segment in each initial merged cluster, combining the inter-cluster distance of the initial merged cluster, obtaining the adjusted distance, merging the initial merged clusters, and finally obtaining a merged cluster with qualified environmental data in each dimension.
[0045] Taking humidity data as an example, the representative points reflect the morphological characteristics of the entire cluster. If the number of humidity environmental data in the cluster is too small, for example, there are only one or two similar humidity data segments, and the corresponding humidity characteristics are highly similar, the cluster characteristics reflected by the representative points are not accurate. The representative points are usually humidity data distributed within a certain range, but are very different from other humidity data. Considering the occurrence of local abnormal data, some data segments will be too short. Therefore, the number of representative points needs to be determined based on the shortest data segment to ensure that the representative points can reflect the overall range of environmental characteristics in similar areas and time periods without being interfered with by abnormal data.
[0046] S400: Analyze the correlation between the data segment and the data segments corresponding to each other dimension environment data to obtain the correlation description vector of each data segment, and then obtain the adjusted intra-cluster distance between any two environment data in the merged cluster, and select the representative point of the merged cluster.
[0047] The merging of data segments is based on the similarity relationship between single-dimensional environmental data (the similarity of environmental data and the similarity of corresponding time periods and regions), and the representative points need to be able to reflect the shape outline of the cluster. The environmental data corresponding to the representative points and the environmental data corresponding to the centroid, in addition to the environmental data reflected themselves, also need to be able to reflect the distribution of each data segment in the cluster, that is, the distribution of environmental data collected in different time periods and regions. Therefore, when quantifying the distance between the representative point and the centroid, it is also necessary to base it on the corresponding correlation of each data segment, that is, the changing trend between different-dimensional environmental data of multi-dimensional environmental characteristics in a specific region and time, as the measurement feature of the representative point, to reflect the numerical difference distribution of environmental data within the cluster, as well as the distribution of the region and time period of the environment corresponding to different data segments in the cluster.
[0048] Based on the above analysis, in an embodiment of the present invention, by analyzing the correlation between a data segment and the data segments corresponding to each other dimension of environmental data, a correlation description vector for each data segment is obtained, and then the adjusted intra-cluster distance between any two environmental data in the merged cluster is obtained, and the representative point of the merged cluster is selected. Where: Analyze the correlation between the data segment and the data segment corresponding to each other dimensional environmental data to obtain the correlation description vector of each data segment, and then obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster. Specifically, for any data segment, analyze the data segment corresponding to each other dimensional environmental data (that is, based on the data segment division of the target dimensional environmental data, divide the other dimensional environmental data in the same way to obtain several reference segments, and use the reference segments corresponding to the same position of any data segment of the target dimensional environmental data in all reference segments as the data segments corresponding to each other dimensional environmental data), that is, calculate the Pearson correlation coefficient between the last environmental data in the two data segments, sort the multiple correlations (which can be arranged in the order of collection of each dimensional environmental data), and obtain the relevant description vector of each data segment; analyze the similarity between the relevant description vectors of the data segments where any two environmental data in the merged cluster are located, and combine the numerical values of the environmental data to obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster. Construct the first Environmental data and The calculation formula for the adjusted intra-cluster distance between environmental data is: Where, Indicates the first Environmental data and Adjusted intra-cluster distance of environmental data; Indicates the The value of the environmental data; Indicates the The value of the environmental data; Indicates the The relevant description vector of the data segment where the environmental data is located; Indicates the The relevant description vector of the data segment where the environmental data is located; Indicates the The relevant description vector of the data segment where the environmental data is located is The cosine similarity between the relevant description vectors of the data segment where the environmental data are located.
[0049] Indicates the number before adjustment Environmental data and The distance between the environmental data, that is, the distance between the first Environmental data and The distance between the environmental data; The closer the value is to 1, the The relevant description vector of the data segment where the environmental data is located is The more similar the related description vectors of the data segments where the environmental data are located are, the more it is necessary to reduce the Environmental data and The intra-cluster distance of environmental data is used to reflect the changing trend between different dimensional environmental data of multidimensional environmental characteristics in a specific region and time, that is, Environmental data and The smaller the distance within the adjusted cluster of environmental data is; The closer the value is to 0, the The relevant description vector of the data segment where the environmental data is located is The less correlation there is between the related description vectors of the data segment where the environmental data is located, the Environmental data and The intra-cluster distance of each environmental data is the true distance; The closer the value is to -1, the The relevant description vector of the data segment where the environmental data is located is The more opposite the related description vectors of the data segment where the environmental data are located, the more the Environmental data and The intra-cluster distance of environmental data is used to reflect the changing trend between different dimensional environmental data of multidimensional environmental characteristics in a specific region and time, that is, Environmental data and The smaller the adjusted intra-cluster distance of the environmental data is.
[0050] Selecting a representative point of a merged cluster based on the adjusted intra-cluster distance between any two environmental data in the merged cluster is a prior art and will not be described in detail here.
[0051] At this point, the representative points in the merged clusters during the cure clustering process are obtained. During the cluster merging process, the influence of the collected data on the area where the corresponding inspection robot is located is taken into consideration. At the same time, during the representative point selection process, the intra-cluster distance is adjusted based on this influence, so that the representative points not only reflect the original data differences within the merged cluster, but also reflect the differences between the change trends.
[0052] S500: Analyze the trend difference between the representative point and the centroid of the merged cluster, adjust the distance within the cluster, obtain the shrinkage factor of the representative point, merge the merged cluster, and obtain abnormal data.
[0053] In the process of selecting representative points, the intra-cluster distance between environmental data is adjusted through the relevant description vector. This is based on the similarity between the correlation changes of multidimensional environmental data between data segments, that is, time periods. In other words, the representative points are selected based on the similarity between the time periods and the corresponding areas. The lower the similarity, the farther the representative point is from the centroid. For representative points that are farther away from the centroid, a larger shrinkage factor is required. At this time, it is necessary to analyze the data change trend of the data segment where the representative point itself is located. On the basis of the original intra-cluster distance, the more similar the change trend is, the smaller the shrinkage factor is. The greater the difference in the change trend is, the less the representative point can reflect the shape outline of the merged cluster, and a larger shrinkage factor is required.
[0054] Based on the above analysis, in an embodiment of the present invention, by analyzing the trend difference between the representative points and the centroid of the merged clusters, combined with adjusting the distance within the cluster, the shrinkage factor of the representative points is obtained, and the merged clusters are merged to obtain abnormal data. The representative points themselves need to reflect the cluster's shape distribution, while the shrinkage factor is needed to prevent abnormal outliers from interfering with the cluster's shape distribution. Based on the intra-cluster distance, the local data segment corresponding to the representative point needs to be analyzed. Abnormal data can significantly change the trend of change. If a single piece of environmental data differs significantly from other environmental data within the segment, it may not be abnormal data but error data, which is used to determine the shrinkage factor. Therefore, a confidence level is introduced for the similarity between trend changes. The greater the difference between the representative point and the data in the data segment, the less credible the corresponding trend is, and the smaller the confidence level.
[0055] Therefore, the trend difference between the representative point and the centroid of the merged cluster is analyzed, and the shrinkage factor of the representative point is obtained in combination with the adjusted intra-cluster distance. The specific implementation method is as follows: first, the reference data segments of the representative point and the centroid of the merged cluster are obtained respectively, that is, 5 environmental data before and after any representative point or centroid are obtained to constitute the reference data segment of the representative point or centroid (within the same data segment, and no more data will be obtained if the data exceeds the same data segment); then, the distance relationship between the reference data segment of each representative point in the merged cluster and the reference data segment corresponding to the centroid is analyzed to obtain the initial trend difference degree between the reference data segment of the representative point of the merged cluster and the reference data segment of the centroid; and, the data difference between the representative point and other data in the corresponding reference data segment is analyzed to obtain the trend confidence; then, the trend difference between the representative point and the centroid of the merged cluster is obtained in combination with the initial trend difference degree and the trend confidence; finally, the shrinkage factor of the representative point is obtained based on the trend difference and the adjusted intra-cluster distance between the representative point and the centroid.
[0056] Construct any merged cluster The calculation formula for the trend difference between the reference data segment of the representative points and the reference data segment of the centroid is: Where, Indicates the The trend difference factor (i.e., trend difference) between the reference data segment of the representative points and the reference data segment of the centroid; Indicates the Reference data segments of representative points Reference data segment with centroid The DTW distance between them; Indicates the The number of environmental data in the reference data segment of each representative point (excluding representative points); Indicates the The values of the environmental data corresponding to the representative points; Indicates the The reference data segment of the representative points The value of the reference data; Represents an exponential function with the natural constant e as the base, used as a linear normalization function.
[0057] The larger the DTW distance value between the reference data segment of the representative point and the reference data segment of the centroid, the greater the difference between the two reference data segments, and the greater the trend difference factor between the two reference data segments; It represents the difference of environmental data in the reference data segment corresponding to the representative point. As the trend confidence, the larger the value, the greater the difference between the representative point and other environmental data in its reference data segment. The larger the trend confidence, the greater the trend difference factor between the two reference data segments.
[0058] Then, construct the first The calculation formula for the shrinkage factor of each representative point is: Where, Indicates the first The shrinkage factor of each representative point; Indicates the The intra-cluster distance between each representative point and the centroid; Indicates the The trend difference factor between the reference data segment of the representative points and the reference data segment of the centroid; represents the linear normalization function.
[0059] Merge the merged clusters to obtain abnormal data. The specific implementation method is as follows: shrink the representative points of the merged cluster according to the shrinkage factor; calculate the mean stability factor of the data segment in each merged cluster, combine the inter-cluster distance of the two merged clusters (the distance between the representative points of the two merged clusters), and iteratively merge the merged clusters; screen outliers based on the merging of the merged clusters, that is, after the merging reaches a certain process, screen out outliers that are not in any merged cluster to obtain abnormal data. This is the specific implementation process of cure clustering in the existing technology and will not be repeated here.
[0060] During the real-time inspection process of the inspection robot, the environmental data collected in the target dimension is merged into clusters according to the above-mentioned data segment and cluster merging method, and anomaly judgment is completed to realize the inspection of municipal underground pipelines and anomaly analysis of multi-dimensional environmental data.
[0061] Based on the same inventive concept as the above method, this embodiment also provides a municipal underground pipeline inspection device.
[0062] Figure 2 A schematic diagram of a municipal underground pipeline inspection device provided by an embodiment of the present invention is shown as follows: Figure 2As shown, the device includes: a memory 10 and a processor 20, wherein: the memory 10 is used to store program code; the processor 20 is used to read the program code stored in the memory 10 and execute the collection of multi-dimensional environmental data in the underground pipeline; analyze the correlation changes between environmental data of different dimensions, segment the single-dimensional environmental data, and obtain several data segments of each dimensional environmental data; for each dimensional environmental data, analyze the length of the data segment and the change characteristics of the data in the data segment, combine the inter-segment distance between the data segments, obtain the adjusted distance between different data segments in each dimensional environmental data, merge the data segments, and obtain a merged cluster of each dimensional environmental data; analyze the correlation between the data segment and the data segments corresponding to each other dimensional environmental data, obtain the correlation description vector of each data segment, and then obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster, and select the representative point of the merged cluster; analyze the trend difference between the representative point and the centroid of the merged cluster, combine the adjusted intra-cluster distance to obtain the shrinkage factor of the representative point, merge the merged cluster, and obtain abnormal data.
[0063] Furthermore, the processor 20 includes an environmental data acquisition module 21, a data segmentation module 22, a data segment merging module 23, a representative point acquisition module 24 and an abnormal data screening module 25. Environmental data collection module 21, used to collect multi-dimensional environmental data in municipal underground pipelines; The data segmentation module 22 is used to analyze the correlation changes between the environmental data of different dimensions and segment the single-dimensional environmental data to obtain a number of data segments of each dimensional environmental data; The data segment merging module 23 is used to analyze the length of the data segment and the change characteristics of the data in the data segment for each dimension of environmental data, and combine the inter-segment distances between the data segments to obtain the adjusted distances between different data segments in each dimension of environmental data, and merge the data segments to obtain a merged cluster for each dimension of environmental data; The representative point acquisition module 24 is used to analyze the correlation between the data segment and the data segments corresponding to each other dimension of environmental data, obtain the correlation description vector of each data segment, and then obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster, and select the representative point of the merged cluster; The abnormal data screening module 25 is used to analyze the trend difference between the representative point and the centroid of the merged cluster, and adjust the distance within the cluster to obtain the shrinkage factor of the representative point, merge the merged clusters, and obtain abnormal data.
[0064] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0065] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A municipal underground pipeline inspection method, characterized in that: The method comprises: Collect multi-dimensional environmental data inside underground pipelines; Analyze the correlation changes between environmental data of different dimensions, segment the single-dimensional environmental data, and obtain several data segments of environmental data of each dimension; For each dimensional environmental data, the length of the data segment and the change characteristics of the data in the data segment are analyzed, and the adjusted distance between different data segments in each dimensional environmental data is obtained by combining the inter-segment distance between the data segments. The data segments are merged to obtain a merged cluster of each dimensional environmental data; Analyze the correlation between the data segment and the data segments corresponding to each other dimensional environmental data to obtain a correlation description vector for each data segment, and then obtain the adjusted intra-cluster distance between any two environmental data in the merged cluster, and select a representative point of the merged cluster; The trend difference between the representative point and the centroid of the merged cluster is analyzed, and the shrinkage factor of the representative point is obtained in combination with the adjustment of the intra-cluster distance, and the merged clusters are merged to obtain abnormal data.
2. The municipal underground pipeline inspection method according to claim 1, characterized in that: Analyze the correlation changes between environmental data of different dimensions, segment the single-dimensional environmental data, and obtain several data segments corresponding to each dimensional environmental data, including: Calculate the Pearson correlation coefficient of the target dimension environment data and any other dimension environment data one by one starting from the first environment data, and record the sign of the difference between the Pearson correlation coefficients of two adjacent target dimension environment data to obtain the inflection point; Using the inflection point as a cutoff point, obtaining initial data segments corresponding to the target dimensional environment data and each other dimensional environment data; Using the initial data segment with the smallest segment length among the multiple initial data segments as the first data segment of the target dimension environment data; In the target dimension environment data, the environment data after the first data segment is used as a starting point to obtain the second data segment of the target dimension environment data; And so on, several data segments corresponding to each dimension of environmental data are obtained.
3. The municipal underground pipeline inspection method according to claim 1, characterized in that: For each dimension of environmental data, the length of the data segment and the change characteristics of the data in the data segment are analyzed, and the distance between the data segments is combined to obtain the adjustment distance between different data segments in each dimension of environmental data, including: Analyzing the length of the data segment and the change characteristics of the data in the data segment to obtain a stability factor of each data segment; For each dimension of environmental data, the inter-segment distances between the data segments are analyzed, and combined with the stability factor, the adjusted distances between different data segments in each dimension of environmental data are obtained.
4. The municipal underground pipeline inspection method according to claim 3, characterized in that: Analyzing the length of the data segment and the change characteristics of the data in the data segment to obtain the stability factor of each data segment includes: Counting the amount of environmental data in the data segment to obtain the length of the data segment; Analyzing the fluctuation characteristics of the environmental data within the data segment and the overall distribution characteristics of the environmental data to obtain the variation characteristics of the data within the data segment; The stability factor of each data segment is obtained by combining the length and the change characteristic.
5. The municipal underground pipeline inspection method according to claim 1, characterized in that: The data segments are merged to obtain a merged cluster of environmental data for each dimension, including: Merging the data segments with the closest distances according to the adjusted distance to obtain a number of initial merged clusters of the environmental data of each dimension; Determining the number of representative points of the environmental data in the merged cluster according to the minimum value of the number of environmental data in all the data segments; Determining whether the amount of environmental data in the initial merged cluster is greater than or equal to 5 times the amount of representative points; If not, calculate the mean of the stability factors of the data segments in each of the initial merged clusters, combine the inter-cluster distances of the initial merged clusters to obtain the adjusted distance, merge the initial merged clusters to obtain a merged cluster of each dimensional environment data.
6. The municipal underground pipeline inspection method according to claim 1, characterized in that: Analyzing the correlation between the data segment and the data segments corresponding to each other dimensional environmental data to obtain a correlation description vector for each data segment, and then obtaining an adjusted intra-cluster distance between any two environmental data in the merged cluster, including: For any one of the data segments, analyzing the correlation between the data segment and the data segments corresponding to each of the other dimensional environmental data, sorting the multiple correlations, and obtaining a correlation description vector for each of the data segments; The similarity between the related description vectors of the data segments where any two environmental data in the merged cluster are located is analyzed, and the adjusted intra-cluster distance between any two environmental data in the merged cluster is obtained in combination with the numerical values of the environmental data.
7. The municipal underground pipeline inspection method according to claim 1, characterized in that: Analyzing the trend difference between the representative point and the centroid of the merged cluster, and combining the adjustment of the intra-cluster distance to obtain the shrinkage factor of the representative point, including: respectively obtaining reference data segments of representative points and centroids of the merged cluster; Analyzing the distance relationship between the reference data segment of each representative point in the merged cluster and the reference data segment corresponding to the centroid to obtain the initial trend difference between the reference data segment of the representative point of the merged cluster and the reference data segment of the centroid; Analyze the data difference between the representative point and other data in the corresponding reference data segment to obtain the trend confidence; Combining the initial trend difference degree and the trend confidence, obtaining the trend difference between the representative point and the centroid of the merged cluster; According to the trend difference and in combination with the adjusted intra-cluster distance between the representative point and the centroid, a shrinkage factor of the representative point is obtained.
8. The municipal underground pipeline inspection method according to claim 1, characterized in that: Merging the merged clusters to obtain abnormal data includes: Shrinking the representative points of the merged cluster according to the shrinkage factor; Calculating the mean stability factor of the data segments in each merged cluster, combining the inter-cluster distances of the two merged clusters, and iteratively merging the merged clusters; Based on the merging of the merged clusters, outliers are screened to obtain abnormal data.
9. A municipal underground pipeline inspection device, characterized in that: The device comprises: a memory and a processor, wherein: The memory is used to store program code; The processor is configured to read the program code stored in the memory and execute the municipal underground pipeline inspection method according to any one of claims 1 to 8.
10. The municipal underground pipeline inspection equipment according to claim 9, characterized in that: The processor includes: Environmental data acquisition module, used to collect multi-dimensional environmental data in municipal underground pipelines; The data segmentation module is used to analyze the correlation changes between environmental data of different dimensions and segment the single-dimensional environmental data to obtain several data segments of each dimension of environmental data; A data segment merging module is used to analyze the length of each data segment and the change characteristics of the data in the data segment for each dimension of environmental data, combine the inter-segment distances between the data segments, obtain the adjusted distances between different data segments in each dimension of environmental data, merge the data segments, and obtain a merged cluster for each dimension of environmental data; a representative point acquisition module, configured to analyze the correlation between the data segment and the data segments corresponding to each other dimensional environmental data, obtain a correlation description vector for each data segment, and further obtain an adjusted intra-cluster distance between any two environmental data in the merged cluster, and select a representative point of the merged cluster; The abnormal data screening module is used to analyze the trend difference between the representative point and the centroid of the merged cluster, combine the distance within the adjusted cluster to obtain the shrinkage factor of the representative point, merge the merged cluster, and obtain abnormal data.
Citation Information
Patent Citations
Pipeline abnormity confirmation method and device and computer readable storage medium
CN112949697A
Data anomaly detection method and device and computer equipment
CN115270986A
Abnormal data analysis method and system for automatic environment monitoring equipment
CN115271003A
Programmed medium for clustering large databases
US6092072A
Abnormal data detection method and apparatus, and storage medium
WO2025016349A1